AMD Advancing AI 2026: Software, Hardware & Framework, Unified AMD's Advancing AI 2026 summit in San Francisco featured a panel with Chris Lattner, Ramin Hasani, and Hassan Akbari. Lattner, creator of LLVM, called AI 'mid,' arguing that the current AI surface layer is a distribution follower, while the real work happens in hardware and compute. He advocated for portable alternatives to CUDA, such as his Mojo language at Modular, to let hardware express its capabilities without vendor lock-in. Where Chris Lattner calls AI "mid" and George Hotz wants to knock a trillion dollars of value off of NVIDIA with tinygrad. IMHO: AMD's Advancing AI 2026 https://www.amd.com/en/corporate/events/advancing-ai.html summit in San Francisco was free. And right now, while the firehose of free AI events and learning is still on, I'm taking every opportunity I can get my hands on. Eventually the free opportunities will slowly disappear, and this age of companies wanting us to adopt their tech and ecosystem will end. Cynical, pessimistic maybe. But think about the money spent on a two-day event at the Moscone Center. Free if you sign up. Free learning, access to vendors, tech talks, keynote speakers, food and lunch provided. That's an investment, and eventually the accountants, or the shareholders, are going to need to see a return. But I digress. This is my rundown of day one. To say I'm a fish out of water at any event is an understatement. AWS, MLH, Google, and now AMD, I've tackled each one as a way to explore this brave new world, but sometimes I just don't have the context. And that's okay. AMD was more technical. More focused on hardware. There was more than one robot. And it changed my mental model of technical ecosystems and the symbiosis within them, and I'm so here for it. Day one started with a three-person panel: Chris Lattner, Ramin Hasani, and Hassan Akbari, hosted by, checks notes, a name I didn't catch. There was no real theme. It was a roundtable discussion, and here's what came out of it. Chris Lattner created LLVM https://llvm.org , the compiler infrastructure a huge amount of modern software quietly runs on. So when he talks, I try to keep up. His point wasn't only about software. It was about how the hardware we run works with the applications we write, and why you can't reason about one without the other. He said he'd been told his approach wouldn't work, because CUDA https://developer.nvidia.com/cuda-zone has the moat. His answer, the way I understood it: CUDA is kind of like GCC. GCC is the decades-old open-source compiler a lot of software gets built with, so, big and foundational and hard to dislodge. A monolith. Powerful, legacy, wrapped in an incredible moat, and also not that good. His point: you don't beat the moat head-on, you go around it with architecture. Let the hardware express what it can do instead of locking everyone into one vendor's language. That's his whole Mojo https://www.modular.com/mojo play at Modular https://www.modular.com , for what it's worth, a portable alternative to CUDA. Then he dropped the line. AI is mid. Appropriate laughter followed. But I don't think he meant it as a burn. The way I read it, he meant AI is a distribution follower. The AI most of us touch every day, the LLMs, the predictive models, those are the product. They ride on top of the work happening behind the scenes: the training, the hardware, the compute. He lives in that granular layer. Most of us live in the public one. So when he says mid, I think he's talking about the surface, not the machine underneath it. Maybe I'm getting his intent wrong. I'm allowed. But that's the read I walked away with, and it stuck. Ramin Hasani talked in abstractions I didn't fully follow. And I mean abstract for me, not abstract in general. The concepts were above where I am right now. But I wrote them down anyway, because future me is going to want them. Here's what I caught, in the order my pen caught it. Think about the problem at the algorithmic level before you reach for kernel optimizations. AI designing AI, and not just attention, look at the other transformers. Liquid foundation models, which are his company Liquid AI https://www.liquid.ai 's whole thing, so that one I could at least place. Matching the best models to the right hardware: CPU, NPU, GPU. I'm not going to pretend I can explain all of that yet. I can't. But it's in the notebook for when I get to the "kernel" level of my studies. Lol. Hassan Akbari https://www.linkedin.com/in/hassan-akbari-48a1b270/ was the one who tied it together. He talked about a unified ecosystem between frameworks, hardware, and software. He mentioned kernel again, and no, I still didn't catch the context, so onto the pile it goes with Ramin's. His point was that the interconnectivity of all three has to be leveraged for optimization. You don't tune one piece in isolation. You use how they connect. Then he asked the question that stuck with me. Large models can now be distilled into smaller ones that are still effective. So are we wasting compute? Scaling models, the way Hassan framed it, should be pragmatic. That one made sense to me. His other points came from the customer side, where reliability, cost per token, and accuracy are what matter. Optimization is the whole goal, and evaluation pipelines are how you measure it: benchmarks, deployments, the numbers. At least that's what I wrote down. Correct me if I mangled it. Picture the stereotypical genius from Silicon Valley, the show. That's George Hotz. It was awesome, it was inspiring, and it was very on the nose. Note to self for the day: I am sharing a room with very, very, very smart people. Even his intro was wild. Jailbroke the iPhone. Reverse-engineered the PlayStation 3. Made a self-driving car , comma.ai https://comma.ai , that got a cease-and-desist letter from the DMV. Got hired to hack a company's own systems to keep them safer. I could not tell if I was watching an inspirational documentary or a cautionary afterschool special. But I already said he was awesome, and I stand by it. His talk was deep. Graphs, lines of code, evaluations. I have pictures, if anyone wants them. At one point he said the Radeon RX 7900 XTX https://en.wikipedia.org/wiki/Radeon RX 7000 series , three years old now, is still a good value at $999. I wrote down "what is a 7900 XTX?" and filed it for later. Back to it. George didn't hold back, and the candor was refreshing. tinygrad https://github.com/tinygrad/tinygrad , which I looked up, is an open-source neural network framework and deep learning library, and on stage he put it at 25,000 lines of pure Python, about 9,000 of that the core, meaning the essential engine with everything else built around it. His stated goal, and I quote, is to "commoditize the petaflop." Translation, delivered deadpan: if tinygrad succeeds, it knocks a trillion dollars of value off NVIDIA https://www.nvidia.com . Insert audience laughter here. Same moat Chris Lattner was talking about. George just comes at it from the scrappy open-source end. His pitch is GPUs for the middle class. He talked about growing up in New Jersey with a parent who was a teacher. He doesn't care about data centers. What he cares about is the highest development velocity, and not the out-of-the-box-fast kind. His argument, the way I understood it: an out-of-the-box TensorFlow https://www.tensorflow.org might be faster on day one, but tinygrad's velocity curve is built to climb higher and faster over time. And here's the part that got me. tinygrad has no dependencies. None. No numpy was used in the making of tinygrad. Which I learned is the whole flex. No dependencies means nothing underneath you can break, bloat, or drift out of version. Fewer moving parts, less that can rot. Very much my kind of principle. He even mentioned running tinygrad as a backend, which I am absolutely going to go try. So the guy who reverse-engineered a PlayStation had my full attention. So glad I made it. Here's what kept pinballing around my head after all of that. I'm an AI-assisted builder. I live on the top-level software end. I write code, I deploy applications. So why haven't I been more thoughtful about what runs underneath all of it? My dev.to articles focus on the end product, the thing I shipped. Not the framework, not the hardware, not the compute it all rides on. Maybe that's the next stage of my development. Or maybe not. Here's my honest worry: if I try to hold every piece at once, framework and hardware and software and infra, I might get so overwhelmed I don't build anything at all. And then the comforting thought. Maybe the three of them were mostly talking about training models, and I can go back to coding with vibes. The event leaned hard into CPU, GPU, NPU, compute, and infrastructure, not so much the SaaS layer I live in. So maybe it isn't my fight yet. Something to ponder. But that pinball is the whole reason I go. For a few hours, something pulled me out of my product-focused viewpoint and made me look at the layer underneath. That's the value. Not the swag bag, necessarily After the opening session I hit the workshops. One was Build Your OpenClaw Agent with Multi-Modal Models, running on AMD GPUs. Another was billed as vibecoding with local models. Good practice, all of it technology I hadn't touched. AMD's learning platforms were new to me. Working out of Jupyter notebooks was new too. The vibecoding workshop, well, the title masqueraded it. It was a Lemonade https://github.com/lemonade-sdk/lemonade and Qwen https://github.com/QwenLM workshop, and it rocked. First time I'd heard of Lemonade. It's a community project, sponsored by AMD with optimizations from their engineers, that runs large language models locally on your own GPU or NPU. We ran Qwen3.6-35B-A3B on it. They walked the framework from the ground up: silicon, engines, routing, models, then the app. Another pyramid. And there I was again, looking down from my product endpoint at the top and seeing how much sits underneath. Same shift as the panel, twice in one day. This one got me excited though. Running Lemonade locally means everything stays with me. A new idea, a new system, a new output to go explore. I'm sure the learning curve will be steep. We were working inside pre-set-up AMD environments, and standing something up from scratch will bring its own challenges. But as always, excited and ready to try something new. Full disclosure: I did day one as a day trip, LAX to SFO and back. Day two happens without me, and I'm bummed to miss it. I'll have to watch Dr. Su's keynote https://www.youtube.com/live/jvtPC28nGsc?si=8PPIaCjFY0HQgmwI on Youtube. What I'm carrying home instead is the Lemonade itch. A from-scratch local setup on my own machine, which will have its own learning curve. I'll report back on how that goes. Until the next event I hope to see my Dev.to fam there, too AI Assisted. Human Approved. Powered by NLP.