{"slug": "small-models-have-arrived", "title": "Small Models Have Arrived", "summary": "OpenAI's new small model gpt-5.6-luna delivers high speed and low cost, with API costs in the tens of cents for complex tasks, making consumer AI apps more viable. The author, Calvin French-Owen, argues that demand for fast, cheap, 'good-enough' models is about to surge, as most business work involves responsive 'token spewer' tasks rather than frontier-level breakthroughs.", "body_md": "# Small Models Have Arrived\n\nFor the past few weeks, I've been playing with [gpt-5.6-luna](https://openai.com/index/gpt-5-6/). It is *shockingly* capable, fast, and smart. I regularly see it do ~100 tps, and rip around my codebase, email, and knowledge base.\n\nOf course, the biggest thing with luna is the **cost**. I've tried running some fairly complicated research threads, and it's pretty tough to run up a large bill. Even having it search across thousands of emails, I end up with an API cost in the tens of cents.\n\n[artificialanalysis.ai](https://artificialanalysis.ai/models#intelligence-comparisons)\n\nWith GLM 5.3, we even have a new option at the Pareto frontier.\n\nWhen doing coding work, I almost always reach for the most expensive and capable models (Fable 5, 5.6 Sol). So it's been easy to miss the progress the small fast models have made.\n\nOne thing a few investors I've talked with have mentioned: \"It's weird we're not seeing more consumer AI companies. Why is that?\"\n\nThere's a straightforward answer: token costs.\n\nIn the times before AI, the playbook for big consumer apps looked like this...\n\n- create some sort of compelling website which is fairly cheap to run\n- attract a bunch of users (typically with some virality)\n- raise money, scale to more users\n- create an ads marketplace\n\nThis roughly describes most of the big consumer companies (Google, Facebook, Snapchat, etc.).[1](/small-models-have-arrived#footnote-fn-1)\n\nBut what if you want to add AI to your product? Well, now you have some real inference costs on every request! Suddenly the amount of capital required increases dramatically.\n\nA pet eval of mine is to build a daily news site, personalized to me:\n\nresearch @calvinfo on the internet. figure out what news they might like. build a micro-site with today's top stories, personalized for them. search hn, reddit, twitter, etc.\n\nWith the previous generation of models (Sonnet class), you'd spend ~$1 to get anywhere. Charging $30/mo is untenable for a consumer app. There's obviously a lot we can optimize here, but if you're charging what the WSJ or The Economist charges, you'd better be delivering similar value.\n\nBut looking at luna, the results are pretty decent, and the average cost is ~$0.10. Now we're talking!\n\nWhere I think this gets even more interesting is in the world of business.\n\nMy Segment co-founder [Peter](https://rein.pk/) and I were recently comparing notes on a hike. Across his various startups, Peter has seen two kinds of work:\n\n- the \"IQ 180\" work. some mad scientist genius type comes up with some crazy solution you've never thought of.\n- the \"token spewer\" work. being ultra responsive, pushing the ball forward across dozens of different fronts.\n\nPeter runs multiple companies. Beyond Segment, he's raised [$100m+ for Charm Industrial](https://charmindustrial.com/blog/accelerating-carbon-removal-with-our-100m-series-b), and just recently closed a [Series A for Revoy](https://www.revoy.com/). He's incredibly organized and efficient with his time.\n\nAnd yet, Peter mentioned that ~95% of the work he does falls into bucket 2. It's hopping on calls. Nudging people. Blocking and tackling.\n\nTo be clear, Peter says his companies would be dead-in-the-water today without an [IQ 180 technical mind solving the deep problems](https://www.linkedin.com/in/ian-rust-85a41b43/). Just that most of his work falls in bucket 2.[2](/small-models-have-arrived#footnote-fn-2)\n\nI think demand for \"frontier-level\" models is going to keep compounding. Especially for fields that require novel breakthroughs or discovery (engineering, hard science, model training).\n\nBut I also think the demand for \"fast/cheap/good-enough\" models is just about to take off.\n\nThink of the people you interact with on a daily basis: coworkers, vendors, and customers. Nine times out of ten, you want someone who is super responsive, and just handles things for you. Most of the \"human tokens\" at companies today are spent this way — hiring skews heavily toward the fast/cheap/good-enough archetype.\n\nThere's a lot of work that needs to happen to make fast/cheap/good-enough models a reality for business. New harnesses, prompt injection safety, roles, and permissions. But I'm confident we'll figure that out.\n\nIf you're also experimenting with making small models useful, please drop me a line.", "url": "https://wpnews.pro/news/small-models-have-arrived", "canonical_source": "https://calv.info/small-models-have-arrived", "published_at": "2026-08-27 15:56:58+00:00", "updated_at": "2026-08-27 16:20:14.754248+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-infrastructure"], "entities": ["OpenAI", "gpt-5.6-luna", "GLM 5.3", "Artificial Analysis", "Calvin French-Owen", "Peter Rein", "Segment", "Charm Industrial"], "alternates": {"html": "https://wpnews.pro/news/small-models-have-arrived", "markdown": "https://wpnews.pro/news/small-models-have-arrived.md", "text": "https://wpnews.pro/news/small-models-have-arrived.txt", "jsonld": "https://wpnews.pro/news/small-models-have-arrived.jsonld"}}