Cheapest LLM APIs for Startups in 2026
A data science student built a calculator to compare LLM API costs after finding that agentic coding sessions on OpenRouter could vary from cents to dollars per loop. The cheapest paid models in mid-2…
A data science student built a calculator to compare LLM API costs after finding that agentic coding sessions on OpenRouter could vary from cents to dollars per loop. The cheapest paid models in mid-2…
Nous Research's open-source Hermes Agent, using Mixture of Agents presets, outperformed Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 on SWE-bench Pro benchmarks. The framework strings multiple lan…
Nous Research released Mixture of Agents presets as virtual models in Hermes Agent, allowing users to select multi-model workflows like any other model. The company claims its MoA presets outperform i…
Developer Scott Spencer released My-Pi Coding-Agent, a curated distribution of the Pi coding agent CLI with prewired extensions including MCP, LSP, skills, recall, redaction, telemetry, team mode, and…
Developers are adopting local proxy routers like Weave and 9Router to cut API costs and bypass the context re-read tax in AI coding assistants such as Claude Code and Cursor. These routers intercept A…
Weave released a smart model routing proxy that works with Claude, Codex, and Cursor, ranking #1 on the RouterArena leaderboard. The open-source tool uses an on-box embedder to select the best model p…
A developer built Aantraa, an AI-powered platform for audio and video translation, caption generation, and viral short clip creation, in one week. The platform relies heavily on AI APIs, using OpenRou…
Per-token prices for large language models are collapsing, but AI bills are exploding as reasoning models consume far more tokens per task. Uber burned through a year's AI budget in four months, and M…
AI search engine Exa raised $250 million in Series C funding at a $2.2 billion valuation, led by a16z, to power AI agents with high-quality web search. Exa already serves over 5,000 companies includin…
NeuralWatt, a US-based AI inference provider, introduced energy-based metering for LLM inference, charging by kilowatt-hour instead of per token. A user reported an average 82.9% cost reduction compar…
Jefferies reported that lower-cost AI models, such as Z.ai's GLM-5.2, are increasing demand for computing infrastructure rather than reducing AI investment, citing Jevons Paradox and rising inference …
Chinese AI models now process over three times the weekly token volume of US models on OpenRouter, handling 18 trillion tokens compared to 5.5 trillion. The shift began in February 2026 when Chinese m…
GitHub announced two harness-level improvements to Copilot agentic sessions—prompt caching achieving 94% cache hit rates and deferred tool loading—plus a new Auto model selection feature using its HyD…
Developer JoeBro released a native macOS AI workspace with a Python backend that requires zero dependencies, no pip install, and no Docker. The app runs entirely locally, storing data in a SQLite file…
OpenLanguage, a free and open-source AI language tutor for iOS, was released on Hacker News. The app uses voice I/O for speaking and listening practice, supports multiple AI providers via user API key…
Chinese open-source AI models from DeepSeek and Alibaba have rapidly closed the performance gap with US rivals, capturing over 50% of global open-source downloads and 30% of AI model usage by early 20…
Microsoft's MAI-Image-2.5 model ranks second in image editing and third in text-to-image generation on the Artificial Analysis Image Arena, outranking Google Gemini offerings and prior Microsoft model…
Chinese AI lab Zhipu AI, under its global brand Z.AI, launched the open-weight GLM-5.2 model with up to 753 billion parameters, rivaling top closed models from Anthropic and OpenAI at roughly one-tent…
A developer warns that building products on top of LLM APIs creates a two-way pipe where usage data signals the platform provider to absorb the product as a feature. The post traces how categories lik…
Anthropic's Claude Opus 4.5 and Zhipu AI's GLM-5.2 are frontier reasoning models competing on cost, context length, and coding benchmarks. GLM-5.2 offers a 1M-token context window and leads by 20.3 po…