{"slug": "tai-218-enterprise-ai-use-is-becoming-more-uneven", "title": "TAI #218: Enterprise AI Use Is Becoming More Uneven", "summary": "OpenAI reports that its top 10% of enterprise customers by output tokens per active user generated 8.3 times as many tokens per active user as the middle decile in June, up from 2.6 in January, while Ramp data shows the median firm in its top 1% AI-spend-per-employee tier paid $7,400.50 per employee in July, an annualized run rate of $88,806, 619 times the median AI-paying firm. The concentration of AI activity and spending is widening rapidly, with API use, GPU cloud, and model serving comprising 72.88% of measured July AI spending, and Grok 4.6, released August 12 by SpaceXAI, scores 61 on Artificial Analysis, matching GPT-5.6 Sol Max at a lower cost of $0.84 per task.", "body_md": "Two datasets published this week show how quickly AI activity and spending are concentrating among a small group of companies. [OpenAI](https://openai.com/index/how-enterprises-put-ai-to-work/) says its top 10% of enterprise customers by output tokens per active user generated 8.3 times as many output tokens per active user as customers in the middle decile in June. [Ramp](https://ramp.com/data/ai-index) shows the median firm in its top 1% AI-spend-per-employee tier paid $7,400.50 per employee in July. That is an $88,806 annualised run rate, 11.4 times the top 10% tier median and 619 times the median AI-paying firm.\n\nThese figures measure activity and spending, not productivity or return. A long-running agent can consume millions of tokens while delivering valuable, checked work. A badly designed agent can spend the same amount repeating a failed plan. Ramp also counts application programming interface (API) use, graphics processing unit (GPU) cloud, and model-serving costs that may power customer products rather than employee tools. The concentration is clear. Which firms are earning a return is not.\n\nThe gap has widened fast. OpenAI’s frontier-to-typical ratio rose from 2.6 in January to 8.3 in June. Token use in the top group grew 319%, against 32% in the middle group. Firms are ranked again each month and selected on the same token measure, so the change is more useful than the absolute gap. Codex produced 64% of combined ChatGPT and Codex enterprise output tokens in June. In a separate small study, the share of sampled Codex users attempting tasks estimated to take a skilled person at least eight hours rose from 2.1% in December to 25.6% in May. Long agent tasks can scale far beyond the amount of chat a person has time to read and answer.\n\nRamp shows where much of the money went. API use, GPU cloud, and model serving or inference made up 72.88% of measured July AI spending. Chat and coding-agent subscriptions made up 10.22%. The largest AI bills increasingly include production computing, which is why dividing the whole bill by employee count can be misleading.\n\nGrok 4.6 shows how quickly model price and performance are shifting. SpaceXAI released it on August 12, only 27 days after Grok 4.5. [Artificial Analysis](https://artificialanalysis.ai/models/grok-4-6) scores Grok 4.6 High at 61, level with GPT-5.6 Sol Max, at a measured $0.84 per task against $1.23 for Sol, $3.14 for Fable 5, and $2.34 for Opus 5. Among the models scoring 61 or higher on the current chart, Grok has the lowest measured task cost. It does not lead every agent benchmark, but this is the first Grok release in some time that clearly competes at the frontier on capability and price. Musk says Grok 4.7 could follow within three to four weeks of August 12 after more training on SpaceX data. SpaceXAI has not published a model card, price, or firm release date, so this remains a founder forecast. The 27-day gap from Grok 4.5 to 4.6 still makes the faster release pace worth watching.\n\nThe choice is also becoming about speed. OpenAI has previewed [GPT-5.6 Sol Ultrafast](https://openai.com/index/previewing-ultrafast/), powered by Cerebras, at up to 750 output tokens per second, or 14 times faster than Standard processing. Codex Spark reached extreme speed with a smaller model designed for low latency. Ultrafast runs the same GPT-5.6 Sol model, with no model downgrade. OpenAI has not published a price and access remains limited. I expect a steep premium, which will be wasteful for most batch work and worth paying for when latency changes the result. This could become an arms race among quantitative funds that put LLM analysis into medium-frequency trading. If that analysis becomes core to a strategy, a fund may pay a huge amount to receive it even a fraction of a second earlier. Grok’s low task cost and Sol Ultrafast’s speed show why firms need to know which part of their model spend changes revenue, risk, or user experience.\n\nMy view is that the OpenAI and Ramp data show two divides, one between firms and one inside them. A small group of power users can assign agents much more ambitious work than the average employee. High use can produce a strong return when users provide valuable tasks, current context, useful tools, good tests, and clear review. It can also waste tokens through stale context, weak task breakdown, repeated retries, and missing stop rules. Without enough expertise or taste, firms can produce AI slop at scale.\n\nUsing frontier LLMs and agents well is much harder than most people think. A short prompt class will not teach it. People need to see complex, long-running examples tied to their role, then practise on work they can check. Study your real power users, find the tasks and methods that work, and turn them into role-specific systems that load company context, connect the right tools, and enforce permissions and checks. Many firms will need AI engineers or forward-deployed engineers to build and maintain this layer.\n\nDo not turn the headline figures into usage targets. Pair every usage number with an accepted outcome and include quality, human review, rework, full cycle time, cost, and latency. For some work, waiting ten seconds instead of one has no effect on the result. For trading, live support, incident response, or voice, the delay may determine whether the work has value.\n\nRoute each task to the cheapest model and service tier that clears its quality and speed target. Grok 4.6 deserves private tests for low-cost frontier work. Sol Ultrafast may justify a large premium where fractions of a second change the outcome. Start new workflows with small budgets and clear tests, then give larger budgets to proven users. When the same failure repeats, require a new plan or a human decision rather than more tokens.\n\n*— **Louie Peters — Towards AI Co-founder and CEO*\n\n*This issue is brought to you thanks to **Unblocked**:*\n\n**[Webinar] Can you prove AI is working?**\n\nAI is in your engineering workflow. While the token spend shows it, the throughput doesn’t. The human is very much still in the loop, and that’s a context problem.\n\n[Join live on Aug 19 (FREE)](https://watch.getcontrast.io/register/unblocked-can-you-prove-ai-is-working?utm_source=towardsai&utm_medium=email&utm_campaign=primary) to learn:\n\n1. [Google AI Just Released Gemini 3.7 Flash](https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/)\n\nGoogle released Gemini 3.7 Flash just three weeks after 3.6 Flash, building on the same model with algorithmic improvements informed by developer feedback. Coding improved substantially, with DeepSWE rising from 49.0% to 65.3% and FrontierCode from 34.4% to 43.6%, but the gains extend beyond code: GDP.pdf rose from 22.0% to 34.0% for complex document understanding and AutomationBench from 17.0% to 30.4% for enterprise workflows. Google also says the model adapts better when it hits roadblocks, clarifies intent more often, and is more deliberate with multi-step planning and tool calls. It accepts text, images, audio, and video; supports function calling, search, and computer use; and maintains a 1M-token context window with 64K output tokens. Pricing is $0.75/$3.75 per million input/output tokens through December 31 before doubling in 2027. Gemini Spark for Pro and Ultra users now runs on 3.7 Flash, with improved Google Workspace tool use for tasks such as consolidating files, drafting emails, and updating status documents.\n\nSpaceXAI released Grok 4.6 with a stronger focus on long-running agents and visual, interactive work. Training included a longer supplemental run with curated model-generated reasoning and engineering data, an improved optimizer, regenerated SFT trajectories from Grok 4.5, and RL environments spanning coding, web development, CAD, kernel optimization, and knowledge work. Its Artificial Analysis score rises from 56 for Grok 4.5 to 61, matching GPT-5.6 Sol Max, while SpaceXAI reports gains over 4.5 on every listed evaluation. The comparison is not an across-the-board lead: Grok scores 69.9% on CursorBench versus Fable 5 Max at 70.5%, 65.9% on DeepSWE versus Sol at 73%, and 26% on Terminal-Bench 3.0 versus Sol at 34.6%, but leads both on GDPVal-AA. SpaceXAI also says longer runs show more self-testing and verification, though that is an internal observation. Grok 4.6 supports 500K context, text and image inputs, and a new xhigh reasoning level; API pricing starts at $2/$6 per million tokens for prompts below 200K tokens and doubles beyond that threshold, with a separate Fast variant at twice the standard price.\n\nZ.AI released GLM-5.3 using the same base model as GLM-5.2, with all improvements coming from post-training built around longer, more realistic units of engineering work. Some training environments represent several days of work and provide the agent with access to codebases, compute clusters, storage, documentation, and experimental results; Z.AI uses research agents to generate these environments and judge agents to verify that the tasks are actually solvable. Terminal-Bench 3.0 rose from 4.6 to 28.3, DeepSWE from 46.2 to 66.9, and Agents’ Last Exam from 23.8 to 28.5, while Z.AI’s private Code Bench shows 5.3 completing more work with fewer output tokens than 5.2. Cybersecurity improved even faster: CyberGym reached 84.5%, while ExploitBench more than doubled from 24.4% to 54.4%, although Mythos 5 and GPT-5.6 Sol remain well ahead on deeper exploit tasks. Z.AI also says expert review of real code found 2,436 vulnerabilities across 269 projects, including 1,097 medium-to-high-severity issues, which it is tracking through a new public disclosure ledger. GLM-5.3 is text-only with 1M context, 128K maximum output, and always-on reasoning at low, high, or max effort. It is available through the Coding Plan; the general Model API is still coming, and Z.AI is holding the weights for two weeks while it completes additional safety evaluation and hardening.\n\n4. [OpenAI Expands Daybreak with GPT-5.6-Cyber](https://openai.com/index/expanding-daybreak-as-the-cyber-defense-window-narrows/)\n\nOpenAI expanded its Daybreak cybersecurity program with Blue and Red access tiers and introduced GPT-5.6-Cyber, a Sol-based model specialized for advanced authorized security research. Blue removes system-level cyber request screening for approved defenders, while Red provides the specialized Cyber model for more sensitive work. On OpenAI’s internal refusal evaluation, GPT-5.6-Cyber responded to 95% of advanced cyber requests, versus 1.5% for safeguarded Sol and 57.3% for GPT-5.5-Cyber; this measures willingness to respond, not task success. OpenAI also used the model to find two previously unknown V8 vulnerabilities that could be chained to escape its heap sandbox, with one fixed as CVE-2026–15903. Daybreak is expanding through major security and consulting partners, while OpenAI has separately tightened controls around Astra after saying it cannot yet rule out Critical cyber capability.\n\n5. [NVIDIA Launches Nemotron 3.5 Lightning](https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/)\n\nNVIDIA released Nemotron 3.5 Lightning, a 30B-parameter MoE model with 3B active parameters, a hybrid Mamba-2, MoE, and Attention architecture, and a 1M-token context window. It is designed for the high-volume execution layer of long-running agents, handling tasks such as tool calls, validation, and subagent work while larger models handle planning. NVIDIA reports up to 4x faster output than similar-sized models and 30% faster PinchBench task completion than Qwen3.6–35B at comparable accuracy. Artificial Analysis scores it at 24, nine points above Nemotron 3 Nano. NVIDIA also released NeMo Switchyard for routing different parts of an agent workflow to different models. BF16 and NVFP4 weights are available under OpenMDW-1.1, with local deployment supported on systems including RTX 5090 and DGX Spark.\n\n6. [MiniMax Releases Music 3 as Open Weights](https://huggingface.co/MiniMaxAI/MiniMax-Music3)\n\nMiniMax released the weights for Music 3.0, a roughly 11.1B-parameter music-generation system capable of producing complete songs up to five minutes long with vocals and full arrangements. Its architecture combines an 8B Global LLM, a 0.6B Local LLM, a 2.4B flow-matching module, and a 123M Flow-VAE decoder. Users provide lyrics with optional section tags plus a detailed music description, and the model outputs 32 kHz stereo WAV audio. Music-3.0 first launched as a hosted API on July 16; downloadable weights are now available on Hugging Face and ModelScope, with support for SGLang, Diffusers, and ComfyUI. Unlike MiniMax’s H3 license, Music 3 has no equivalent geographic exclusions, although products that use the model and generate more than $20M in annual revenue require separate written authorization.\n\nHybrid search is supposed to cover two different retrieval failures: keyword search catches exact strings such as product codes and error messages, while semantic search catches different wording with the same meaning.\n\nBut combining the two does not automatically preserve both advantages.\n\nSuppose a user searches for order #8821. Keyword search may rank the exact page first, while semantic search returns several broader pages about orders. If you merge both lists immediately and keep only the highest-ranked results, those broader pages can occupy most of the final set, pushing out the exact match.\n\nWe ran into this while building the AI tutor in our [Full Stack AI Engineering](https://towardsai.staging.tempurl.host/academy/full-stack-ai-engineering/?utm_source=newsletter&utm_medium=email&utm_id=AItip) course. The fix was to preserve the strongest candidates from each retriever before combining them: keep the top 5 keyword results and the top 5 semantic results, then merge and deduplicate the two sets.\n\nIf exact identifiers matter in your application, you can go further by reserving part of the final context for strong keyword matches.\n\n1. [Inside Agent Skills: A Structured Workflow Framework for AI Coding Agents](https://pub.towardsai.net/inside-agent-skills-a-structured-workflow-framework-for-ai-coding-agents-9ee700f411ff?sk=b4c1039f26ec8e87a1e1999bcd154889)\n\nThis article breaks down Agent Skills, an open-source framework that gives AI coding agents a structured way to work through the full software development lifecycle. You will learn what a skill actually is, how the six phases fit together, and how the skill file format keeps an agent on track from a rough idea to shipped code. It also walks through the architecture by running a small application through the entire pipeline.\n\n2. [The Harness Is the Product: An End-to-End Guide to Harnessing in Agentic AI](https://pub.towardsai.net/the-harness-is-the-product-an-end-to-end-guide-to-harnessing-in-agentic-ai-fcc0a9931526?sk=5491fd1934989f872367cfbf60346bf1)\n\nThe piece breaks down seven harness components: prompts, tools, context management, memory, guardrails, verification, and observability, and covers multi-agent orchestration, checkpointing, and evals. It also walks through case studies on Claude Code, Deep Research, Manus, and Cursor that show identical models producing different products through scaffolding choices.\n\n3. [Adding Cost Metering and LLM Spend Visibility to a Multi-Agent System](https://pub.towardsai.net/adding-cost-metering-and-llm-spend-visibility-to-a-multi-agent-system-38e2d8591fb1?sharedUserId=tai-tech)\n\nMulti-agent LLM systems create a billing blind spot: provider dashboards show total spend but reveal nothing about which agent, workflow, or retry caused it. This piece details a metering layer that captures token counts at the call site, prices them against a versioned rate card collection, and writes attributed usage documents tagged with traceId and agent name. Aggregation pipelines then surface cost by agent, model, or outcome, feeding Atlas Charts dashboards and materialized views built for scale.\n\n4. [Procedural Memory in AI Agents: Why Knowing the Answer Is Not Enough](https://pub.towardsai.net/procedural-memory-in-ai-agents-why-knowing-the-answer-is-not-enough-fc072c8f8acf?sk=190570f41df21cd572e68306f167a96d)\n\nProcedural memory gives AI agents a reusable method for familiar tasks, distinct from knowing facts or recalling past events. The author uses the example of solving a Rubik’s Cube to show how a learned method turns scattered moves into steady progress, applying the idea to a leave-request assistant and a coding agent guided by an AGENTS.md file. It also uses LangGraph and LangMem to turn user corrections into lasting instructions.\n\n5. [Explaining Markov Chain Monte Carlo using Wildfire Forensics](https://pub.towardsai.net/explaining-markov-chain-monte-carlo-using-wildfire-forensics-a334fecaefb3?sharedUserId=tai-tech)\n\nThis article explains the Markov Chain Monte Carlo algorithm by investigating the ignition source of a wildfire. It walks through the Metropolis algorithm’s mechanics: proposing a neighboring cell, computing the acceptance ratio relative to the current posterior, and accepting or rejecting via a random draw. It shows visit frequencies converging on the true posterior using Python simulations and covers Metropolis’s limitations and Hastings’s correction factor toward Hamiltonian Monte Carlo.\n\n1. [Anthropic Cybersecurity Skills](https://github.com/mukul975/Anthropic-Cybersecurity-Skills) contains 817 structured cybersecurity skills spanning 29 security domains, each following the agentskills.io open standard.\n\n2. [Llmfit](https://github.com/AlexsJones/llmfit) is a Rust CLI that detects your hardware and scores hundreds of LLMs on fit, speed, quality, and context, telling you which ones will actually run on your machine.\n\n3. [oMLX](https://github.com/jundot/omlx) is an LLM inference server for Apple Silicon with continuous batching, SSD-backed KV cache offloading, and Metal-accelerated decoding.\n\n4. [Needle 2](https://github.com/cactus-compute/needle) is a 45M-parameter tool-calling model compressed into a single 14MB binary that runs a full session in 28MB of RAM.\n\n1. [BDH-CQ: In-Context Learning with Recurrent Latent Reasoning](https://arxiv.org/abs/2608.09888)\n\nBDH-CQ unifies in-context learning with latent reasoning: demonstrations update a recurrent memory, and the model solves queries through iterative computation in a continuous latent space without verbalizing intermediate steps. A 150M-parameter configuration reaches 29.5% pass@2 on ARC-AGI-1 at $0.0007 per task, breaking the previously reported cost-accuracy Pareto frontier for the benchmark.\n\n2. [A Controlled Study of Attention-Only Transformers](https://arxiv.org/abs/2607.18363)\n\nFeed-forward networks hold two-thirds of a transformer’s non-embedding parameters, but are they necessary? This paper pretrains attention-only transformers against standard transformers that are matched in parameters, FLOPs, and depth. Deleting FFN layers in place is costly, but reallocating the freed budget into attention depth closes the gap to 0.006 nats (0.27% of loss), reproducible across seeds and shrinking with scale. The residual deficit concentrates entirely on low-context factual recall, not reasoning.\n\n3. [AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses](https://arxiv.org/abs/2608.12307)\n\nThis paper asks whether a stronger model can improve a weaker model at test time without touching its weights. A builder model constructs inference-time harnesses (code scaffolding for routing, parsing, and answer enforcement) that the target model runs inside. The gains come from offloading unstable reasoning into deterministic code rather than encouraging the target to think harder. Weaker models receive the largest improvements.\n\n4. [DarwinX: Evolving Agent Harnesses Through Natural Selection](https://arxiv.org/abs/2608.07545)\n\nSingle-lineage harness self-improvement is path-dependent: local wins often regress other tasks. DarwinX evolves a population of harnesses with the model frozen, allowing only variants that extend coverage without regressing, while maintaining alternative lineages for recombination. One evolution loop adds roughly 17 points on average across four benchmarks, with a Terminal-Bench harness transferring unchanged to SWE-bench Verified. What evolves is general agent competence, not benchmark-specific patches.\n\n1. [OpenAI previewed Ultrafast](https://openai.com/index/previewing-ultrafast/), a new API service tier that runs GPT-5.6 Sol at up to 750 output tokens per second, or up to 14x faster than Standard processing. Cerebras, which powers the tier, reports a 5.6x end-to-end speedup on GDP-Val without any reduction in quality. In another Cerebras-run test, Sol Ultrafast completed all 2,500 Humanity’s Last Exam questions in 11 hours 11 minutes, versus 78 hours 27 minutes for Fable 5 at comparable accuracy. However, the models were run through different agent harnesses.\n\n2. [Liquid AI releases LFM2.5-VL-3B](https://www.liquid.ai/blog/lfm2-5-vl-3b), a 3.1B-parameter open-weight vision-language model for on-device deployment. It is a non-reasoning model that answers directly, keeping latency low. LFM2.5-VL-3B extends the vision-language capabilities with improvements such as screen/UI understanding, function calling, grounding, and multi-image input. On the text-only ToolSandbox benchmark, its score rose from 26.4 to 59.5. Liquid reports 228 tokens/s on an M5 Max in roughly 3GB of memory. Native, GGUF, ONNX, and MLX checkpoints are available on Hugging Face under Liquid’s LFM license.\n\n3. [Dyna Robotics introduces Dyna-2](https://www.dyna.co/dyna-2), a World-Action Model pretrained on more than one million hours of egocentric human video without robot data, roughly 170 years of continuous experience. Rather than learning only actions, the model jointly learns to predict future video and actions, which Dyna says improves transfer to robot embodiments. In one customer deployment, Dyna-2 achieved an 87% pass rate versus 46% for Dyna-1, while a separate benchmark found that the World-Action architecture achieved 1.55x the success rate of a VLA baseline. Dyna also reports that roughly 10 minutes of robot demonstrations were enough to teach two five-fingered hands to open a bottle cap. The company claims its experiments demonstrate the first human-to-robot transfer scaling law, though the results are not independently verified.\n\n4. [DeepSeek AI releases DeepSeek Harness in developer preview](https://x.com/deepseek_ai/status/2087887408440164663), an MIT-licensed agent framework built around one idea: nearly every capability is a plugin. Powered by Cordis, models, tools, skills, sessions, sandboxes, storage, agent loops, and UI components can be composed or replaced without rebuilding the core harness. It supports DeepSeek, Anthropic, OpenAI, and major cloud providers, as well as custom compatible endpoints, while optional integrations can delegate tasks to Codex or Claude Code as subagents.\n\n**Applied AI Engineer, Digital Natives @OpenAI (São Paulo, Brazil)**\n\n**Solutions Architect, Applied AI @Anthropic (Sydney, Australia)**\n\n**Software Engineer @Microsoft Corporation (Redmond, WA, USA)**\n\n**Manager, AI Engineering @HubSpot (Remote/USA)**\n\n**Senior AI ML Engineer @UnitedHealth Group (Remote/USA)**\n\n**AI Content Writer @Wix (Tel Aviv, Israel)**\n\n**Research Fellowship (Applied AI/ML) @ixigo (New Delhi, India)**\n\n*Interested in sharing a job opportunity here? Contact **sponsors@towardsai.net**.*\n\n*Think a friend would enjoy this too? **Share the newsletter and let them join the conversation.*\n\n[TAI #218: Enterprise AI Use Is Becoming More Uneven](https://pub.towardsai.net/tai-218-enterprise-ai-use-is-becoming-more-uneven-ee84965cb7bd) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/tai-218-enterprise-ai-use-is-becoming-more-uneven", "canonical_source": "https://pub.towardsai.net/tai-218-enterprise-ai-use-is-becoming-more-uneven-ee84965cb7bd?source=rss----98111c9905da---4", "published_at": "2026-08-19 11:38:57+00:00", "updated_at": "2026-08-19 12:11:13.970461+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-products", "ai-infrastructure", "ai-research"], "entities": ["OpenAI", "Ramp", "SpaceXAI", "Artificial Analysis", "Grok 4.6", "GPT-5.6 Sol Max", "Cerebras", "Codex"], "alternates": {"html": "https://wpnews.pro/news/tai-218-enterprise-ai-use-is-becoming-more-uneven", "markdown": "https://wpnews.pro/news/tai-218-enterprise-ai-use-is-becoming-more-uneven.md", "text": "https://wpnews.pro/news/tai-218-enterprise-ai-use-is-becoming-more-uneven.txt", "jsonld": "https://wpnews.pro/news/tai-218-enterprise-ai-use-is-becoming-more-uneven.jsonld"}}