[AINews] SpaceXAI Grok 4.6 and Grok @Bot XAI released Grok 4.6, a 1.5T-parameter model that scores 61 on the Artificial Analysis Intelligence Index, roughly in line with GPT-5.6 Sol Max but behind Claude Opus/Fable, with 88.4% on Terminal-Bench v2.1 and 1753 GDPval-AA v2 Elo, at $2/$6 per 1M input/output tokens. The model powers the new Grok @Bot in Cursor, which has drawn positive reviews, positioning xAI as a top contender in the AI teammate category. AINews SpaceXAI Grok 4.6 and Grok @Bot AI teammate category just had its most significant new entrant yet One of our top recurring themes of the year has been coding agents breaking containment into knowledge work https://www.latent.space/p/ainews-agents-for-everything-else , and it’s clear that the AI teammate/multiplayer/multiagent space is the next big AI battleground. With Claude Tag https://www.latent.space/p/ainews-claude-tag-multiplayer-proactive launching to mixed reviews and Block’s Buzz https://block.xyz/inside/introducing-buzz-where-humans-and-agents-work-together requiring a more technical user, the space was still open for a new category leader, which the now Cursor→SpaceX https://x.com/Techmeme/status/2085810949563543786 team has adroitly shipped to very https://x.com/GergelyOrosz/status/2087636651329618108?s=20 positive https://x.com/kunchenguid/status/2087567139318477117 reviews https://x.com/SherryYanJiang/status/2087317738436125070?s=20 : This is powered by their newest model, Grok 4.6, released today as arguably the second best knowledge work model in the world as both competitor Cognition https://x.com/cognition/status/2087579582492987881 and Elon acknowledges https://x.com/elonmusk/status/2087606260539777263?s=20 … though it is surely the top by efficiency https://x.com/elonmusk/status/2080723860073091158 : Grok 4.6 is a confirmed 1.5T model https://x.com/amanrsanger/status/2087567861040750810 that “builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work”. The only training disclosure can be reproduced in full emphasis ours : Grok 4.6 underwent a longer supplemental training run than Grok 4.5, with curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and animproved optimizerandtraining recipe. This produced a stronger foundation for the SFT and RL stages that followed.We then used Grok 4.5 to regenerate the SFT trajectories across reasoning efforts, agent harnesses, and domainssuch as STEM, software engineering, and knowledge work, and filtered out problematic traces with model-based checks. The resulting SFT checkpoint shows strong performance and improved behavior.Grok 4.6 is trained on a wide range of agentic RL tasks, including knowledge work, general coding, and domain-specific environments for kernel optimization, web development, computer-aided design, and more. It is currently unclear if the lack of sandbox escape incidents when it came to training Grok 4.6 is a testament to their infra engineers or an indictment of their researchers. that is a joke https://x.com/swyx/status/2085790995569090966 about current events, don’t get mad AI News for 8/11/2026-8/12/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space . You can opt in/out of email frequencies AI Twitter Recap Frontier Model Day: Grok 4.6, Qwen3.8-Max, DeepSeek V4 Pro, and Microsoft’s MAI-Thinking-1 Grok 4.6 reaches the frontier on price/performance : xAI released, described as a major step up from 4.5 at the same price. Independent evaluations from Grok 4.6 https://x.com/SpaceXAI/status/2087562800982077492 Artificial Analysis https://x.com/ArtificialAnlys/status/2087564648325530099 place it at 61 on the Intelligence Index , roughly in line with GPT-5.6 Sol Max , behind Claude Opus/Fable, with strong agentic results including 88.4% on Terminal-Bench v2.1 , 1753 GDPval-AA v2 Elo , and competitive AA-Briefcase performance at far lower cost AA-Briefcase note https://x.com/ArtificialAnlys/status/2087598780086632522 . Early arena data from Code Arena https://x.com/arena/status/2087566422390231534 also slots it near GPT-5.6 Sol and Claude Fable on webdev tasks. Pricing is a central theme: AA highlights $2/$6 per 1M input/output tokens , materially below frontier peers, while practitioners immediately framed it as the new default for coding and bug-finding workloads Pawel Huryn https://x.com/PawelHuryn/status/2087600689337835811 , Cognition availability in Devin https://x.com/cognition/status/2087579582492987881 . xAI says the gains came from a longer supplemental training run, regenerated SFT traces, and agentic RL over coding, web, CAD, and kernel optimization; they also report more self-testing behavior during long tasks @kimmonismus summary https://x.com/kimmonismus/status/2087563670054211704 . Elon also said Grok 4.7 https://x.com/elonmusk/status/2087604711767896527 is already in flight https://x.com/elonmusk/status/2087604711767896527 , with initial training complete and supplemental training on SpaceX internal data planned. Qwen3.8-Max open weights are out : Alibaba’sdropped as an open-weight Qwen3.8-Max https://x.com/ClementDelangue/status/2087562019788697818 2.4T total / 95B active MoE . Community notes emphasize its scale, day-0 serving, and long-context/agent orientation: Yuchen Jin https://x.com/Yuchenj UW/status/2087566479558394360 called it one of the largest open-weight releases to date; vLLM https://x.com/vllm project/status/2087571359413281049 shipped day-0 support plus vendor-specific 4-bit checkpoints for NVIDIA B300 and AMD MI355X ; Together AI https://x.com/togethercompute/status/2087649685129318585 and Baseten https://x.com/baseten/status/2087654112338817278 also announced immediate support. One important caveat from users: the released open-weights variant appears to be text-only , with no vision input in the initial drop skalskip92 https://x.com/skalskip92/status/2087578544801010075 . DeepSeek V4 Pro GA undercuts the market : DeepSeek’simmediately drew attention less for “best benchmark in every column” than for economics. Multiple observers highlighted pricing around V4 Pro GA rollout https://x.com/synthwavedd/status/2087558842271813860 $0.435/M input and $0.87/M output kimmonismus https://x.com/kimmonismus/status/2087577624180637806 , with Cline https://x.com/cline/status/2087602193205694891 calling it roughly 57× cheaper than Fable 5 while reporting meaningful gains over the preview, including a 15.8% Terminal Bench increase . Reaction was mixed on capability: some early users found it solid but not clearly ahead of Kimi/Flash on all tasks Yuchen Jin’s roundup https://x.com/Yuchenj UW/status/2087577925919068639 , scaling01 https://x.com/scaling01/status/2087569635612778655 , teortaxesTex https://x.com/teortaxesTex/status/2087582179039563836 , suggesting DeepSeek’s next gains may depend more on RL environment and agent work than raw scale. Microsoft enters with its own reasoning model : Mustafa Suleyman announced, Microsoft’s first reasoning model “built from scratch,” now available in Foundry. The initial ask from the team is notably practical— MAI-Thinking-1 https://x.com/mustafasuleyman/status/2087570047967408396 Finbarr Timbers https://x.com/finbarrtimbers/status/2087593173501771987 specifically requested feedback on tool use —which suggests Microsoft is positioning it as an applied reasoning model rather than just a benchmark entrant. Solar Pro 4 also moved up a tier : Artificial Analysis https://x.com/ArtificialAnlys/status/2087590023742775472 reported that Upstage’s Solar Pro 4 jumped from 14 to 42 on the Intelligence Index, with especially large gains on agentic and long-context tasks, though still behind the current top frontier and open leaders on both raw score and price. Open-Weight Multimodal and Edge Models: Video, Vision, Voice, and Local Inference LTX-2.5 and the open video stack keep improving : @RisingSayak https://x.com/RisingSayak/status/2087457946770850274 highlighted that Lightricks’ LTX-2.5 landed in Diffusers with several practical features that matter for local workflows: joint video + 48 kHz audio generation , prompt-controlled clip length, a 2-pass quality mode , tile rendering for lower memory usage, and preprocessing that re-compresses input images to better match training. Ostris AI Toolkit https://x.com/ostrisai/status/2087507808984199668 added support the same day. More broadly, several accounts framed the week as an unusually strong run for open multimedia releases, including MiniMax H3 , LTX-2.5 , LFM2.5-VL-3B , and North Micro Vision victormustar https://x.com/victormustar/status/2087551400377037062 , multimodalart https://x.com/multimodalart/status/2087576052457513234 . Small VLMs and local multimodal are getting serious : Cohere launched, an Apache-2.0 open-source small VLM aimed at North Micro Vision https://x.com/cohere/status/2087571573947392419 document understanding , with claims of outperforming Gemma 4 E2B and Ministral 3 3B on a broad visual benchmark mix results thread https://x.com/cohere/status/2087571579517489581 . Liquid AI’s LFM2.5-VL-3B was also repeatedly cited as a strong compact vision model, and users demonstrated hybrid local/remote agent stacks—for example, Hermes Agent using DeepSeek V4 Flash for planning plus LFM2.5-VL-3B for local vision https://x.com/noctus91/status/2087559912687862240 . Speech and sign-language releases were unusually substantive : Google DeepMind announced, a sign-language-to-text system powering ASL input on Android/Pixel 11. The follow-up notes are technically interesting: SL2T https://x.com/GoogleDeepMind/status/2087541213284946191 body pose tracking happens on-device , translation runs server-side, and the system is optimized for real-world constraints like one-handed signing detail https://x.com/GoogleDeepMind/status/2087541217965809850 . Separately, Deepgram launched, a low-latency conversational TTS model claiming Flux TTS https://x.com/deepgramscott/status/2087533416849838386 ~80 ms response time and mid-call adaptation for voice agents. Inference, Compression, and Systems: vLLM, Quantization, CUDA Scheduling, and Ranking Infra vLLM added important infra for giant models and long prompts : vLLM https://x.com/vllm project/status/2087543021844017182 now supports Azure Blob paths for both model loading and KV connectors. The Microsoft/NVIDIA recipe matters operationally: faster weight loading via Dynamo ModelExpress up to 7.3× faster on H100/A100 and blob-backed KV caching via LMCache + NIXL , trading recomputation for fetches on long-prompt workloads follow-up https://x.com/vllm project/status/2087543024213737527 m . Compression work is extending the useful life of very large models : LLM Compressor v0.13.0 https://x.com/RedHat AI/status/2087519343349305528 added REAP expert pruning for MoE models—dropping whole experts based on calibration saliency before quantization—as well as arbitrary 3/5/6/7-bit quantization . On the more extreme end, Unsloth https://x.com/UnslothAI/status/2087569665652580797 claimed to shrink Qwen3.8-2.4T-A95B from 4.9 TB to 397 GB via dynamic 1-bit quantization, making local execution conceivable on 410 GB+ RAM/VRAM systems. They also showed a 2-bit Nemotron 3.5 Lightning setup https://x.com/UnslothAI/status/2087598047589196052 sustaining long tool-use sessions in 22 GB VRAM . GPU kernel authoring is getting safer and more declarative : maharshii https://x.com/maharshii/status/2087553144184258961 highlighted CuTeDSL 4.7.0 Task Scheduling kernels , which let developers explicitly declare warp roles, resources, dependencies, and schedules, enabling static checks for deadlocks, races, and barrier initialization before lowering to GPU code. The same author also posted a concise explainer on the prerequisites behind TMA async copy —acquire/release semantics, mbarriers, and CuTe arithmetic tuples—for people trying to reason about modern NVIDIA memory movement primitives thread https://x.com/maharshii/status/2087495927313629516 . Classic recommender/ranking stacks are still quietly delivering wins : François Chollet pointed to Expedia’s migration to a modern Keras 3 setup, reporting 30% faster training and 70% lower inference latency for ranking models tweet https://x.com/fchollet/status/2087519531547701335 . His follow-up stresses a more strategic point: Keras’s backend-agnostic APIs reduce lock-in if teams later need PyTorch or JAX kernels note https://x.com/fchollet/status/2087557096736702699 . Agents, Harnesses, and Developer Tooling: Reliability, Memory, Plugins, and Security The stack above the model is becoming the main product surface : Several tweets converged on the same theme: many practical gains are coming from harness engineering , memory, approvals, evals, and tools more than bespoke model training. Scott Stevenson restated the argument that RAG and harness engineering beat training most of the time because they personalize per customer, avoid privacy risks, improve in real time, and inherit base-model progress thread https://x.com/scottastevenson/status/2087511232169308371 , follow-up https://x.com/scottastevenson/status/2087555212470853655 . Random Walker added a useful product distinction between delegation agents and collaboration agents , with very different optimization targets around verifiability, latency, and human control tweet https://x.com/random walker/status/2087598781436944399 . Tooling releases reflected that shift : GitHub’s @code https://x.com/code/status/2087640853783232562 introduced Agent Plugins 1.0 , packaging skills, MCP servers, and AI extensions together, and separately shipped UX improvements like sticky scroll and better session handling release thread https://x.com/code/status/2087591365357998136 . OpenAI/Codex-side momentum showed up too, including Codex for Linux https://x.com/reach vb/status/2087639484275863830 . LangChain rebuilt LangSmith dashboards https://x.com/LangChain/status/2087557830408626639 for more useful trace analysis and reporting. Memory and portable agent state are becoming baseline expectations : Hermes Agent got multiple ecosystem updates, from Raspberry Pi deployment https://x.com/witcheer/status/2087509716746326124 to easy profile export/import https://x.com/tonbistudio/status/2087642578128921068 and new skills like generating reusable APIs from observed web traffic Teknium https://x.com/Teknium/status/2087686461822996905 . Managed Deep Agents examples from LangChain focused explicitly on durable memory and recurring workflows such as social-media agents hwchase17 https://x.com/hwchase17/status/2087607611097264579 . Security and governance for agents is becoming concrete : W&B showed a side-by-side agent email example where one agent leaked SSN/card info while another blocked prompt injection and redacted secrets before the model saw them thread start https://x.com/wandb/status/2087524765548577209 . The Turing Post raised a more architectural issue around delegated identity : if an agent uses your SaaS credentials directly, revocation and auditing become muddy tweet https://x.com/TheTuringPost/status/2087555136864289032 . Benchmarks, Research Directions, and AI-for-Science AI-assisted math and science claims are getting harder to ignore : The most engaged technical tweet was Steven Strogatz sharing a story that a neurosurgery resident reportedly used ChatGPT 5.6 to solve a significant open problem in numerical linear algebra tweet https://x.com/stevenstrogatz/status/2087474852814880960 . Relatedly, multiple accounts noted another EpochAI open problem apparently falling scaling01 https://x.com/scaling01/status/2087534845937189235 . New benchmarks target less gamed capabilities : Princeton/MIT collaborators released, a text-based benchmark for DiG-bench https://x.com/jcrwhittington/status/2087535497480388729 discovery rather than standard QA or code tasks; tri Dao specifically praised it for having some of ARC’s flavor without confounding vision issues tweet https://x.com/tri dao/status/2087677140410290302 . Redwood + Anthropic introduced the, targeting AI-risk-relevant argumentation and conceptual reasoning where feedback is sparse and hard to automate. Vals announced Conceptual Reasoning Index https://x.com/emwcooper/status/2087584904905114064 , focused on binary reverse engineering rather than source-level cyber tasks. SRE-Bench https://x.com/ValsAI/status/2087682813743317396 Post-training efficiency and long-context research stood out : Lewis Tunstall summarized, where RL is done on a smaller model and the resulting policy shift is transferred to a larger model using a dense implicit reward, roughly halving pipeline cost in the cited setup. Separately, Direct On-Policy Distillation https://x.com/ lewtun/status/2087530369306288300 dair.ai’s summary https://x.com/dair ai/status/2087600513441546589 of new OLMo/Llama/Qwen long-context work argues that four architecture choices —normalization, GQA, pretraining context length, and sliding-window attention—can together cost up to 47% of long-context performance , even when short-context validation looks fine. Clinical and domain-specific RL is maturing : A thread summarizing Google’s ResidencyRL work reports that training Gemini 3.5 Flash over 49,870 simulated telehealth encounters increased diagnostic accuracy under adversarial conditions from 81% to 88% and reduced missed red flags by 31% kimmonismus https://x.com/kimmonismus/status/2087532555277115604 . Snowflake also shared a good counterexample to “bigger always wins”: a new 4B SQL autocomplete model https://x.com/StasBekman/status/2087690011433164807 beat their previous 30B-A3B MoE , improving user acceptance while cutting median latency 71% . Top tweets by engagement Grok 4.6 release : @SpaceXAI https://x.com/SpaceXAI/status/2087562800982077492 announced the model; @elonmusk https://x.com/elonmusk/status/2087565020158992709 amplified it; Artificial Analysis https://x.com/ArtificialAnlys/status/2087564648325530099 provided the most useful independent breakdown. Qwen3.8-Max open weights : @ClementDelangue https://x.com/ClementDelangue/status/2087562019788697818 , @Yuchenj UW https://x.com/Yuchenj UW/status/2087566479558394360 , and @UnslothAI https://x.com/UnslothAI/status/2087569665652580797 captured the release, deployment, and aggressive quantization angle. DeepSeek V4 Pro GA : @synthwavedd https://x.com/synthwavedd/status/2087558842271813860 on rollout; @cline https://x.com/cline/status/2087602193205694891 and @kimmonismus https://x.com/kimmonismus/status/2087577624180637806 on the unusually strong price/performance profile. AI-for-math headline : @stevenstrogatz https://x.com/stevenstrogatz/status/2087474852814880960 shared the numerical linear algebra story involving ChatGPT 5.6. Accessibility milestone : @GoogleDeepMind https://x.com/GoogleDeepMind/status/2087541213284946191 announced SL2T for ASL-to-English input on Android. AI Reddit Recap /r/LocalLlama + /r/localLLM Recap 1. Claude Text Watermarking Rollout Activity: 2077 : Claude now embeds invisible watermarks in all text outputs + signed metadata on files https://www.reddit.com/r/singularity/comments/1vkzjln/claude now embeds invisible watermarks in all/ Anthropic says Claude marks some AI-generated/edited content via metadata/provenance signals, not a visible text watermark; the mechanism and persistence depend on file type/workflow and may be lost after editing, export, or platform handling Commenters are skeptical of usefulness for text because paraphrasing through another model or local LLM could likely remove detectable signals, and some view any Claude-linkable marking as a privacy/control reason to prefer open-source models. support article https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content . For plain text, commenters question whether this implies statistical linguistic watermarking versus attached metadata; based on Anthropic’s description, the robust claim is metadata/provenance marking, not an undeletable watermark embedded in arbitrary copied text. Anthropic/Claude rollout details: commenters cite the submission statement that Claude models launched on or after August 2, 2026 will embed an imperceptible model-level text watermark intended to survive copy-paste and some editing without changing readability or semantics. Supported file outputs such as .png , .jpg , and .svg will also carry digitally signed C2PA provenance metadata , with third-party detection tooling still forthcoming and older models expected to be updated during a transition period.A technical concern raised is robustness: for text, users argue the watermark may be removable by paraphrasing through another model, especially a local/open-source one, because rewording can destroy token-level statistical patterns. Another commenter notes this is not unique to Anthropic and points to OpenAI’s provenance/watermarking work : Understanding the source of what we see and hear online https://openai.com/index/understanding-the-source-of-what-we-see-and-hear-online/ . Activity: 878 : How would an “invisible watermark” in AI-generated text actually work? https://www.reddit.com/r/ClaudeAI/comments/1vl9gq5/how would an invisible watermark in aigenerated/ The thread asks how an invisible text watermark could be embedded in Claude-style LLM output without hidden Unicode; the technical answer is a keyed generation-time scheme that slightly biases token sampling toward pseudo-randomly selected “favored” tokens based on prior context and a secret key, then detects overrepresentation via a statistical score such as a z-score . Commenters note this is robust to copy/paste and minor edits, but degrades under substantial paraphrasing, sentence restructuring, or regeneration by another LLM; Google’s SynthID-Text approach, described in The main skepticism is epistemic: Nature https://www.nature.com/articles/s41586-024-08025-4 , uses a related tournament-sampling watermarking method. “how would anyone know if it was watermarked?” —i.e., detection depends on access to the secret rule/key or a trusted detector, and robustness claims are limited once the text is heavily rewritten.A commenter describes LLM text watermarking as a keyed sampling bias : during next-token generation, the model slightly boosts a secret, context-dependent subset of tokens, producing a hidden statistical pattern while preserving fluency. Detection then recomputes the same secret rule over the text and checks whether favored tokens occur above chance, often via a z-score-like statistic ; copy/paste and light edits may preserve the signal, while heavy paraphrasing can destroy it.One linked technical reference is the Nature paper “Scalable watermarking for identifying large language model outputs” https://www.nature.com/articles/s41586-024-08025-4 , which is relevant to production-grade schemes such as Gemini-style tournament sampling . The discussion also notes that implementations vary by provider, with claims that watermarking can be added at the sampling/model-output layer rather than requiring visible text markers.A key unresolved technical concern raised is false positives : if detection is purely statistical, naturally written text could coincidentally overuse the “green-list” or favored tokens. This implies practical detectors need calibrated thresholds, long-enough samples, and measured false-positive/false-negative tradeoffs rather than treating watermark detection as a deterministic yes/no signal. 2. Frontier Model Security and Governance Flashpoints Keep reading with a 7-day free trial Subscribe to Latent.Space to keep reading this post and get 7 days of free access to the full post archives.