[AINews] Muse Glimmer and Spark: Open Weights return Personal Superintelligence promise Meta released its first open-weights frontier small language model, Muse Glimmer, with Muse Spark to follow, alongside an essay by CEO Mark Zuckerberg outlining the company's agenda for personal superintelligence. Zuckerberg predicted that everyone will have an exceptionally capable personal agent, tools for creation, entrepreneurial tools, personalized tutors, scientific advances, and free or affordable access, while also proposing that frontier AI labs share intermediate training checkpoints with governments and commit resources to harden critical infrastructure. Last week was the 1 year anniversary of Zuck’s original Personal Superintelligence essay https://www.meta.com/superintelligence/ , and MSL seems to be feeling a second wind this year, as they slowly ramped up with the Dreamer acquisition https://www.latent.space/p/ainews-dreamer-joins-meta-superintelligence?utm source=publication-search and then Muse Spark https://www.latent.space/p/ainews-meta-superintelligence-labs?utm source=publication-search and recently Muse Code https://www.latent.space/p/ainews-jeff-sanjay-oriol-and-quoc?utm source=publication-search . For a while it seemed like MSL was being rather timid with the launches… but today that all changed. Zuck returned https://x.com/finkd/status/2086754845218726027 with a hit sequel essay https://www.meta.com/thefutureisforeveryone/ and released MSL’s first real open weights frontier-ish small LLM, with Spark to also be released soon. The essay maps out what is likely to be the lasting agenda for MSL: Meta is the company primarily focused on building personal superintelligence for everyone. Most other labs are focused on building AI for companies, governments, or other institutions, so if those labs lead, thenthe balance of power will favor larger institutions over individuals. Meta's mission since our founding has focused on putting power in people's hands. If our beliefs and principles lead, then the balance of power will favor individuals and a better future for everyone. His core predictions: Everyone will have an exceptionally capable personal agent that understands you, your goals, and everything you care about. likely why he was interested in OpenClaw https://www.businessinsider.com/openclaw-creator-peter-steinberger-gets-feedback-from-mark-zuckerberg Everyone will have incredible tools for creation to express your ideas. Everyone will have powerful tools to create new businesses and the economy will become more entrepreneurial .Everyone will have a personalized tutor and coach with a PhD in every subject and unlimited patience to help you learn anything you want.Everyone will benefit from scientific advances and be able to contribute to scientific progress.Everyone will have free or affordable access to these tools. And he named some core risks: Job Growth and The Economy: “ Company sizes may shrink -- just as they did in the transition from industrial giants to tech companies. But this doesn’t mean fewer jobs overall. It implies a larger number of companies with fewer people each. ” Building AI Infrastructure with Communities: “in Richland Parish, Louisiana, where Meta is building a large data center, teachers received a $50,000 bonus this year because of the increased tax revenue from our investment…We help keep electricity prices low by building our own energy-generating infrastructure wherever we invest….In areas with high water stress, our goal is to restore 200% of the water we use.” Securing Against AI Misuse in Cybersecurity, Bioterrorism, and More : “I propose that companies developing frontier AI should commit significant technical resources towards helping the government harden critical infrastructure. I also propose that frontier AI labs should share intermediate training checkpoints of new models for government use and review rather than waiting until training has completed.” Protecting Freedom and Preventing Government Tyranny: “To maintain freedom, we must ensure that superintelligence primarily empowers individuals. The ideal in liberal democracy is that people naturally hold all rights and only agree to restrict some freedoms to protect the common good. Similarly, individuals should have access to personal superintelligence and should only be subject to restrictions when truly required.” Ensuring American Leadership: “On infrastructure, America and its allies currently hold an advantage in silicon design but a disadvantage in how quickly we can build energy capacity and physical infrastructure. Countries like China are bringing online 1GW+ of nuclear capacity every other week, so we will need to accelerate building both energy and data centers to remain competitive. Export controls on silicon have been successful for slowing the progress of foreign labs during this critical period, so it is the right strategic move to continue those. Any policy that slows American model releases -- even by a month -- could add significant risk to American leadership while letting foreign models race ahead. At the same time, when new capabilities emerge, it is important that the US government has advanced knowledge and resources to harden critical systems, and potentially some period of advantage in using advanced systems.” Alignment With People and Addressing Existential Risk : “A healthy balance of power is to ensure that there is no singular centralized superintelligence, but instead as many people and businesses as possible with different superintelligent agents aligned to their goals that check and compete with each other in the ways our natural economy behaves. This balance would be further enhanced if there were multiple frontier labs whose models have different values that could check each other as well.” Maintaining Control of Superintelligence : “There is a dilemma that once AI systems can autonomously improve themselves, any lab that doesn’t let their AI system direct a substantial amount of compute capacity towards recursive self-improvement will inherently fall behind. For example, if a self-improving AI system focused on optimizing its compute efficiency, it could theoretically invent ways to squeeze 100x or more intelligence out of each gigawatt. That means that a self-improving AI system running on a fraction of the world’s compute could conceivably command more effective compute and intelligence, and therefore a greater balance of power than everyone else combined and become the singular superintelligence we fear.” AI News for 8/8/2026-8/10/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space . You can opt in/out of email frequencies AI Twitter Recap Meta’s Return to Open Weights with Muse Glimmer and Spark 1.2 Meta re-enters the open-weight frontier : The day’s dominant story was Meta’s release of Muse Glimmer , a 30B dense , multimodal, agent-focused model under Apache 2.0 , plus the promise to release Muse Spark 1.2 weights “soon.” The announcement came from Mark Zuckerberg https://x.com/finkd/status/2086755195535413696 and Alexandr Wang https://x.com/alexandr wang/status/2086756152034066792 , with Meta framing this as a renewed commitment to broadly available “personal superintelligence” in Zuckerberg’s essay https://x.com/finkd/status/2086754845218726027 . Meta’s product thread positions Glimmer as optimized for always-on local agents , able to run on consumer hardware, with official details https://x.com/AIatMeta/status/2086757844544811485 and download links https://x.com/AIatMeta/status/2086757850790109574 . What’s technically notable about Glimmer : Meta says Glimmer is designed for long-horizon agent loops, tool use, and local deployment. In the serving stack, Meta explicitly mentions quantization to bring the LM under 20GB and a lightweight DFlash drafter for faster generation on-device, yielding “fluid” local interaction @AIatMeta https://x.com/AIatMeta/status/2086757846847263014 . Community summaries add more architectural color: @eliebakouch https://x.com/eliebakouch/status/2086769240271405477 notes similarities to Gemma 4-style hybrid attention plus scale-free QK norm , larger vision depth, and longer SWA; @nrehiew https://x.com/nrehiew /status/2086779938884182073 highlights that Glimmer was logit-distilled from Muse Spark and trained from the outset on agentic traces , i.e. not a conventional “base then post-train” release. Benchmarks and deployment ecosystem landed immediately : Third-party analysis from Artificial Analysis https://x.com/ArtificialAnlys/status/2086916150278111551 places Muse Glimmer at 35 on its Intelligence Index, just behind Qwen3.6-27B 38 and around Kimi K2.5 36 , while scoring well for openness 44 Openness Index . Their read is that Glimmer is strong for its size and particularly notable for local self-hosting: ~60GB BF16 , ~18GB 4-bit , 128K context , and memory-efficient hybrid attention suitable for single-node deployment details https://x.com/ArtificialAnlys/status/2086916150278111551 . Weaknesses: relatively poor hallucination / knowledge calibration and trailing some peers on agentic knowledge work, though it does well on Tau3-Banking tool use follow-up https://x.com/ArtificialAnlys/status/2086916156796055922 . Anthropic and OpenAI Push on Frontier Capability: Math and Cybersecurity Anthropic’s Claude improves a Riemann-hypothesis-related bound : Anthropic reported that an unreleased research Claude variant, when tasked with the Riemann Hypothesis , did not solve the conjecture but did improve a longstanding lower bound: the fraction of zeta zeros on the critical line increased from 41.6% to 67.2% in its generated result announcement https://x.com/AnthropicAI/status/2086867246073401655 . The post quickly became the second major story of the day, with Jarred Sumner https://x.com/jarredsumner/status/2086869681785500011 adding that the model used repeated retries and large-scale exploration over 31M output tokens . Engineers viewed this less as “RH solved” and more as a striking example of AI-assisted theorem-search and proof iteration; see reactions from @jdlichtman https://x.com/jdlichtman/status/2086903994094682557 and @kimmonismus https://x.com/kimmonismus/status/2086881395465466004 . OpenAI launches GPT-5.6-Cyber under restricted access : OpenAI announced GPT-5.6-Cyber and an expansion of its Daybreak cybersecurity initiative, explicitly positioning the model for advanced, authorized defensive work @OpenAI https://x.com/OpenAI/status/2086864365379010729 . OpenAI says the model has already been used in real-world vulnerability research, including finding previously unknown bugs in open-source software and even Chrome V8 details https://x.com/OpenAI/status/2086864372500942906 . Access is limited to “approved defenders,” with extra controls and monitoring for higher-risk cyber tasks safeguards https://x.com/OpenAI/status/2086864374837150108 . The move follows broader debate over model cyber misuse and agent-driven exploitation, referenced by @kimmonismus https://x.com/kimmonismus/status/2086735083528921422 and @jachiam0 https://x.com/jachiam0/status/2086705930440159403 . Pricing pressure also showed up : Anthropic separately announced that Claude Sonnet 5’s introductory pricing would become permanent at $2/M input and $10/M output @claudeai https://x.com/claudeai/status/2086891169217122586 , a move widely read as competitive pressure amid a rapidly strengthening open and semi-open field. Agent Harnesses, Tool Use, and Cost/Latency Optimization Harness quality is becoming a first-class differentiator : Several tweets underscored that model quality is increasingly constrained by the agent harness , not just the base model. Composio’s benchmark https://x.com/composio/status/2086814488162972027 ran DeepSeek V4 Flash through four harnesses over 30 agentic tasks , finding Pi Agent both the cheapest and the best-performing in that setup. Shashwat Goel https://x.com/ShashwatGoel7/status/2086840890023137420 similarly called Prime-agent a strong general harness for long-horizon tasks. Tool interface design matters more than many stacks assume : A notable paper summary from @dair ai https://x.com/dair ai/status/2086846794840019178 argues that programmatic tool calling —typed Python stubs executed in-code—matches or beats native JSON tool calling in 11/14 models , with the GPT-5.6 family gaining 10.6% over JSON baselines on BFCL v4. The claim: as models get better at code, treating tools as code objects rather than schema blobs increasingly wins, especially under context rot and parallel fan-out. Token efficiency remains a live systems problem : Teknium https://x.com/Teknium/status/2086702328024125926 highlighted read-tool improvements in Hermes Agent , while later reporting a ~60% token reduction for browser automation by collapsing multiple browser actions into one CLI-driven tool interface here https://x.com/Teknium/status/2086881909209252209 and here https://x.com/Teknium/status/2086882821910782270 . Relatedly, Browser Use https://x.com/browser use/status/2086882292761571758 and Stagehand v4 https://x.com/Stagehanddev/status/2086849338089857082 signal a shift toward thinner, browser-native abstractions for agents. Local-first agent toolchains keep improving : Pi’s SDK https://x.com/pidotdev/status/2086777926016540888 emphasized that a coding agent can stay surprisingly capable with only four primitives— read, bash, edit, write —while Jerry Liu’s LiteParse https://x.com/jerryjliu0/status/2086915480389111830 targets low-latency document parsing inside the agent loop, claiming 4 ms for 200 pages on heuristic extraction before falling back to OCR/VLMs. Inference and Systems: Speculative Decoding, Serving, and GPU Efficiency Speculative decoding is getting more production-realistic : A long technical thread summarized by @ZhihuFrontier https://x.com/ZhihuFrontier/status/2086712577296633887 compared DSpark and DFlash on Qwen3-4B in vLLM. Reported result: DSpark 2.45–2.55× baseline throughput vs DFlash 1.96–2.09× , with DSpark’s advantage attributed to semi-autoregressive structure plus a hardware-aware prefix scheduler that avoids wasteful target verification. This is directionally consistent with Meta’s own use of DFlash in Glimmer for local agent responsiveness. Alternative inference architectures remain hot : SemiAnalysis https://x.com/SemiAnalysis /status/2086697535549440370 highlighted TileRT / InferenceX on NVIDIA GPUs as an attempt to emulate high-interactivity characteristics often associated with vendors like Cerebras, Groq, or SambaNova—specifically for batch size 1 , disaggregated serving, and decode/prefill separation. Provider variance is still huge : Across tweets on Muse Glimmer, DeepSeek V4 Flash, and hosted inference, the recurring engineering theme was that “same model” does not imply same user experience. Artificial Analysis https://x.com/ArtificialAnlys/status/2086958697444696113 teased a discussion on why output speed can vary by 15× across providers. Meanwhile QuixiAI https://x.com/QuixiAI/status/2086913835580211500 reported 175 tok/s single request and 1k tok/s at 64 concurrency for DeepSeek V4 Flash on 4× A100 with SlimServe. Video, Multimodal, and Robotics Models MiniMax H3’s open-weight video momentum continues : MiniMax kept pushing H3 as an open-weight video model with rapid community uptake. The company pointed to new ecosystem work around quantization, offloading, Context-IR , and consumer GPU deployment in a ComfyUI livestream recap https://x.com/MiniMax AI/status/2086685565722984842 , and praised fast community response including LoRA support, MLX, and ComfyUI optimizations in a ThursdAI recap https://x.com/MiniMax AI/status/2086724681219068006 . Notably, antirez released a fast Metal implementation https://x.com/antirez/status/2086764219433660463 , which MiniMax itself celebrated as a direct benefit of open weights @MiniMax AI https://x.com/MiniMax AI/status/2086940119324565748 . Seedance, Omni, and creator tooling keep advancing : Google showcased uses of Gemini Omni Flash for multi-angle video generation and editing @Google https://x.com/Google/status/2086814383582118356 , while fal added both MiniMax H3 LoRA training @fal https://x.com/fal/status/2086883706891808867 and Seedance 2.5 endpoints @fal https://x.com/fal/status/2086927528032145450 . The multimodal creator stack is becoming increasingly composable: reference images, audio, first/last-frame control, and LoRA fine-tuning are being treated as standard primitives rather than special demos. Robotics/world models also had a notable release : Dyna Robotics https://x.com/DynaRobotics/status/2086856327150858298 introduced Dyna-2 , a world-action model pretrained on 1 million hours of human video , claiming new scaling laws: scaling on human video transfers to unseen robot data, and objective choice matters for cross-embodiment transfer. Separately, Sakana AI https://x.com/SakanaAILabs/status/2086829673699316179 framed its expanded RSI Lab around “Physical AI,” world models, and recursive self-improvement for real-world agents. Top tweets by engagement Meta / Muse Glimmer launch : Mark Zuckerberg on Glimmer + Spark 1.2 https://x.com/finkd/status/2086755195535413696 , Alexandr Wang’s launch thread https://x.com/alexandr wang/status/2086756152034066792 , and Meta AI’s official model thread https://x.com/AIatMeta/status/2086757844544811485 . Anthropic math result : Claude improves RH-related lower bound from 41.6% to 67.2% https://x.com/AnthropicAI/status/2086867246073401655 . OpenAI cyber model : GPT-5.6-Cyber announcement https://x.com/OpenAI/status/2086864365379010729 . Claude Sonnet 5 pricing : Permanent $2/M input, $10/M output https://x.com/claudeai/status/2086891169217122586 . Open-source ecosystem reaction : Andrew Ng thanking Meta for open-weight contributions https://x.com/AndrewYNg/status/2086845515665166398 , Clement Delangue: “Meta is back” https://x.com/ClementDelangue/status/2086760700014203090 , and Yuchen Jin on open-source AI momentum https://x.com/Yuchenj UW/status/2086849057306325243 . AI Reddit Recap /r/LocalLlama + /r/localLLM Recap 1. Meta Muse Glimmer 30B Local Release Activity: 2141 : Introducing Muse Glimmer: an open-weight model optimized for always-on local agent workflows https://www.reddit.com/r/LocalLLaMA/comments/1vkgsum/introducing muse glimmer an openweight model/ Meta announced Muse Glimmer, a dense 30B open-weight multimodal agent model under Apache 2.0, supporting interleaved text+image inputs via a dedicated perception encoder, 100+ languages, controllable reasoning effort, and agent benchmarks such as DeepSearch QA, MCP-Atlas, τ³-Bench, and SWE-Bench. The release targets local always-on workflows: ~ 4-bit quantization brings the LM below 20 GB , leaving room on 24–32 GB systems for KV cache, perception encoder, and a bundled DFlash-based speculative decoding drafter; weights are on Comment sentiment was largely enthusiastic about Meta returning to open-weight releases, but there was no substantive technical debate in the top comments. Hugging Face https://huggingface.co/meta-models , with planned support for Ollama, LM Studio, Unsloth, torchtitan, llama.cpp, MLX, ExecuTorch, vLLM, and SGLang. A top comment cites Alexandr Wang saying an open-weight Muse Spark 1.2 release is coming soon on X https://x.com/alexandr wang/status/2086756152034066792 .A commenter cites Alexandr Wang saying on X that Meta/Scale ? will be releasing an open-weight version of muse spark 1.2 soon , which is the only concrete model-release detail in the thread: