Less than a month ago we had just featured Poolside’s Model Factory with Eiso Kant on the pod (following our Paper Club coverage):
It appears that Jensen really, really liked Poolside too, as he went from investor to doing licensing their factory and hiring 109 of their employees:
Unless things changed drastically, this accounts for the overwhelming majority of the technical Poolside employees:
Eiso Kant [01:52:31]: We are hiring on every possible role in applied research and engineering in the company, from training all the way to evals to post-training architecture. Like, we are still in a world where, individuals can have massive impact. And I think our pitch to join us —I think we are one of the places where it’s the highest ratio to individual to impact, Right? L
ess than 70 people built this model. Less than 115 between engineering and researchers, like, together did this effort, and that’s a very broad definition ‘cause I put myself in the 115 list.
As the founders say, this is “not an acquisition and not an acquihire”:
We’ve been calling the Windsurf-Google and Character-Google and Scale-Meta and Instacart-OpenAI deals execuhires because usually the executives go leaving the employees with a rich payout but holding the company remaining, but this is a first time it is happening the other way around. The action amounts to founders pivoting the company extremely hard to SOMETHING, and finding an EXTREMELY comfortable golden parachute for investors and employees to continue on with the original mission or stay aboard for the new pivot:
For the last 3 1/2 years we’ve been directionally correct in a race where capital requirements went vertical. At the end of last year, we had a 6 week window in which to raise $2 billion dollars to pay for a 40,000 GB300 cluster coming online in January.We didn’t close it in time, and we lost the cluster.
and:
We also know that at 10,000-20,000 GB300s we would produce a great model that could rival the current frontier.But the scale of next year’s frontier models requires far more than an order of magnitude larger cluster. And for this the constraint today is not only capital, it isphysical data center space and contracted compute.
The compute needed to be at the frontier of the current model recipe is going vertical, and as the world accelerates along the axis of Recursive Self Improvement this will only become more evident.
To this end, the PIC infraco, spun out in Jan 2026, is also interesting in its ambitions…
We’re confused too, and the founders say they are “not ready to share the updated vision”, but everyone here is coming out with a lot of money so we’re just interested to see what’s next for everyone on the 3 different directions emerging from OG Poolside.
The only hints left to us:
We wholeheartedly believe that everything economically valuable, scientifically interesting and a lot of what will be personally meaningful is going to be underpinned by Al.
The world has not yet reached 0.1% of this transition….… We believe
human level capabilities of intelligence will be fully commoditized by open source models, while super intelligence will likelynot be.The world has two types of economically valuable problems, those that are intelligence bound, and those that are
experiment bound. The first are problems which we can solve by scaling up intelligence e.g. building software, doing accounting, solving a math theorem. The second are ones that require real world experimentation to progress, andno amount of increased intelligence without experimental results will make progress. We could put 100,000 of the world’s brightest minds together to solve cancer but without a real world experimental feedback loop, they likely never will.Today’s model revenue is from coding and soon from all of knowledge work. In the future, companies who can go beyond human level capabilities will tap into
revenue coming from scientific discoverieswhere there is a true data moat derived from real world experimentation. In our humble opinion, Al’s ultimate value will not derive from the first kind, that will become a low margin commodity, but it will from the second.Al will become the world’s most valuable scientific discovery engine.
Fascinating. Sounds like we could not have timed our AI for Science podcast better.
AI News for 8/19/2026-8/20/2026. We checked 12 subreddits,
[544 Twitters]and no further Discords.[AINews’ website]lets you search all past issues. As a reminder,[AINews is now a section of Latent Space]. You can[opt in/out]of email frequencies!
AI Twitter Recap
OpenAI and Anthropic Expand the Agent Product Surface
OpenAI pushed several desktop and builder features in one wave:@ChatGPTlaunched an** Apple Messages pluginfor ChatGPT Work/Codex on Mac, enabling message search, catch-up, drafting, and sending from the desktop app.@OpenAIDevsalso addedcollaborative editing for ChatGPT Sites**, with teammates sharing a project while Codex manages git/CI;shared read-only conversation linksandPR-context sharingfurther push ChatGPT/Codex toward being a coordination surface, not just a chat UI. On the API side,transparent backgrounds in GPT-Image-2are now in preview for reusable design assets.OpenAI’s desktop memory/workflow features continue rolling out geographically:@OpenAIDevssaid** Computer Historyand cross-app memory are now available in the EEA, UK, and Switzerlandfor Pro/Business/Enterprise Mac users, withRecord & Replayalso live there. Together, these features point to a product strategy of capturing user workflows on-device and turning repeated actions into reusable skills.Anthropic made its agent platform more composable and production-ready:@ClaudeDevsannounced general availability for computer use, browser tool, Skills API, and Files APIon the Claude Platform. TheSkills APIadds versioned reusable procedures; theFiles APInow supports expiration control,5x higher rate limits to500 RPM**, and** 1 TB/org**. Anthropic also published anAG-UI adapter for Claude Managed Agents, mapping chat threads to managed sessions and streaming text, tool calls, and thinking into custom UIs.
Model Economics, Usage Limits, and the Enterprise Shift Toward Open Models
AT&T became the clearest public case study yet for hybrid routing: the most consequential enterprise datapoint in the set came via@Hesamation, summarizing AT&T’s internal AI deployment:40% of employee AI usage already routes to open models, with a target of** 60–70%; coding costs are down 56%for only a 2% quality drop**, at** 45B tokens/day**. That supports the increasingly common view that frontier closed models remain reserved for the hardest tasks, while “good-enough” open models eat the broad middle of enterprise demand.@amirexplicitly framed this as a warning sign for OpenAI/Anthropic’s enterprise moat, while@ollamawelcomed AT&T to open models.Pricing pressure is intensifying across closed-model distribution:@eglymanannounced** GPT-5.6 Sol at 50% offthrough Router, and both@githuband@codeamplified the temporary discount for GitHub Copilot / VS Code users. At the same time, user sentiment suggests supply constraints are surfacing as usage caps rather than degraded quality:@bridgemindaicomplained that a$200/mo OpenAI Pro plan** could be exhausted in a single heavy Codex day, and@theonoted it was possible to continue consuming substantial tokens after hitting the stated cap. The broader signal: labs are still searching for the right product boundary between high-end model access and economically sustainable agentic usage.Open-weight adoption and distribution continue to broaden:@ollamasaid** Kimi K3is now rolled out to over half its subscription base with US/EU hostingand zero data retention**. On the open ecosystem side,@Googleand@osansevierohighlightedGemma surpassing 1B downloads, while@_philschmidlaunched an** Awesome Gemma**repo aggregating variants, deployment guides, and fine-tuning recipes.
Multimodal and Agent Benchmarks: Muse Spark, GLM-5.3, Gemini 3.7 Flash
Meta’s Muse Spark 1.2 had a strong benchmark day across multimodal/agentic evals:@AIatMetapresented demos spanning** visual coding, robotics planning, and audio-visual understanding**, and previewedWildArtifactBench, an internal eval using** win rates and Elo from human/agentic judgesfor practical multimodal tasks. Third-party measurements were favorable:@arenareported+2.1% net improvement** in Agent Arena, up from0.9% in v1.1, with particularly strongBash Recovery (+11.4%);@DesignArenaplaced Muse Spark 1.2**#1 for Video-to-Website**,#2 for Image-to-HTML, and**#3 for Image-to-Frontend**, while noting it sits on the** price-preference Pareto frontier**.** Zhipu’s GLM-5.3 keeps showing up in agentic/code evals**:@AutoClawAIerannounced** GLM-5.3 integration into AutoClaw**, Z.ai’s work agent. More importantly,@arenasaid** GLM-5.3 Maxshifts the Code Arena: WebDev Pareto frontier**, projecting to**#2 among open models** and**#8 overall** at1597 pts and**$3.65/M**. Separately,@ZixuanLi_resurfaced** SAO (Single-Rollout Asynchronous Optimization)as a key GLM-5.2/5.3 RL advance for stable asynchronous agentic RL. Gemini 3.7 Flash keeps accumulating “cheap and strong” evidence**:@arcprizereported** ARC-AGI-2: 84.6% at $0.25/taskand ARC-AGI-1: 95.5% at $0.12/task**, making Gemini 3.7 Flash stand out on cost-adjusted reasoning performance.@JonathanJarvisseparately called it excellent foragentic vision tasks.
Infra, Hardware, and Systems Work: Rubin, Cerebras, Linux Agents, Caching
OpenAI’s next pretraining stack is moving onto Rubin:@udayruddarrajuposted that OpenAI’s first** NVIDIA Vera Rubin racksare now installed and running the training stack, explicitly tied to next-generation frontier pre-training**.@gdbcalled it a major milestone in the OpenAI-NVIDIA partnership.** Cerebras’ CS-4 drew attention for inference scaling without a node shrink**:@kimmonismussummarized the launch as essentially doubling performance on the same5nm wafer,** 4T transistors**, and** 900k AI cores**, via redesigned power delivery and cooling. Reported specs include** 250 PFLOPs per WSE-3 Turbo**,** 43.2 PB/s memory bandwidth**, and a** 3-wafer CS-4 rackat 750 PFLOPs**. The notable claim for practitioners:** 4,400+ tok/s per user on GPT-OSS-120B**, up to** 30x fasterthan GPU-based systems. Agent runtime ergonomics are becoming a systems bottleneck**:@theoargued that** Linux materially outperforms macOS for agent workloads**, especially on filesystem-heavy operations.@Qdrant_engineshared a practical semantic-caching writeup showing57.1% hit rate,** 55.7% fewer tokens**, and**~15 ms** hit latency.@MParakhinpushedgisting as an underused production technique, citing**~40% lower end-to-end latency** and**~15% higher throughput** with better results, and linked aShopify engineering writeup.
Agents, Memory, and Harness-Centric Learning
Chroma launched a research preview of self-improving memory:@jeffreyhuberannounced** Foundation**, Chroma’s approach to agent memory, built from prior agent sessions. This landed amid a broader shift from “single-shot agent” thinking toward persistent harnesses with accumulated state, skills, and memories.The most interesting agent research in the set was about harness evolution, not model weights:@omarsar0highlighted a paper on** harness continual learning**, where prompts, memories, skills, and routing rules evolve independently of the model. The key failure mode is** harness-level forgetting**: improving one component can silently break previously reliable behavior. The proposed solution,** guarded harness evolution**, separates proposing updates from committing them, with reported**>10% gains** across textual, multimodal, and open-world tasks.Related negative results matter too:@dair_aiflagged a study showing that memory-based self-improving agents look worse once you control fortask order effects andevaluation variance.@omarsar0also summarized a paper arguing post-training agents tend tolock into an initial strategy early and spend the remaining budget on local refinement rather than revisiting the strategic choice itself.
Top Tweets (by engagement) ChatGPT desktop + Messages:@ChatGPT’s Apple Messages plugin launchwas the single biggest product tweet in the set and reflects the shift toward desktop-native, action-taking assistants.AT&T’s open-model routing economics:@Hesamation’s summaryis arguably the most strategically important enterprise datapoint:40% open now, 60–70% later, 56% coding cost reduction.** OpenAI’s Rubin racks**:@udayruddarrajuprovided a rare concrete infrastructure signal about frontier pretraining scale-up.Claude Platform GA for computer use / Skills / Files:@ClaudeDevsmarked a significant maturity step for Anthropic’s agent platform.Gemini 3.7 Flash on ARC-AGI:@arcprizereinforced Google’s positioning around strong low-cost reasoning.
AI Reddit Recap
/r/LocalLlama + /r/localLLM Recap
1. Qwen3.8-27B Quantization and Coding Benchmarks
(Activity: 2059):Introducing Qwen3.8-27B Dynamic v3 Unsloth GGUFsTheimageis a technical announcement graphic for Unsloth Dynamic v3.0 GGUF quantizations of Qwen3.8-27B, claiming>10%
**better top-1 accuracy at the same model size versus other quant providers. It highlights post-training quantization only—no QAT/QAD and no training on the imatrix calibration dataset—plus memory targets from 1-bit quants runnable on ~8GB RAM up to BF16, with evaluation framed around Divergence-300 @32, KLD, and top-1% accuracy comparisons. The linked release points to the Unsloth blog and Hugging Face GGUF repo:**Commenters were broadly positive but asked for more comparative data, especially adding the priorhttps://unsloth.ai/docs/basics/dynamic-3.0-ggufsandhttps://huggingface.co/unsloth/Qwen3.8-27B-GGUF.Qwen 3.8 27B UD 2.0 quants to the chart so users can judge whether upgrading is worthwhile. One user also noted practical hardware interest: whetherIQ4XS
can now run on16GB
VRAM without MTP.Users requested
comparative quantization metrics against the priorQwen 3.8 27B UD 2.0 GGUFs, specifically asking forKLD and/ortop-1 error lines on the graph so existing local files can be directly compared to the new Dynamic v3 quants.A technical point was raised that the new
IQ4XS quant may fit within16 GB
VRAM without MTP, which would be significant for single-GPU local inference if quality degradation remains low. Another user noted the apparent~15 GB
size for Q4_K_M, asking whether it preserves quality well enough to be practically useful.One commenter asked for more granular evaluation now that
oobabooga is involved, specificallyper-category KLD andKV-cache quantization KLD metrics similar to those shown bylocalbench.substack.com, to better understand where quantization loss appears across tasks and cache settings.
Qwen3.8-27B took a serious hit toknowledge(Activity: 758):vs 3.6Users report Qwen3.8-27B regresses vs Qwen3.6-27B on offline/weight-only factual recall, aligning with lower scores on Artificial Analysis’sCommenters generally frame this as an intentional or acceptable specialization tradeoff: Qwen 3.x may be shifting toward coding/agentic use where external retrieval is expected, while models likeOmniscience knowledge benchmark. The observed tradeoff is that Qwen3.8 appears stronger for tool calling, coding, and agentic workflows, but weaker when web/search tools are disabled for obscure trivia, historical/location identification, or airgapped knowledge retrieval.Gemma may be preferable for broad “mini Google” factual recall. One commenter explicitly preferred not allocating parameters to niche trivia if it improves coding performance.Several commenters converged on the view that
Qwen 3.8-27B appears optimized away from memorized factual recall and toward coding/agentic workflows. One user reported that with web search/fetch disabled, Qwen 3.8 regressed on niche knowledge tasks such as identifying stamps, historical locations, and old photos, while tool calling and coding were*“impressive”*when retrieval tools were available.The discussion framed the regression as a deliberate parameter-capacity tradeoff for a
27B
model: reduce obscure memorized knowledge while preserving reasoning, coding, and tool-use competence. Commenters suggested using other models such asGemma for trivia or broad factual recall, while positioning Qwen 3.x as better suited to agentic tasks that retrieve information externally before acting on it.One technically interesting speculation was around future
modular model knowledge/skill extensions, described as “neural plugins” similar to** LoRAs**. The proposed architecture would keep the base model lean while adding native domain or language competence—e.g. Japanese support or financial-services knowledge—through optional plugins rather than baking all knowledge into the base model.
(Activity: 422):I ran Qwen3.8-27B against Opus, Sonnet, GPT and others. Results inside.The image is a benchmark dashboard for the author’s home-built coding eval comparing Qwen3.8-27B, DS4 0731, GPT-5.6-sol, Opus 5, Sonnet 5, and Haiku 4.5 across algorithm tests, repo bugfix/feature tasks, wall-clock completion time, and blind-judged code quality (image). The main technical takeaway is that GPT-5.6-sol leads overall with perfect repo-task performance and near-perfect algorithms, while local models are surprisingly competitive: Qwen3.8-27B xhigh scores strongly on hard algorithms and “surgical fixes” but is much slower, and DS4 0731 achieves8/8
on both repo tiers despite being a2-bit
**local quantization. The author notes a practical tradeoff: higher “thinking” improves some hard reasoning/code-quality cases but can overthink, increase latency, and even reduce repo-task accuracy compared with medium thinking.**Commenters questioned benchmark saturation and task difficulty, arguing that if nearly all models score near the top then the eval may not distinguish frontier/local capability well. Others asked for more detail on the definitions of “algorithm” and “repo work” tasks, expected outputs, and hidden test design to make the results more reproducible and interpretable.Several commenters argued the benchmark appears
saturated, with*“all models at the top”*, making it hard to distinguish Qwen3.8-27B from Opus, Sonnet, GPT, and others. One analogy framed it as testing stronger models on tasks too easy to separate capability, implying the suite needs harder or more discriminative evaluations.A commenter requested more precise methodology for the
“algorithm” and**“repo work”** tasks, specifically asking for expanded task descriptions and expected results. This points to reproducibility concerns: without clear prompts, grading criteria, and target outputs, cross-model comparisons are difficult to interpret.One technically relevant question asked what
“DNF” means forQwen3.8 medium, in the context of a comparison between** Qwen3.8 xhighand medium**settings. This suggests the benchmark table included incomplete or failed runs, but the failure semantics were not defined clearly enough for readers to assess the result.