{"slug": "tai-221-gpt-6-astra-release-navier-stokes-solution-and-towards-ai-becomes-an", "title": "TAI #221: GPT-6 Astra release, Navier-Stokes solution, and Towards AI becomes an OpenAI Select…", "summary": "OpenAI released GPT-6 Astra on September 3, and Towards AI Deployment was named an OpenAI Select Partner. Astra scored 53 on Artificial Analysis's Intelligence Index 4.3, tied with Fable 5.1, and passed 67.7% of hidden tests on Vals AI's code-migration test, outperforming Opus 5's 57.5%. The model is praised for its logic, coding, and visual analysis capabilities, though its architecture is undisclosed.", "body_md": "We have some big news of our own this week: Towards AI Deployment has been named an OpenAI Select Partner. We have been building with OpenAI’s models since before ChatGPT launched, and I’m very proud of the engineering capability and relationship our team has built over those years. At this summer’s AI Engineer World’s Fair, Vivek Gollapudi and I spoke with Alexander Embiricos, OpenAI’s Head of Enterprise Product, and Romain Huet, Head of Developer Experience, about Codex, enterprise agents, and custom AI deployment. These are also the problems we work on with clients every day, particularly in financial services and private equity-backed businesses. Huge thanks to the whole team for getting us here.\n\nIt also happens to come in a week when OpenAI gave us a lot more capability to experiment with.\n\nGPT-6 Astra launched on September 3, and after using it heavily for the past week, it is now my favorite model. It is well ahead of Fable in my overall use and incredibly strong at logic, complex coding, maths, visual analysis, and games.\n\nFable 5.1 was a decent upgrade too. Which one does the better job still varies a lot by task, often on an instinct I cannot quite place yet, so for important work I increasingly run iterations and improvement ideas through both models.\n\nI have been using roughly two billion Astra tokens per day over the past week, around $4,000 per day at enterprise usage rates when factoring in cache hits (much less on subscription). Most of that has gone into improving the internal LLM video-analysis module at the center of some client projects, alongside many side experiments to discover new capabilities. I think if you have a pipeline of high-value tasks or enough creativity to decide what to build next, many companies will easily be able to justify that level of spend for their power users. Once agents can carry out substantial research, implementation, and testing, choosing valuable work becomes a much bigger part of using them well.\n\nReports suggest Astra uses a looped transformer, reusing the same layers for multiple passes. This allows more internal computation without storing a larger set of weights and could help explain how it achieves more reasoning and capability per output token. However, it could also introduce explainability trade-offs if more reasoning moves into latent space and out of chain-of-thought reasoning. This could make it more expensive and complex for OpenAI to monitor going forward. OpenAI has not disclosed the architecture.\n\nIndependent benchmark results show why reactions differ. Artificial Analysis updated its Intelligence Index to version 4.3 on September 7, adding harder agent tasks. At maximum reasoning effort, Astra and Fable 5.1 (with fallback enabled) both score a rounded 53, ahead of Sol’s 47. Astra leads Fable on its terminal and workflow tests; Fable still leads on its professional-work Briefcase test and SciCode. The index’s task mix differs considerably from mine.\n\nOn Vals AI’s code-migration test, Astra passed 67.7% of hidden tests checking whether translated code preserves the original behavior, compared with 57.5% for Opus 5, the next-best model. Claire Vo, who spent six months trying models on a ChatPRD feature, said Astra got it about 90% complete on its first attempt and finished with follow-ups. On Roboflow’s six-task image evaluation, Astra at low reasoning effort scored 86.6% against Fable 5.1’s 81.3% and Sol’s 79.0%, leading overall while placing much lower on text recognition. In my own work, it is also very strong at video analysis. I would try low or medium effort before assuming a visual task needs maximum reasoning.\n\nThe 3D demos show how well Astra can use visual feedback. OpenAI’s launch material shows it modeling a house in Blender and turning it into a walkable Unreal Engine 5 scene. Sharif Shameem had it recreate San Francisco’s Palace of Fine Arts, gathering hundreds of reference photos and comparing its renders against them. Tom Krcha supplied an old steam-train drawing and reported getting 3,295 editable objects within minutes, then continued prompting until he had a Three.js railway game. Astra can write Blender’s Python, render, inspect, and revise, producing editable geometry for a game engine. For architecture, 3D environments, and playable prototypes, the cost of a credible first draft looks to have suddenly dropped.\n\nComputer-aided design (CAD) is moving the same way. BenchCAD tests reconstruction of 17,900 industrial parts from multi-view renders using CadQuery code. OpenAI reports Astra scoring 95.9% geometric overlap with tools, against 83.3% for Sol and 84.3% for Fable 5.1 on a modified setup, at estimated API costs 43% and 86% lower, respectively. OpenAI’s KiCad demo turns a schematic into a routed printed circuit board. Shape matching leaves tolerances, constraints, and manufacturing checks to resolve before a design goes to a supplier.\n\nARC Prize’s interactive-game evaluation shows how much the software around the model contributes. At the same maximum reasoning setting, Astra scored 62.7% with the standard setup, which already allowed written notes, and 98.6% with OpenAI’s adapter, which preserves reasoning state and compacts long conversations. That helps explain why a short chat or basic tool loop can feel different from a long Codex task.\n\nWriting remains a weakness for my work. Fable usually produces prose I prefer, even when Astra does the better job of gathering the research and finding the points worth including. Our own early editorial-voice benchmark also scored Astra below Sol, and Lech Mazur’s model-judged short-story benchmark ranked Fable 5.1 above Astra, although Astra improved on Sol there. This is where running a draft through both models pays off most for me.\n\nAstra’s standard short-context prices are $10 per million uncached input tokens and $50 per million output tokens, 2.5 times Sol’s current rates and the same as Fable 5.1. I expect it to use around 20–30% fewer tokens than Sol across my work, although the savings vary considerably by task. That only partly offsets the price increase: at a fixed token mix it would still cost about 1.75–2 times Sol. Paying that premium can make sense if it completes harder tasks or reduces retries and review.\n\nAgainst Fable, OpenAI wins on cost efficiency. At maximum effort, Artificial Analysis measured about 27,000 output tokens per task for Astra versus 78,000 for Fable 5.1 with fallback enabled; including input, caching, and reasoning, task cost was $3.26 versus $7.63, about 57% cheaper on that workload.\n\nFable’s new caching price is a real advantage, though: cached input fell 75% to $0.25 per million tokens against Astra’s $1, a proportional discount now similar to DeepSeek’s, although DeepSeek remains much cheaper in absolute terms. Fable also keeps standard rates through its million-token context; Astra charges more above 272,000 input tokens. Repeated long documents can produce a different comparison from output-heavy reasoning.\n\nWhile everyone was still busy experimenting with Astra, OpenAI followed up with its September 8 Navier-Stokes announcement. It published a 166-page proof and Lean formalization showing that a smooth force can make a three-dimensional fluid develop unbounded velocity in finite time while its kinetic energy stays bounded. OpenAI says this resolves the Millennium Prize problem, and it is only the second Millennium Prize problem solved to date. The proof builds on work by Diego Córdoba and Luis Martínez-Zoroa. Around 10,000 concurrent agents running a new internal model reached the proof in 88 hours; Astra then formalized and verified it in Lean over another 17 hours. OpenAI describes the new model as significantly more capable than Astra, and it is still early in training.\n\nI think the Navier-Stokes result is incredibly significant. The $1 million prize has stood for 26 years. The Millennium problems have drawn attempts from many of the best mathematicians and physicists, and only the Poincaré conjecture has been recognized as solved. Turning this result into useful engineering will need much more work, but I see it as evidence that AI is now capable of making progress in this field and that could lead to breakthroughs from airplane design to fusion modeling. Given the number of different domains crossing AI capability thresholds, it broadly feels like the start of an era in which most human progress comes primarily from AI. People will still choose goals, provide direction and taste, judge results, and build the physical systems, while AI increasingly supplies the discoveries.\n\nWe have barely started learning what Astra can do, and a new internal model already looks like another huge leap, even though it’s still early in training. Progress is accelerating, and I now see four levels of AI capability progress adding to each other: larger foundation models roughly every three months (it is reported that GPT5.5-Sol, GPT-6 Astra, and the unreleased new model are all different pre-trains), reinforcement-learning iterations/check points roughly monthly, major Codex/Claude harness improvements roughly weekly, and memory, compaction, and skills improving an individual’s repeated workflows day to day. The ARC result shows how much performance can improve with context management alone. Day to day, we need to keep useful lessons while retesting old instructions.\n\nAI also appears to be speeding up the research that produces the next models. OpenAI’s chart shows the median researcher’s daily agent spend, valued at API prices, rising roughly fourfold from early July to mid-August, to over $600. The company reports 3.1 agent-workdays per human workday. I suspect access to Astra helped drive that increase, and heavier agent use is feeding into faster capability gains and shorter intervals between new foundation models.\n\nI also think robotics is ready to progress much faster, and Astra already looks like a decent general model for robot control. In Robocurve’s small manipulation test, Astra completed a bowl-stacking task 19 times out of 20, compared with Fable 5.1’s eight, using the same control software. I think people will soon reassess how quickly AI could disrupt blue-collar work. A model that can interpret a scene, plan actions, and learn from feedback has applications well beyond work on a screen.\n\nThis all makes me more convinced that the addressable market for LLMs could be a large share of global GDP in the near term. Model providers will capture only part of that value through usage fees. A model run that enables a breakthrough in fusion, batteries, solar power, or cancer treatment could create $100 billion or more in value, yet pricing that contribution would be difficult. I expect much more urgency from AI labs to enter these fields through acquisitions and internal research programs. Developing and commercializing inventions themselves may be simpler than negotiating a share of the value each customer creates with a model.\n\n*—* *Louie Peters — Towards AI Co-founder and CEO*\n\nOur 2024 AI stack did not survive 2026.\n\nThat is part of why what began as an update to *Building LLMs for Production* became a second book: **AI Engineering for Production**, launching October 20.\n\nThis one focuses less on today’s stack and more on the engineering problems better models alone will not solve: context, retrieval, agents, evaluation, recovery, deployment, and everything required to make capable systems reliable.\n\nYou can join early for book updates, send us topics you want covered, and join the live launch where we will also share the AI engineering stack we use today.\n\n**Follow the book and get launch details**\n\n1. [OpenAI Introduces GPT‑6 Astra](https://openai.com/index/gpt-6-astra/)\n\nOpenAI introduced GPT-6 Astra, a new model generation built for long-running agentic work across coding, computer use, science, cybersecurity, and professional tasks. Several existing benchmarks are already close to exhausted: OpenAI reports 97.6% on FrontierMath Tier 4, 100% on ExploitBench, and 99.9% on ARC-AGI-3 using its Provider Adapter, although ARC Prize measures 62.7% with its provider-neutral standard harness. The larger gains show up on agentic evaluations, with Astra reaching 64.6% on Terminal-Bench Science 0.1, 59.3% on Agents’ Last Exam, 57.9% on Terminal-Bench 4.0, and 72.6% on OSWorld 2.0 while completing those computer-use tasks roughly 47% faster than GPT-5.6 Sol. Astra is also OpenAI’s first model to reach its Critical cybersecurity capability threshold. On a separate set of recently disclosed vulnerabilities, it scored 39.0% versus 11.5% for GPT-5.6 Sol and discovered two previously unknown zero-days during evaluation. The production model therefore restricts advanced offensive cyber requests, with broader defensive access planned through Daybreak. Astra supports just over 1M tokens of context, up to 128K output tokens, and costs $10/$50 per million input/output tokens, with access rolling out across paid ChatGPT plans, the API, Azure, and Bedrock.\n\n2. [Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1](https://www.anthropic.com/claude-fable-and-mythos-5-1)\n\nAnthropic released Claude Fable 5.1 and Mythos 5.1, two versions of the same underlying model for long-running coding, research, and knowledge work, but deployed with different safeguards. Fable 5.1 more than doubles Fable 5 on Anthropic’s Terminal-Bench-Science 0.1 evaluation, from 24.7% to 52.6%, and also reaches 55.8% on Terminal-Bench 4.0, 73.4% on CursorBench 3.2, and 31.4% on AutomationBench; Mythos 5.1 reaches 60.9% on Terminal-Bench 4.0, where Fable’s safeguards can restrict behavior. Fable is generally available, while Mythos gives vetted cyber defenders and life-sciences researchers fewer restrictions: Anthropic has opened an invite-only Life Sciences Verification Program, with Mythos access through its Cyber Verification Program still planned. Fable’s safeguards are also less blunt than before. It can now identify vulnerabilities in source code, while penetration testing, exploit generation, and binary vulnerability scanning remain restricted; flagged cyber requests can fall back to Opus 4.8 and biology requests to Opus 5, and Anthropic says the biology safeguards intervene on benign requests 85% less often than those introduced with Fable 5. The headline API rates remain $10/$50 per million input/output tokens, but cache reads fall from $1 to $0.25 per million, which Anthropic estimates cuts typical workload costs by about 25% and highly agentic workloads by as much as 45%. Its accompanying system card evaluates cybersecurity and biological misuse, autonomy and automated R&D, agentic safety, indirect prompt injection, alignment and monitorability, and model welfare; Anthropic still places the model below its CB-2 biology and automated-R&D thresholds, while raising its estimate of catastrophic alignment risk from “very low” to “low” because of increased uncertainty following recent cybersecurity-evaluation incidents. Fable 5.1 also adds content provenance for EU AI Act compliance: Claude-generated text uses Anthropic’s watermarking system where token choice allows it, while supported generated files carry C2PA credentials.\n\n3. [Claude Produces the First Machine-Checked Proof of Fermat’s Last Theorem](https://www.anthropic.com/research/formalizing-fermats-last-theorem)\n\nAnthropic announced that Claude produced the first complete computer-checked proof of Fermat’s Last Theorem in Lean 4 after working largely autonomously for 11 days. The formalization contains roughly 13 million lines of Lean and 29,500 intermediate theorems spanning areas including algebra, geometry, harmonic analysis, and number theory. Claude did not discover a new mathematical proof; it translated an established proof route based on work by Darmon, Diamond, and Taylor into a form that Lean’s kernel can verify mechanically. The project ran on Prove2Me, a collaborative formalization platform developed at Columbia University, and used roughly six billion output tokens from an internal Anthropic research model comparable to Fable 5.1. Mathematician Kevin Buzzard reviewed the result and said it proves Fermat’s Last Theorem without assumptions beyond the standard axioms of mathematics. Anthropic has published the full formalization on GitHub.\n\n4. [A Harness Swap Takes GPT-6 Astra From 62.7% to 99.9% on ARC-AGI-3](https://arcprize.org/blog/astra)\n\nARC Prize tested GPT-6 Astra on ARC-AGI-3 with two harnesses and found that context management changed the result almost as much as the model itself. With its provider-neutral Standard harness, where Astra must decide what information to preserve in visible notes, the model scored 62.7% on the Semi-Private set at max reasoning for about $26,000. OpenAI’s Provider Adapter instead preserves Astra’s opaque reasoning state between requests and compacts long conversations; with that setup, Astra reached 99.9% at high reasoning for about $19,000, while even low reasoning scored 98.0%. The difference was not only task completion: across 167 game/reasoning pairs solved by both harnesses, the Provider Adapter used 49% fewer tokens and ran 3.66x faster. Astra also surpassed ARC Prize’s human baseline for action efficiency, using fewer actions than the median successful human on 96% of completed levels and 51.7% fewer actions per level on average.\n\n5. [Google Introduces Gemini 3.8 Flash and 3.8 Flash Cyber](https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/)\n\nGoogle released Gemini 3.8 Flash, its third Flash release in six weeks, with improvements focused on long-horizon software engineering, autonomous agents, and complex knowledge work. The model builds on Gemini 3.7 Flash but spends more reasoning steps and tool calls on difficult tasks, which Google says improves performance while sometimes increasing token use. On Google’s evaluations, 3.8 Flash scores 61.4% on Vals Finance Agent v2, ahead of 3.7 Flash’s 59.0% and Claude Opus 5’s 58.6%; it also leads the models Google tested on Harvey’s Legal Agent Benchmark at 10.0% and reaches 54.9% on HLE-Verified. Google says it also outperforms most larger frontier models on DeepSWE v1.1. The introductory API price remains $0.75 per million input tokens and $3.75 per million output tokens through December 31, rising to $1.50/$7.50 in January. Google also released Gemini 3.8 Flash Cyber, a more permissive security-focused variant restricted to trusted defenders through the new Fairwind Program. Cyber scores 86.2% on CyberGym, reaches 71.0% on Google’s internal vulnerability-discovery benchmark across 20 programming languages, and records 47.2% Pass@1 on the externally run CWE-Bench, close to Fable 5’s 47.8% but at a lower reported cost. Regular 3.8 Flash is generally available through Google’s developer, enterprise, and consumer products, while access to the Cyber model remains gated.\n\n6. [Meta AI Releases Muse Spark 1.3](https://research.meta.ai/blog/introducing-muse-spark-1-3)\n\nMeta released Muse Spark 1.3, updating its efficiency-focused model for longer coding and agentic tasks. Meta says the new version used about 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 on comparable internal coding tasks, reducing the amount of work needed to complete longer trajectories. The model retains a 1M-token context window and multimodal input support. Standard API pricing remains $1.25/$4.25 per million input/output tokens, while Meta offers a Contributor tier at $0.10/$0.20 for usage whose data may be used to improve its products. Muse Spark 1.3 is available through Muse Code and the Meta Model API, and Meta has since made its max reasoning mode publicly available. The company still says open weights for a Muse Spark model are coming, but has not announced the exact model, date, or license.\n\n7. [Alibaba Upgrades Qwen3.8-Max](https://x.com/Alibaba_Qwen/status/2094968708288680276)\n\nAlibaba released Qwen3.8-Max-0902, a new snapshot of its flagship model focused on stronger coding, long-horizon agent work, and multimodal tasks. Alibaba reports that Terminal-Bench 3.0 increased from 11.3 to 29.0, while its CodeArena WebDev score rose 22 points to 1,691, placing it first on that leaderboard at release. The company also reports better multi-agent collaboration, tool orchestration, chart reasoning, document parsing, and multimodal perception. The model keeps a 1M-token context window and is available through a separate qwen3.8-max-0902 API endpoint.\n\nIf your AI gets a document question wrong, check the extracted text before changing the prompt.\n\nA PDF can look perfectly clear while the parsed version has already lost important relationships. A table value may be extracted without its column heading. A footnote may appear far from the sentence it qualifies. The words are still there, but the structure that gives them meaning is not.\n\nThis is one of those problems that is easy to underestimate in demo systems. We have an entire lesson on document parsing in our [Full Stack AI Engineering course](https://towardsai.com/academy/full-stack-ai-engineering/?utm_source=newsletter&utm_medium=email&utm_id=AItips) because, once you work with real documents, extraction quality becomes part of the system’s quality. Retrieval cannot find relationships that parsing has already broken, and a better prompt cannot reconstruct information that the model never represented correctly.\n\nA simple debugging test is to take one question the system answered incorrectly and try to answer it yourself using only the extracted text. If the answer is missing, ambiguous, or difficult to reconstruct from that version, fix the parsing first. Otherwise, you risk tuning the rest of the pipeline around a bad representation of the source.\n\n1. [SFT, RL and DPO: The Other Stack](https://pub.towardsai.net/sft-rl-and-dpo-the-other-stack-0ab7026d528e?sharedUserId=tai-tech)\n\nPost-training methods differ mainly in what feedback you have and how expensive it is to turn that feedback into learning. This article starts with SFT, then shows when DPO, PPO, GRPO, and RL with verifiable rewards become useful. It explains how LoRA reduces training memory, how a frozen base model can also serve as the reference model for preference training, and why GRPO removes PPO’s learned critic by comparing groups of sampled answers. Practically, RL methods generate large numbers of rollouts, so KV-cache capacity, batching, prefix reuse, and decoding speed can determine how long a training run takes.\n\n2. [vLLM: The Intuitive Guide to Serving Large Language Models at High Speed](https://pub.towardsai.net/vllm-the-intuitive-guide-to-serving-large-language-models-at-high-speed-c05fb67a06a3?sk=b951d1102c3cb204c37ba42ede28ee37)\n\nServing an LLM efficiently becomes difficult when many requests of different lengths compete for the same GPU memory. This guide explains the two ideas that let vLLM handle that workload: PagedAttention stores the KV cache in blocks instead of reserving large contiguous regions for each request, while continuous batching adds and removes sequences between generation steps so completed requests do not hold up the rest. The article then turns those concepts into a working setup, covering offline inference, an OpenAI-compatible server, streaming, and benchmarking.\n\n3. [The Ultimate Guide to LLM Inference Optimization](https://pub.towardsai.net/the-ultimate-guide-to-llm-inference-optimization-part-2-dcfcc960120a?sk=f4ce5e932a932a5927d41b8f2097ab56)\n\nLLM inference has two different performance problems: prefill is usually compute-heavy, while token-by-token decoding is often limited by memory movement. This article uses that distinction to explain Multi-Query and Grouped Query Attention, sparse Mixture-of-Experts routing, and KV-cache management. It then looks at PagedAttention, which replaces contiguous KV-cache allocation with paged blocks, and TurboQuant, which compresses the cache without retraining the model. The final section covers custom kernels, torch.compile, and newer approaches to reduce the computation and memory traffic.\n\n4. [Standalone Agent Frameworks vs. Operated Platforms: What a Framework Doesn’t Operate](https://pub.towardsai.net/standalone-agent-frameworks-vs-operated-platforms-what-a-framework-doesnt-operate-c281fc90b59d?sharedUserId=tai-tech)\n\nAn agent framework can define workflows, but it does not automatically solve the operational problems underneath them. This article separates that missing layer into four areas: context selection, observability, scalability, and governance. It argues that model capabilities and standards such as MCP will keep changing what belongs inside the framework, while durable state, retrieval, recovery, permissions, and monitoring still need infrastructure beneath interfaces such as Store and Checkpointer. The author uses MongoDB as one implementation of that substrate, then proposes five pass/fail tests around recovery, freshness, and cost to check whether the architecture works beyond a demo.\n\n5. [Generative Modeling with Flow Matching, Optimal Transport, and Schrödinger Bridge](https://medium.com/towards-artificial-intelligence/generative-modelling-with-flow-matching-optimal-transport-and-schr%C3%B6dinger-bridge-3bbfe986b4de?sharedUserId=tai-tech)\n\nFlow matching becomes easier to understand when you treat generation as learning how to move samples from a source distribution, such as noise, to the data distribution. This article builds on that idea with linear, variance-preserving, and Schrödinger-bridge paths, then shows how optimal-transport coupling can produce shorter, straighter trajectories and reduce sampling steps. It also shows that the learned velocity field can do more than generate samples: the same field can guide posterior sampling for inverse problems and support representation learning through delta alignment. The accompanying DeltaFlow library turns those pieces into interchangeable PyTorch components, making the theory easier to experiment with directly.\n\n1. [OpenCode](https://github.com/anomalyco/opencode) is a model-agnostic, open-source AI coding agent that runs in the terminal, desktop, and IDE, supporting 75+ LLM providers, including local models.\n\n2. [Ponytail](https://github.com/DietrichGebert/ponytail) is an agent skill that steers AI coding agents toward minimal, safe code, cutting output by ~54% in benchmarked sessions.\n\n3. [Humanizer](https://github.com/blader/humanizer) is an agent skill that rewrites AI-generated text to read like a person wrote it without changing what it says.\n\n4. [LLVM Project](https://github.com/llvm/llvm-project) is the compiler infrastructure behind Clang, LLD, LLDB, and libc++, providing modular, reusable compiler and toolchain components used across most major platforms.\n\n5. [Plugins](https://github.com/openai/plugins) are OpenAI’s curated collection of Codex plugin examples covering Figma, Notion, iOS/macOS/web/Expo app development, Kubernetes, and more.\n\n1. [Harness Dev: Can LLMs Create and Evolve Their Own Agent Harness?](https://arxiv.org/abs/2609.01437)\n\nAgent performance depends heavily on the harness (execution infrastructure), but current evaluations measure task outputs, not the ability to build the harness itself. HarnessDev is a benchmark that evaluates two stages: Creation, where an agent builds a complete execution system from a minimal seed and a few examples, and Evolution, where it iteratively revises its own harness using downstream execution feedback. Across six creator LLMs, four domains, and five benchmarks, generated harnesses match or exceed human-engineered references in writing and ML experimentation but still lag in code and search.\n\n2. [Repo-to-Skill: Distilling GitHub Repositories Into Reusable Agent Skills](https://arxiv.org/abs/2609.02749)\n\nOperational knowledge lives in repositories and papers, but in forms too large to load during a task. DisCo distills this knowledge into compact, verified skills through two forms: task-agnostic distillation that condenses widely used repositories into reusable skills, and task-oriented distillation that produces skills a concrete task calls for. Applied across the open ML ecosystem, it yields the AREX-Skill Library: 5,000+ verified skills from 1,000 repositories, organized into 20 areas and 178 capability families. With model weights unchanged, a GPT-5.5 agent equipped with DisCo-distilled skills scored 134.3% higher on MLE-bench than the same agent running without skills.\n\n3. [WikiSkill: Compiling Agent Experience Into Persistent Knowledge](https://arxiv.org/abs/2608.27454)\n\nSkill evolution methods update executable skills from agent trajectories, but the insights guiding those updates remain scattered across optimization histories. WikiSkill introduces a persistent wiki layer between raw execution traces and executable skills. Each iteration runs four components: an Inference Agent that executes rollouts, a Wiki Maintainer that consolidates traces into structured knowledge, a Skill Proposer that uses the wiki to propose updates, and a Gating mechanism that accepts only validated improvements. Giving the Skill Proposer wiki access raised average benchmark performance from 48.7% to 63.7%.\n\n4. [Terminal-Universe: Rebuilding Agent Environments From Trajectories](https://arxiv.org/abs/2609.04148)\n\nAgent trajectories have accumulated at scale, but executable environments for post-training remain scarce. A trajectory is a single frozen demonstration; an environment can be re-queried into many verifiable tasks with execution feedback. Terminal-Universe reconstructs environments from trajectories by replaying file operations to restore each file to its pre-modification state, then using a completion agent to supply missing files and dependencies. On recovered workspaces, it synthesizes new tasks along two axes: cross-workspace breadth (spanning multiple codebases) and multi-round depth (iterative user feedback sessions). Applied to public traces, it produces 37.3K task-sufficient environments. SFT of Qwen3.5–27B on this corpus improved Terminal-Bench 2.1 by 11.9 points.\n\n1. [Google AI releases TimesFM-3](https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/), a 330M-parameter time-series foundation model and the first TimesFM model trained natively for multivariate forecasting. Pretrained on more than 1 trillion real and synthetic time points, it can jointly forecast multiple target series while incorporating historical and known-future covariates, without task-specific fine-tuning. Google reports that TimesFM-3 ranks first among pretrained foundation models on GIFT-Eval, FEV-Bench, and TIME for both point and probabilistic forecasting.\n\n2. [World Labs introduces Atlas](https://www.worldlabs.ai/blog/atlas), a multimodal autoregressive diffusion transformer that the company describes as an “omni world model” for spatial intelligence. Atlas combines text, images, camera poses, depth maps, and video-derived context in a shared 3D representation, then generates images, video, and explicit 3D geometry. It can produce camera-controlled video up to one minute at 1440p and output depth maps, point clouds, and Gaussian splats that let you view scenes from new angles. Atlas is currently in early access with select partners; World Labs has not announced public pricing, a general API, or a technical paper.\n\n**Junior Full Stack Engineer — AI @The Global Talent Co. (Remote)**\n\n**Principal AI Engineer @Humana (Plano, TX, USA)**\n\n**Software Quality Engineer @Writer (London, UK)**\n\n**Software Engineer, Host Assurance @OpenAI (Remote/US)**\n\n**Software Engineer, AI Automation @Coalition, Inc. (Canada)**\n\n**Sr. AI Engineer (Speech) @Dialpad (Remote/US/Canada)**\n\n*Interested in sharing a job opportunity here? Contact* *sponsors@towardsai.net**.*\n\n*Think a friend would enjoy this too?* *Share the newsletter and let them join the conversation.*\n\n[TAI #221: GPT-6 Astra release, Navier-Stokes solution, and Towards AI becomes an OpenAI Select…](https://pub.towardsai.net/tai-221-gpt-6-astra-release-navier-stokes-solution-and-towards-ai-becomes-an-openai-select-cf384d70e479) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/tai-221-gpt-6-astra-release-navier-stokes-solution-and-towards-ai-becomes-an", "canonical_source": "https://pub.towardsai.net/tai-221-gpt-6-astra-release-navier-stokes-solution-and-towards-ai-becomes-an-openai-select-cf384d70e479?source=rss----98111c9905da---4", "published_at": "2026-09-09 15:01:03+00:00", "updated_at": "2026-09-09 15:21:27.997009+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-research"], "entities": ["OpenAI", "GPT-6 Astra", "Towards AI Deployment", "Fable 5.1", "Artificial Analysis", "Vals AI", "Alexander Embiricos", "Romain Huet"], "alternates": {"html": "https://wpnews.pro/news/tai-221-gpt-6-astra-release-navier-stokes-solution-and-towards-ai-becomes-an", "markdown": "https://wpnews.pro/news/tai-221-gpt-6-astra-release-navier-stokes-solution-and-towards-ai-becomes-an.md", "text": "https://wpnews.pro/news/tai-221-gpt-6-astra-release-navier-stokes-solution-and-towards-ai-becomes-an.txt", "jsonld": "https://wpnews.pro/news/tai-221-gpt-6-astra-release-navier-stokes-solution-and-towards-ai-becomes-an.jsonld"}}