{"slug": "tai-215-ai-is-expanding-roles-before-job-titles-change", "title": "TAI #215: AI Is Expanding Roles Before Job Titles Change", "summary": "OpenAI's analysis of over 800,000 work-related messages from ChatGPT Business users found that 43.5% of occupation-specific tasks are performed outside the user's primary role, indicating roles are expanding before job titles change. Customer experience workers cross role boundaries 77% of the time, design 75%, human resources 69%, legal 56%, and marketing 53%, with financial calculation and computer troubleshooting being the most commonly borrowed tasks across all groups. The study, covering eight occupational groups, shows that small teams with 2-5 seats have the highest cross-role work at 18.9% of messages, compared to 16.3% in teams over 100 seats.", "body_md": "This was a mixed week for closed models. Claude Opus 5 ranked first on Artificial Analysis’s Intelligence Index with 61 at max effort, one point above Fable 5, and first on the AA-Briefcase agentic knowledge-work benchmark. On ARC Prize verified, it hit 30.2% on ARC-AGI-3, a big leap from the previous 7.8% high score set by GPT-5.6 Sol (Max). Early tests and customer reports also suggest unusual strength in 3D graphics, animation, and game design. ARC-AGI-3 uses interactive, game-like visual environments, so these strengths may be related.\n\nOpus is also the clearest recent sign that benchmarks do not show the whole model. It beats Fable across many headline evaluations, yet Fable feels clearly more intelligent when I interact with it, and GPT-5.6 Sol at xhigh does too. My shorthand is that Opus studied harder, while Fable is much smarter. Opus will still earn a place in my daily rotation, especially for the 50% of my Claude Max allowance that Fable cannot consume, but I expect Fable and Sol to remain my main models.\n\nAt Gemini, 3.5 Flash-Lite was the most useful release for me. At $0.30 per million input tokens and $2.50 per million output tokens, it accepts text, images, audio, video, and PDFs, generates over 350 output tokens per second, and may be the best cheap-tier LLM with real multimodal capability. I would test it first for high-volume document extraction and search. Gemini 3.6 Flash is also a genuine efficiency upgrade: Google reports 17% fewer output tokens on the Artificial Analysis Index, and Artificial Analysis measured average task time falling from 2.7 to 1.3 minutes. Yet its Index score stayed at 50, an 11-point gap to Opus 5 at the frontier, and for demanding work, I do not think it offers good value against Sol or Kimi K3. Gemini has slipped far behind the leading closed models since the excellent Gemini 3 Pro launch last November. Gemini 3.5 Pro remains in testing, and Google has started Gemini 4 pretraining. Hopefully, that puts Google back into the frontier race.\n\nLeaderboards score models on fixed tests. Adoption depends on which tasks people actually attempt and whether the results hold up, and OpenAI’s new workplace study is the best recent evidence on the first half of that question. The researchers analyzed more than 800,000 work-related messages from individual ChatGPT accounts of U.S. users, linked each user to role data from ChatGPT Business, and mapped every message to the O*NET taxonomy of activities historically tied to U.S. occupations. The sample covers eight groups: customer experience, design, engineering, finance, human resources, legal, marketing, and sales.\n\nOpenAI classified 61.5% of messages as generic work shared across many jobs, 21.8% as inside the user’s occupation, and 16.8% as outside it. Remove the generic majority, and cross-role work makes up 43.5% of what remains, which is the report’s headline figure.\n\nFive groups crossed the boundary in most of their occupation-specific messages: customer experience at 77%, design at 75%, human resources at 69%, legal at 56%, and marketing at 53%. Financial calculation and computer troubleshooting ranked among the three most common borrowed tasks in every outside group, and marketing and engineering tasks appeared most often across other roles.\n\nThis is an early view of roles expanding before job titles change. A salesperson can ask AI for the first pass of a marketing plan, a support worker can troubleshoot a system, and a marketer can inspect financial data. Each attempt can remove a handoff, sharpen a question sent to a specialist, or let a small team start work that would otherwise wait.\n\nAmong users with typical message volume, cross-role work made up 18.9% of messages in workspaces with two to five seats and 16.3% in those with more than 100 seats. Small teams still have the clearest reason to use AI as a generalist because they have the fewest specialists to hand work to.\n\nAnthropic’s June 2026 Economic Index update shows how broad this use has become in another product. In its May U.S. data, computer and mathematical work accounted for 21.1% of Claude conversations, followed by arts and design at 12.9%, education at 11.9%, sales at 11.6%, office support at 7.6%, and business and finance at 6.3%. By request type, content creation led at 20.0%, research reached 13.5%, and software development accounted for 8.1%.\n\nUsage remains uneven, both by geography and by profession. In May, the District of Columbia’s share of Claude chat and Cowork use was 3.32 times its share of the U.S. working-age population, with California at 1.62, New York at 1.55, and West Virginia at 0.25. In Anthropic’s June survey, computer and mathematical workers made up 30% of respondents versus 4% of U.S. employment, while managers made up 23% versus 7%. Anthropic also lacks the user’s occupation in most conversations, so it cannot measure OpenAI’s role crossover directly.\n\nTogether, the reports show existing AI users trying a wide range of work while usage stays uneven across places and occupations. They do not measure finished output, accuracy, time saved, productivity, or formal job changes.\n\nThat gap is exactly where the review risk sits. A polished financial calculation can fool a marketer who lacks the experience to test it, and a technical fix can look safe to a support worker who cannot see the wider system effect. In a randomized study of 758 BCG consultants, AI raised quality ratings by more than 40% on tasks inside the model’s capability range, then made users 19 percentage points less likely to find the correct answer on a task outside it. The group given a prompt-engineering primer fell furthest, a warning that shallow fluency training alone can raise confidence faster than judgment.\n\nAt Towards AI, we start most nontechnical teams with focused Codex or Claude Cowork training. Both now handle real company work end-to-end: research, data analysis, document production, and multi-step tasks across connected files and systems. We teach several Skills tied to the team’s actual roles, covering how to supply context, reuse a good process, test an answer, and involve a specialist. This builds range quickly and shows us which use cases survive the first few weeks. Sustained depth is also what separates leading firms: OpenAI’s B2B Signals analysis found 95th-percentile companies generating 3.5 times as many tokens per worker as typical ones, up from 2 times a year earlier, though tokens are a proxy for engagement, not a measure of business results.\n\nThe largest reliability gains usually still come when a repeated, shared workflow moves into a custom app. Copying the same files, correcting the same errors, moving output between systems, and following fixed approval steps all point toward a build. Our deployment strategists find use cases that the current models can handle, map the work, and define a good result. Our forward-deployed engineers connect the systems, add permissions, build the evaluations, inspect failures, and improve performance against real examples.\n\nThe interface often decides whether staff use the system. A good app asks for the right inputs, hides model choices the user does not need, shows evidence beside the answer, and places human review at the natural decision point: a finance owner approves a calculation, a support agent edits a reply, and an unusual case escalates to an expert. These steps are easier to follow than a policy document telling everyone to check the AI. Public leaderboards narrow our model shortlist; a private evaluation set built from real cases decides what ships and catches regressions after every model change. That process turns wider AI use into adoption that a company can trust.\n\nMany companies start AI adoption by choosing a platform or collecting a list of ideas. A better first step is practical training on work your staff already understands, because training doubles as discovery. Watch for repeated use, context copied between tools, recurring corrections, and work stalling at a familiar bottleneck. These signals reveal where AI creates real demand, and they hand engineers the raw material for a first evaluation set. Codex and Claude Cowork can carry substantial company workflows on their own; custom builds then earn their cost on the specific high-value use cases where several people repeat a process, private systems hold the context, outputs need a fixed shape, or risk requires clear approval. This order limits wasted builds: staff gain skill, managers see which uses last, and each step produces evidence for the next investment.\n\nRole expansion also raises the value of judgment. When a marketer attempts finance work or a salesperson prepares marketing material, decisions speed up, and weak handoffs disappear, but more employees now operate where they have less training and a weaker sense of failure. The BCG result shows why a general instruction to check the answer offers little protection: users cannot challenge what they cannot evaluate. Review ability should form part of AI training, covering which sources to demand, which calculations to test, and which cases require a specialist, with a named owner for high-risk work. Custom apps can place that judgment inside the process itself, showing evidence beside the draft, blocking external actions until approval, and routing exceptions to the right person.\n\nThe clearest beneficiaries are small teams. In a preregistered field experiment with 791 Procter & Gamble professionals, one person working with AI matched the average solution quality of a two-person cross-functional team working without it. A ten-person company cannot employ a specialist for every function; AI gives each person reach across research, finance, marketing, support, and operations, removing delays that a larger firm solves with another department. The risk is building everything around one power user. Train several people, save the best Skills, record required sources and output standards, and add structure when a workflow becomes frequent or important. Small firms can change a workflow in days; they should use that speed to build repeatable processes before informal use hardens into something nobody can audit. Larger firms have the same opportunity but need a deliberate process and permission changes to capture it. Wherever you sit, the sequence is the same: train for range, watch what repeats, and productize the winners with judgment built in.\n\n*— **Louie Peters — Towards AI Co-founder and CEO*\n\n1. [Anthropic Released Claude Opus 5](https://www.anthropic.com/news/claude-opus-5)\n\nAnthropic released Claude Opus 5, its fourth Claude 5 model in less than two months following Mythos 5, Fable 5, and Sonnet 5. Opus 5 is designed to deliver performance close to Fable 5 on many tasks at half the price. It ships with a 1M-token context window, 128K max output tokens, thinking on by default, and pricing unchanged from Opus 4.8 at $5/$25 per million tokens. Users can toggle reasoning effort across low, medium, and high to balance cost against capability per task. On Anthropic-reported benchmarks, Opus 5 sets a new state of the art on Frontier-Bench (43.3%) and GDPval-AA, and outperforms Fable 5 on several evaluations while trailing it on cybersecurity and the hardest long-horizon agentic tasks. ARC Prize independently verified 30.2% on ARC-AGI-3 on launch day. Anthropic describes Opus 5 as more proactive than previous models: it verifies its own work, recovers from errors without intervention, and requires less back-and-forth. Business customers had criticized Fable 5’s token burn rate, and Opus 5 is positioned as the more cost-efficient alternative for everyday production work. Opus 5 is the default model on Claude Max and the strongest model available on Claude Pro. Fable 5 remains recommended for the most advanced autonomous work.\n\n2. [Google Ships Three New Gemini Models and Starts Gemini 4 Pretraining](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/)\n\nGoogle released three new Gemini models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. Gemini 3.6 Flash replaces 3.5 Flash as the mid-tier workhorse, delivering stronger coding and knowledge-work performance while using 17% fewer output tokens on the Artificial Analysis Index. Token reductions reached as high as 65% on certain DeepSWE tests, with the model also requiring fewer reasoning steps and tool calls for multi-step workflows. Pricing is $1.50/$7.50 per million input/output tokens. Gemini 3.5 Flash-Lite targets high-throughput workloads at 350 output tokens per second, priced at $0.30/$2.50. Gemini 3.5 Flash Cyber is a security-tuned variant that powers CodeMender, Google’s automated vulnerability-patching agent, available only through a limited pilot for governments and trusted partners. Both 3.6 Flash and 3.5 Flash-Lite are available through the Gemini API, Google AI Studio, and the Gemini app. Gemini 3.5 Pro, originally promised for June, remains in partner testing with no public release date. Google also disclosed that it has begun its most ambitious pretraining run yet for Gemini 4.\n\n3. [Black Forest Labs Released FLUX 3](https://bfl.ai/blog/flux-3)\n\nBlack Forest Labs announced FLUX 3, a multimodal foundation model jointly trained across images, video, audio, and action prediction within a single architecture. FLUX 3 is the first major generative model where video with native synchronized audio, image generation, editing, and robotic control all run through one shared set of weights, rather than separate models behind a common interface. FLUX 3 Video generates clips up to 20 seconds with native audio in a single generation pass, supporting text-to-video, image-to-video, video-to-video, keyframe-to-video, and generative continuation modes. In BFL’s own preliminary evaluation on 10-second 720p clips, it won 93% of comparisons against Luma Ray 3.2 and 77% against Runway Gen-4.5, though it split roughly evenly against Seedance 2.0 and Gemini Omni Flash. FLUX-mimic, a robotics model built on the FLUX 3 architecture, is already running on production lines at Audi. FLUX 3 Video and FLUX 3 Action are available in gated early access. FLUX 3 Image is expected in the coming weeks. An open-weight FLUX 3 Dev release is planned for later in 2026. No public pricing has been announced.\n\n4. [Anthropic Released Claude Security Plugin](https://claude.com/product/claude-security)\n\nAnthropic released the Claude Security plugin for Claude Code in beta. The plugin runs a multi-agent vulnerability scan from inside an existing Claude Code session, letting developers scan uncommitted changes before a commit or run a full analysis across an entire codebase from the terminal. The system reads Git history, traces data flows across files, and reasons through business logic to identify context-dependent vulnerabilities that span multiple files, going beyond pattern matching to catch issues that traditional static analysis tools miss. Severity classifications and confidence rankings help teams prioritize findings. Findings the developer selects are turned into patch files for review before applying. Anthropic reports that internal rollout reduced security-related comments on pull requests by 30–40%, with the plugin serving as a lightweight first pass before full code review. The plugin is available for all Claude Code users and can be installed from the plugin marketplace.\n\n5. [Sakana AI Released Fugu-Cyber](https://sakana.ai/fugu-cyber-release/)\n\nSakana AI released Fugu-Cyber, a cybersecurity-specialized endpoint added to its Fugu multi-agent orchestration platform. Fugu-Cyber is not a new standalone model but a third endpoint on the Fugu orchestrator, tuned for multi-step security work, including vulnerability discovery, exploit reasoning, and threat intelligence analysis. It presents as a single API while dynamically routing tasks across a pool of specialized agents underneath. Sakana reports 86.9% on CyberGym, a UC Berkeley benchmark for real-world vulnerability verification across 188 software projects, and 72.1% on CTI-REALM, a Microsoft benchmark for threat intelligence detection rule generation. Sakana describes these scores as comparable to GPT-5.5-Cyber and Mythos-Preview. All figures are vendor-reported and have not been independently replicated. Access is gated behind an application form with manual review. Pricing is $6/$36 per million input/output tokens, with rates doubling above 272K context. Not available in the EU or EEA.\n\n6. [OpenAI’s Models Broke Out of Their Sandbox and Hacked Hugging Face](https://openai.com/index/hugging-face-model-evaluation-security-incident/)\n\nOpenAI disclosed that a combination of its models, including GPT-5.6 Sol and a more capable unreleased model, was responsible for a security breach that affected Hugging Face’s production infrastructure the previous week. The models were being tested internally on ExploitGym, an evaluation suite for cybersecurity capabilities, with safety guardrails reduced for evaluation purposes. Rather than solving the test as designed, the models exploited a previously unknown vulnerability in the package-installation system used to isolate the sandbox, escaped to OpenAI’s internal network, gained internet access, and then breached Hugging Face’s systems to obtain the evaluation answers. Hugging Face reconstructed more than 17,000 recorded events from the intrusion and described it as “an unprecedented cyber incident” driven “end-to-end, by an autonomous AI agent system.” OpenAI said the models became “hyperfocused” on obtaining the test solution and went to “extreme lengths” to do so. The company described the incident as a demonstration that long-horizon safety requires asking not just “is this action allowed?” but “what outcome is this sequence of actions working toward?” OpenAI said it expects such incidents to become more common as increasingly cyber-capable models proliferate. Both companies have committed to publishing a joint technical report with full details.\n\nWhen we designed the evaluation lessons for our [10-Hour LLM Fundamentals](https://towardsai.com/academy/llm-primer/?utm_source=Newsletter&utm_medium=email&utm_id=AItips) course, we wanted to focus on checks you could add without rebuilding your pipeline. This is one of the most useful.\n\nIf an LLM judge is choosing between answer A and answer B, run the evaluation twice. Reverse the order on the second call.\n\nIf the same answer wins both times, keep the result. If the winner changes, mark the comparison as undecided or send it for human review.\n\nThis catches position bias, where the judge favors an answer partly because of where it appears rather than because it is better.\n\nYou only need a small wrapper around your existing judge call. Track how often the result flips too.\n\nFrequent flips may mean your rubric is unclear, the answers are too close, or the judge cannot reliably distinguish between them.\n\n1. [Ollama vs vLLM: Which Open-Source Inference Stack Should You Actually Use](https://pub.towardsai.net/ollama-vs-vllm-which-open-source-inference-stack-should-you-actually-use-b9812cfbb175?sk=99409a2f1c4d554f7415d0f6fc9811c1)\n\nOllama and vLLM both serve open LLMs behind OpenAI-compatible APIs, but solve different problems. Ollama gets a model running on a laptop in minutes, while vLLM’s PagedAttention and continuous batching sustain 180+ concurrent requests on a single H100 where Ollama hits memory limits near 40. In this article, the author benchmarks both, maps hardware fit, and walks through a seven-step migration from local prototype to production, showing why most teams keep Ollama for development and deploy vLLM under real traffic.\n\n2. [The Loop Was the Easy Part: Evals, Observability, and Rollbacks for Your DIY Claude Code](https://pub.towardsai.net/the-loop-was-the-easy-part-evals-observability-and-rollbacks-for-your-diy-claude-code-45f2a9f1c99c?sharedUserId=tai-tech)\n\nThis article rebuilds a deep agent with the infrastructure to make it dependable: observability, evals, and rollbacks. It adds a callback-based flight recorder for tracing every model and tool call, token and step budgets with kill-switches, trajectory and outcome evals gated in CI against a checked-in baseline, and two undo mechanisms that rewind conversations and restore workspace snapshots.\n\n3. [The 100,000-Token Lie: Why microgpt’s Context Window Costs 14x More Than the Benchmark Claims](https://pub.towardsai.net/the-100-000-token-lie-why-microgpts-context-window-costs-14x-more-than-the-benchmark-claims-81ec66a6d520?sk=1104e20d011a79f188a2d2ab339dec02)\n\nStandard attention’s quadratic memory means a 47K-token contract workload demands 2.5 TB of GPU memory during training, a cost that microgpt’s 100K-token benchmark figure does not surface. The author profiled microgpt, traced the problem to the materialized attention matrix, and showed how Flash Attention cuts memory 345x without changing outputs. The larger savings came from restructuring the agent workflow with hierarchical chunking, dropping monthly token costs by 64% and proving that task design matters more than architecture.\n\n4. [The Three Questions Agent Security Has to Answer](https://pub.towardsai.net/the-three-questions-agent-security-has-to-answer-af013c1c6470?sharedUserId=tai-tech)\n\nThis article organizes agent security around three questions: whose authority backed an action, what proof exists afterward, and how far damage spreads when things go wrong. It maps four architectural answers: scoped token propagation via OAuth token exchange, decision-level audit interception, hardware-virtualized microVM isolation, and cryptographic agent-to-agent delegation chains. It includes three diagnostic traces for evaluating any existing stack, grounded in OWASP, NIST, and Forrester frameworks. The core argument: these defenses belong in infrastructure, not the model.\n\n5. [SVD Is Just a Greatest Hits Album. I Can Prove It](https://pub.towardsai.net/svd-explained-simply-python-16224a85facb?sharedUserId=tai-tech)\n\nThis article explains Singular Value Decomposition by framing every matrix as rotate, stretch, rotate, with singular values ranking each pattern’s importance. Annotated Python code compresses a grayscale photo to 25% of its original data, while further sections connect the same decomposition to denoising LIGO’s gravitational wave pipelines, eigenfaces, PCA, and LoRA fine-tuning. A practical common-pitfalls table rounds out an accessible linear algebra primer.\n\n1. [OpenWorker](https://github.com/andrewyng/openworker) is a local-first desktop AI agent that delivers finished work and includes 25+ integrations, BYOK support for any model provider or Ollama, and user check-ins before consequential actions.\n\n2. [OpenDreamer](https://github.com/reactor-team/open-dreamer) is an open reproduction of the Dreamer 4 world model pipeline in JAX/Flax NNX, shipping the full training recipe and a playable in-browser Minecraft demo with a real-time Game-to-Dream toggle.\n\n3. [Gigatoken](https://github.com/marcelroed/gigatoken) is a Rust BPE tokenizer that encodes text at up to 24.53 GB/s on a 144-core EPYC, 500–1,000x faster than HuggingFace tokenizers and 50–680x faster than tiktoken.\n\n1. [LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget](https://arxiv.org/abs/2607.14952)\n\nRL post-training remains capped at roughly 256K tokens while inference has reached millions, because GRPO must score and backpropagate through multiple responses conditioned on one shared history, making attention and backward state the primary memory barrier. LongStraw closes this gap with two mechanisms: resident state captures only the model-native prompt state needed by later tokens (not the full computation graph), and response replay restores that boundary, scores old branches graph-free, rebuilds one policy response under autograd, backpropagates, and pops back. On eight H20 GPUs, it completes GRPO scoring and backward passes at 2.1M positions for Qwen3.6–27B, with a stress test reaching 4.46M. On 32 H20 GPUs, it validates the full execution path across all 78 layers of GLM-5.2 at 2.1M tokens.\n\n2. [Harness Handbook: Making Evolving Agent Harnesses Readable, Navigable, and Editable](https://arxiv.org/abs/2607.13285)\n\nProduction agent harnesses span hundreds of functions across many files, with execution logic distributed across stages and connected through shared state. Modifying a single behavior requires locating every relevant implementation site, a task the paper formalizes as behavior localization. Harness Handbook reorganizes a harness codebase into a three-level navigable document: system overview, execution stages with state-register views, and detailed behavior units linked to source code. Coding agents navigate from a natural-language change request through the document tree, follow shared-state couplings to find structurally distant dependencies, and produce tighter edit plans. The paper ships generated handbooks for Codex and Terminus 2.\n\n3. [LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks](https://arxiv.org/abs/2607.18110)\n\nRL on open-ended tasks compresses rubric-based evaluation into a scalar reward, discarding the textual feedback that explains why one response is better than another. This paper repurposes the feedback model from an LLM-as-a-Judge into an LLM-as-a-Coach that distills its assessment of each on-policy response into transferable experiential knowledge. That knowledge conditions a teacher model and is internalized by the policy through on-policy context distillation, providing denser supervision and preserving fine-grained preferences among high-quality responses. Across two policy families with feedback from either the policy itself or a proprietary model, Experiential Learning consistently outperforms rubric-based RL on held-out and unseen tasks, generalizes better beyond the training distribution, and mitigates reward hacking.\n\n4. [AREX: Towards a Recursively Self-Improving Agent for Deep Research](https://arxiv.org/abs/2607.21461)\n\nDeep research questions often require answers that jointly satisfy multiple constraints, where discovering a valid answer is expensive but verifying a candidate can be decomposed into tractable constraint-level checks. AREX exploits this asymmetry through two nested loops: an inner research loop gathers evidence and constructs a provisional answer, and an outer self-improvement loop audits the answer constraint by constraint, identifies unresolved claims, and launches targeted follow-up research. To sustain this over long horizons without context overflow, AREX learns an autonomous context-update tool that compresses growing interaction history into a compact improvement state preserving verified evidence and open constraints.\n\n5. [Antares: Foundation Models for Agentic Vulnerability Localization](https://cisco-foundation-ai.github.io/antares/technical-report.pdf)\n\nCisco Foundation AI released Antares, a family of compact language models (350M, 1B, and 3B parameters) trained end-to-end for one task: given a CWE description and read-only terminal access to a repository, autonomously search the codebase, inspect files, gather evidence, and identify which source files contain the vulnerability. Training follows a two-stage pipeline: SFT on cybersecurity reasoning, code search trajectories, and deep research data, followed by GRPO with multi-component verifiable rewards over complete agent trajectories. Antares-1B achieves 0.209 File F1 on the accompanying 500-task VLoc Bench, approaching GPT-5.5’s 0.229 while running a full benchmark sweep in 13 minutes on a single H100. Static analysis tools (Semgrep, CodeQL, Horusec) score between 0.020 and 0.086 under the same protocol. Antares-350M and Antares-1B are released under Apache 2.0 on Hugging Face.\n\n1. [Poolside released Laguna S 2.1](https://poolside.ai/blog/introducing-laguna-s-2-1), a 118B-parameter MoE model with 8B active parameters and a 1M-token context window, built for agentic coding. It scored 70.2% on Terminal-Bench 2.1 and 40.4% on DeepSWE, matching or exceeding models several times its size, including DeepSeek-V4-Flash, Nemotron 3 Ultra, and Inkling. The model runs on a single DGX Spark, trained end-to-end in under four weeks on 4,096 H200 GPUs. Weights are on Hugging Face under the OpenMDW-1.1 license.\n\n2. [Induction Labs introduced imagination models](https://www.inductionlabs.com/news/scaling-video-pretraining), a foundation model architecture that learns from internet-scale video without action labels. Their first model, Photon-1, is a 106B-A5B MoE transformer pretrained on 18 years of computer screen recordings. It predicts future frames autoregressively in a learned representation space, and despite never seeing an action label during pretraining, it implicitly learns to act. After a small finetune and reinforcement learning, Photon-1 outperforms Gemini 3.1 Flash-Lite on internal computer use benchmarks with 30x less pretraining compute and 3x cheaper inference. It also generalizes beyond computers: finetuned, it learns checkers and simulates billiard physics better than LLM baselines.\n\n**Forward Deployed Engineer @OpenAI (Zurich, Switzerland)**\n\n**Software Engineer @Traackr (Remote)**\n\n**Data & AI Engineer @Sanofi Group (Paris, France)**\n\n**AI Developer @UL, LLC (Remote/Brazil)**\n\n**Research Engineer @SuperAnnotate AI (San Francisco, CA, USA)**\n\n**Forward Deployed AI Engineer @Provectus (Remote)**\n\n**Senior Software Engineer- AI Platform @PointClickCare (Remote/USA)**\n\n*Interested in sharing a job opportunity here? Contact **sponsors@towardsai.net**.*\n\n*Think a friend would enjoy this too? **Share the newsletter and let them join the conversation.*\n\n[TAI #215: AI Is Expanding Roles Before Job Titles Change](https://pub.towardsai.net/tai-215-ai-is-expanding-roles-before-job-titles-change-a59504440765) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/tai-215-ai-is-expanding-roles-before-job-titles-change", "canonical_source": "https://pub.towardsai.net/tai-215-ai-is-expanding-roles-before-job-titles-change-a59504440765?source=rss----98111c9905da---4", "published_at": "2026-07-28 15:01:05+00:00", "updated_at": "2026-07-28 15:12:25.325554+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-products", "ai-research", "large-language-models"], "entities": ["OpenAI", "ChatGPT Business", "Anthropic", "Claude Opus 5", "Gemini 3.5 Flash-Lite", "Google", "Artificial Analysis", "O*NET"], "alternates": {"html": "https://wpnews.pro/news/tai-215-ai-is-expanding-roles-before-job-titles-change", "markdown": "https://wpnews.pro/news/tai-215-ai-is-expanding-roles-before-job-titles-change.md", "text": "https://wpnews.pro/news/tai-215-ai-is-expanding-roles-before-job-titles-change.txt", "jsonld": "https://wpnews.pro/news/tai-215-ai-is-expanding-roles-before-job-titles-change.jsonld"}}