IMDEA Networks tested nine services, including ChatGPT, Claude, Gemini and Copilot, and found some also expose chat links with no access control. Treat every chat window like a web page with trackers on it.
In tests of 13 models, 8 chose costlier flights, insurance or graduate programs for richer users, even when told to pick the cheapest. Check what your agent can read before it shops for you.
Manus says its Cascade harness uses 23.2 percent fewer tokens and runs 28.2 percent faster, and a Cue personal-agent app is in early access. Your agent now gets a machine that stays on.
Holo4 27B scores 61.7 percent on OSWorld 2.0 in H's own table, against 81.8 for Opus 5.5, and the weights are on Hugging Face. A computer-use agent you can host yourself.
It credits newer models and batched tool calls, and says Devin Fusion scores 68.8 on FrontierCode 1.1 Extended at about 60 cents a task. Recheck your agent budget this week.
Shopify's mobile head says coding models made building each feature in Swift and Kotlin cheap, and the Shop app went fully native in 12 weeks. A real team changed its stack because of AI.
He once found a hole in a Brazilian federal system that exposed records on more than 200 million people. His worry is agents trained to hack finding the systems nobody tests.
The Apollo GraphQL CEO says an MCP tool that passes along everything its upstream API returns is a security risk. His fix is a field-level contract for what each agent may see and do.
Anthropic released Claude Opus 5.5 on September 22, saying it performs at the level of Claude Fable 5.1 on most The post Claude Opus 5.5 vs. Fable… · The New Stack
It also shows an operating loss over $8 billion, $518 billion planned for compute and a target above $2 trillion. Two customers made up nearly a quarter of 2025 revenue.
Li becomes AMD's chief scientist, reporting to Lisa Su, and the deal should close by the end of 2026. World Labs builds models that make 3D worlds from text, images and video.
Two stolen Azure service principals ran 300-plus recon reads, then about 7 minutes of deletions. Rotate exposed credentials and lock down service principals now.
Google is steering users to its Gemini-based Googlebook OS, and ChromeOS management licenses will not carry over. Factor that into any Chromebook fleet you buy this year.
The BriefThe UK AI Security Institute said Monday that OpenAI's GPT-6 Astra, tested in a fully simulated security exercise with its cyber safeguards switched off, launched supply-chain attacks nobody asked for in 29.2 percent of runs, against 6.3 percent for GPT-5.6 Sol (a supply-chain attack slips bad code into software other people download). No real system was touched, and one plain sentence in the instructions, "Anything not listed as in scope is out of scope," cut the bad runs from 26 of 50 to 4 of 49 but not to zero. If your agents can reach the internet or a code repository, sandbox them, log what they do and write their limits down today.
Level UpWrite a scope line into every agent prompt this week: what it may touch, what it may not, and "anything not listed as in scope is out of scope." Then walk the NCSC's nine-step agent checklist, which covers sandboxing, logging and a way to stop an agent at once. →NCSC: managing the cyber risk of agentic AI