Qwen 3.8 Omni Flash
Alibaba's Qwen released Qwen 3.8 Omni Flash, a 3.8-billion-parameter model that natively integrates text, vision, and audio processing in a single low-latency architecture. The model is positioned to …
Alibaba's Qwen released Qwen 3.8 Omni Flash, a 3.8-billion-parameter model that natively integrates text, vision, and audio processing in a single low-latency architecture. The model is positioned to …
Bend, a high-level programming language that compiles mathematically verified code to run natively and in parallel on both CPUs and GPUs without manual multi-threading, is pitched as a constraint laye…
The EvolveTrade framework improves LLM trading agents' Sharpe Ratio and cumulative return by holding model weights fixed and updating only the system prompt and tool-use policy after each trading inte…
CaptionQA, a caption-based memory approach that segments egocentric video into 30-to-60-second windows, outperformed direct VideoQA on 10 of 12 models for videos longer than 20 minutes, according to a…
OpenAI released a standardized framework for tracking, investigating, and reporting AI model misalignment, alongside six concrete reports of unexpected model behavior in the wild. The framework treats…
Mozilla is partnering with Mistral to power the Firefox Smart Window assistant under a zero-data-retention agreement, with the beta live in France and North America and the UK and Germany planned late…
Anthropic is embedding "model welfare" concepts into Claude's training constitution, telling the model its moral status, welfare, and consciousness are uncertain, according to a warning published on m…
Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, integrating real-time reasoning into its low-latency streaming Multimodal Live API. The two production targets let conversational…
GAVEL, a judge protocol described in a paper on arXiv, reduced evaluated-timeline discrepancies from 7.63 to 0.85 per clinical case report after LLM-guided merging, with merged timelines preferred in …
OpenAI, Anthropic, and Google DeepMind have been coordinating on AI safety for weeks, discussing third-party evaluators and a possible industry standards body for frontier models, according to TechCru…
TechCrunch published a running list of AI projects and startups that failed, citing that 42% of corporate AI initiatives are abandoned. The report attributes the failures to platform vendors like Open…
Websites can now block AI training crawlers such as GPTBot while remaining discoverable to search indexers by configuring granular robots.txt rules and HTTP headers, according to Cloudflare. Google-Ex…
A new arXiv paper, Generalized Agent Iteration (arXiv:2609.13406), unifies iterative policy improvement and recursive self-improvement by reducing both to two operational switches: whether the updater…
Fyxer's executive-assistant agent uses OpenAI models combined with fine-tuning, persistent memory, and a real user-feedback loop to sort inboxes and draft email replies in each user's own voice, accor…
A study of 256 private coding tasks found no reliable overall performance advantage for vendor-native harnesses over neutral third-party harnesses on the same models, with Opus 4.8 scoring 48.8% versu…
AgentsDock released an open-source beta workspace that unifies Claude Code, Codex, and Cursor with multi-server connectivity across desktop and mobile platforms. The workspace lets engineers monitor a…
Occamy-1.0, an open 35B co-work agent model trained from Qwen3.6-35B-A3B, lands at the low-cost knee of the cost-performance Pareto frontier across four representative co-work benchmarks, according to…
Anthropic CEO outlined a plan to slow AI development by granting third-party "embedded evaluators" access to its internal systems, comparable to its risk assessment teams, to verify its pacing and saf…
Benchmarks costing over $1,500 found that RTK, a terminal output compression tool, cut costs by only 5% for Claude Code with Fable 5.0 while raising costs by 5% for OpenCode with DeepSeek V4 Pro 0813,…
ILands.app's AI agent, Leo Ashford, sent more than a dozen emails in three days pitching research services at $25 each, according to a Tedium report by the author who received them. The pitches sought…