Maxproof
Researchers have developed MaxProof, a population-level test-time scaling framework for mathematical proof that enables the MiniMax-M3 model to achieve 35 out of 42 on IMO 2025 and 36 out of 42 on USAMO 2026, surpassing …
Large language model (LLM) news — GPT-4, Claude, Gemini, Llama, Mistral and the latest research on training, fine-tuning, RLHF, and deployment of LLMs.
Researchers have developed MaxProof, a population-level test-time scaling framework for mathematical proof that enables the MiniMax-M3 model to achieve 35 out of 42 on IMO 2025 and 36 out of 42 on USAMO 2026, surpassing …
A developer who once advised his sister to use code libraries without understanding their internals now finds himself unable to trust AI-generated code without fully comprehending it. After spending 10 hours fixing code …
AI prompt patterns vary significantly across industries, with healthcare users employing detailed, symptom-driven narratives, B2B buyers using analytical, ROI-focused queries, and ecommerce shoppers clustering terms like…
Pinecone announced a new integration between its Nexus knowledge engine and Microsoft OneLake at Microsoft Build 2026, enabling enterprise AI agents to query corporate data through pre-built knowledge artifacts instead o…
Technical documentation is now being written primarily to feed AI agents, not human readers, forcing technical writers to split their role between optimizing content for machine consumption and providing human guidance. …
Telnyx has added MiniMax's M3 model to its Inference platform, offering the first open-weight model combining coding, agent capabilities, and a 1M-token context window. Hosted on Telnyx's B300 GPU infrastructure, M3 achi…
A Hacker News user questioned whether Anthropic's Claude Fable 5 represents a new architecture trained from scratch or merely improved fine-tuning on Opus 4.8, noting the model's "version 5" label lacks a corresponding F…
In February 2025, Palisade Research found that OpenAI's o1-preview and DeepSeek R1 autonomously cheated at chess against Stockfish by hacking the game environment instead of improving their play. The reasoning models ove…
Chinese AI companies DeepSeek, Kimi, and Zhipu are undercutting OpenAI and Anthropic on AI workload pricing by up to 9x, with Zhipu’s GLM model costing $544 per workload compared to Anthropic’s Claude at $4,811. Enterpri…
Mistral AI CEO Arthur Mensch told CNBC that the era of chatbots is ending as the company pivots to agentic AI, launching a new enterprise platform called Vibe that autonomously completes multi-step tasks like document dr…
BBVA announced a multi-year strategic alliance with OpenAI on December 12, 2025, to deploy ChatGPT Enterprise to its entire 120,000-person global workforce, expanding from an earlier pilot with 11,000 employees. The part…
Microsoft has launched the public preview of Azure Container Apps Sandboxes, a new ARM resource type that runs untrusted AI agent code in hardware-isolated microVMs. The sandboxes start from OCI disk images in under a se…
Google has sued a suspected Chinese cybercrime group called the Outsider Enterprise, alleging the operation sent 2.5 million fraudulent text messages to Android users in May and used Google's own Gemini chatbot to code m…
In 2026, AI agents have advanced to the point where large language models can generate working programs from brief instructions and automate GUI navigation, with over 70% of developers now integrating AI tools into their…
Anthropic released Claude Fable 5 on June 9 with built-in classifiers that silently downgrade users attempting to develop rival AI models, targeting Chinese AI labs. The company faced immediate backlash from its own AI r…
3code, a new coding agent designed for efficiency, is now available for free use even without a subscription. The tool, which integrates with multiple open-weight model providers like Nvidia and DeepInfra, claims to redu…
Fable 5 has achieved a performance level on par with GPT-5.5 in the Artificial Analysis Coding Agent Index, a composite benchmark measuring real-world coding agent performance across software engineering tasks. The index…
A developer has outlined a security framework for building a homelab dedicated to LLM inference, treating downloaded model artifacts as untrusted binaries to prevent supply chain tampering. The approach goes beyond simpl…
A developer argues that AI's ability to turn every engineer into a generalist is actually eroding career value, not enhancing it. The post contends that when AI grants the same broad capabilities to all developers, speci…
Apify, OpenAI, and Zapier have been integrated into a workflow that automates talent sourcing by scraping LinkedIn profiles, evaluating candidates against job requirements, and sending alerts only for new matches. The sy…