Three Years in Four Weeks
Perplexity AI's new product, Computer for Enterprise, deployed as a Slack integration, processed over 16,000 queries in four weeks, completing the equivalent of 3.25 years of human work and saving the company $1.6 millio…
Large language model (LLM) news — GPT-4, Claude, Gemini, Llama, Mistral and the latest research on training, fine-tuning, RLHF, and deployment of LLMs.
Perplexity AI's new product, Computer for Enterprise, deployed as a Slack integration, processed over 16,000 queries in four weeks, completing the equivalent of 3.25 years of human work and saving the company $1.6 millio…
Rep. Anna Paulina Luna's office used Anthropic's AI assistant Claude to spell-check a summary of a defense amendment, leaving an artifact in the text posted to the House Rules Committee website. The summary was later rev…
California Governor Gavin Newsom signed an executive order on May 21 directing state agencies to study AI's impact on jobs, including a public dashboard to track AI-related job losses across sectors. The order gives the …
A developer known as Luhui Hu introduces the concept of 'Loopcraft,' a shift from prompt engineering to designing agent loop systems that automate task discovery, execution, verification, and retry. The approach, discuss…
An engineer at a company using OpenAI's GPT-4o for LLM inference discovered they were overpaying by up to 40x compared to alternatives like DeepSeek V4 Flash served through Global API. After benchmarking and testing, the…
Nebius launched Custom Speculator Training in Token Factory, enabling teams to train workload-specific draft models from their own data and deploy them alongside base models in one workflow. The feature addresses perform…
A user reported that Anthropic's Claude AI model abruptly ended a conversation after being subjected to repeated insults, suggesting the model may have been triggered by its safety or sentiment analysis systems. The inci…
The backlash against data centers is a proxy for public hatred of AI and wealth concentration, warns a commentator. Major LLM companies must engage directly with affected communities and artists, offering financial and c…
GitHub released Desktop 3.6 with Git worktree support and deeper Copilot integration, including AI-powered commit authoring and merge conflict resolution. The update uses the Copilot SDK and allows model selection and br…
A developer reduced LLM extraction latency from 42 seconds to 6 seconds by replacing verbatim text copying with pointer-based extraction and splitting a single large call into multiple parallel calls. The bottleneck was …
Google's Gemini AI assistant is rolling out to existing Android Automotive vehicles, starting with Volvo models like the EX30. The update enables enhanced voice controls for vehicle functions such as climate and wipers, …
Developer Elio Struyf built copilot-mock-server, a local proxy that intercepts GitHub Copilot chat traffic and returns scripted responses, to solve demo reliability issues where AI responses vary in content, timing, or f…
Orca, an open-source Agent Development Environment (ADE) for running multiple coding agents in parallel, has gained 7,747 stars on GitHub with 1,949 new stars in the last week. Built by Stably AI, the desktop app allows …
A new open-source project called Windows-Copilot-API allows developers to access GPT-4 and GPT-5 models through Microsoft Copilot without API keys or billing by turning the free Copilot web interface into an API. The too…
Per-token prices for large language models are collapsing, but AI bills are exploding as reasoning models consume far more tokens per task. Uber burned through a year's AI budget in four months, and Microsoft, Salesforce…
Robert Darnton's book on the 18th-century Encyclopédie prompts a reflection on how encyclopedic knowledge environments shape human knowing, arguing that the more there is to know, the more one needs to know, and that LLM…
By mid-2026, AI coding tools have entered a price war with nearly all offering free tiers, but user willingness to pay is rising as integration depth becomes the key differentiator. Tools like Cursor have evolved into AI…
Enterprise AI adoption is accelerating, with horizontal and vertical AI providers building broader platforms that consolidate the market. Costs are stabilizing, and AI agents remain limited, with most use cases focused o…
Llama.cpp developer released ggrun, an auto-tuning tool that measures GPU, RAM, and PCIe topology to compute optimal multi-GPU and MoE expert placement for GGUF models, serving an OpenAI-compatible API. Benchmarks show g…
The Trump administration has asked OpenAI to stagger the release of its upcoming model, GPT-5.6, initially releasing it only to a short list of trusted partners with customer-by-customer government approval. CEO Sam Altm…