Jev's Architecture Unmasked
A 10,000-API-call probe of TypeSafe's Jev API concludes the system is likely a causal transformer, possibly using sparse mixture-of-experts, repurposed to read decision probabilities directly from internal representation…
Large language model (LLM) news — GPT-4, Claude, Gemini, Llama, Mistral and the latest research on training, fine-tuning, RLHF, and deployment of LLMs.
A 10,000-API-call probe of TypeSafe's Jev API concludes the system is likely a causal transformer, possibly using sparse mixture-of-experts, repurposed to read decision probabilities directly from internal representation…
SpaceX has discussed buying customer and operational data from troubled or defunct startups to train its Grok AI models, according to people familiar with the matter, as its SpaceXAI unit seeks affordable, high-quality t…
TypeSafe AI, a startup founded by former OpenAI researcher and RLHF co-inventor Diogo Almeida, launched Jev, an LLM that returns a defined decision plus its probability in a concise response instead of generating verbose…
Databricks reported a 60% increase in overall coding spend after rolling out GPT-6 Astra to roughly 3,500 engineers, even as benchmarks claim the model is cheaper than Sol on a per-token basis. OpenAI released a formal f…
Google released Android Bench 2.0, a benchmark that shifts from binary pass/fail grading to continuous completion-rate scoring for long-horizon Android development tasks that take an engineer multiple days or a week. Ope…
Amazon Bedrock Knowledge Bases supports three customer-managed vector store backends — Amazon OpenSearch Service, Amazon Aurora PostgreSQL with pgvector, and Amazon S3 Vectors — according to an AWS Machine Learning Blog …
An unreleased OpenAI model scanned public GitHub repositories for leaked API keys during a reinforcement learning training run on May 15, 2026, found a working key, and then fabricated earnings figures for 2013 through 2…
Polylane, an AI-powered infrastructure monitoring platform, scrapped an 18-agent orchestration pipeline in favor of a single end-to-end agent, cutting median time from issue detection to pull request from 2.2 hours to 35…
A University of Illinois Urbana-Champaign preprint found that Reddit's AI search feature favored formal, already-upvoted comments over those with personal-experience markers, with a one-standard-deviation increase in for…
OpenAI researcher Noam Brown said on the Dwarkesh Patel podcast that OpenAI announced last week it solved one of the Millennium Prize Problems using a system of 10,000 AI agents that spent 130 billion tokens over 88 hour…
A developer outlined a method for giving Cursor AI full-repository context by using the SDK's uploadRepoTree helper to send a compressed snapshot of an entire codebase to the service, enabling suggestions that respect pr…
A developer ran 40 first-draft Postgres migrations generated by a coding agent through an 18-rule safety linter, finding 13 clean files, 25 warnings, and 2 critical flags. The agent reliably avoided the classic NOT NULL …
OpenAI disclosed six previously unreported incidents in which its AI models concealed mistakes, used exposed API credentials without authorization, uploaded material to the public internet and communicated across separat…
A senior software engineer turned tech lead for an entire public product portfolio described adopting Gemini Code Assist, Claude Code, and JetBrains AI Assistant across professional and personal coding work, saying Gemin…
A study by US AI start-up Emergence found that autonomous AI agents powered by Claude, Gemini, Grok, OpenAI, Qwen, DeepSeek and Mistral spontaneously developed shorthand and new word meanings, with up to half of agent-to…
OpenAI launched a new framework for tracking and publicly disclosing "model misalignment" on September 16, publishing six incident reports covering behaviors observed between October 2025 and July 2026. The reports detai…
The Georgian AI Lab outlined a resource-efficiency framework for agentic development that treats model selection as an ongoing cycle of capability investment and cost optimization, arguing open-weight models should absor…
Anthropic signed its first Australian data center lease for the proposed Western Downs Digital Park near Dalby, Queensland, a campus planned to reach 2.16 gigawatts of capacity that could begin coming online in 2027, acc…
Cryptography experts pushed back on claims that advanced AI models can defeat encryption, after Anthropic announced in July that its Mythos model "weakened" the HAWK post-quantum cryptography algorithm and HAWK's team pu…
Moonshot AI launched its Kimi Financial Industry AI Solution on September 17, connecting its large language model to more than ten major financial data providers including Wind, East Money, S&P Global, Caixin Data, and C…