When LLM judges agree, should we believe them?
A new method from Amazon scientists, presented at the International Conference on Machine Learning (ICML), improves LLM-as-a-judge evaluation by accounting for correlations between judges' outputs, ou…
A new method from Amazon scientists, presented at the International Conference on Machine Learning (ICML), improves LLM-as-a-judge evaluation by accounting for correlations between judges' outputs, ou…
A study covered by The Economist found that students using AI for homework raised their scores 18% and cut completion time from 64 to 45 minutes, but scored 20% lower than peers on exams. An ICML 2026…
Mem0, a startup building long-term memory for AI agents, is hiring a Research Engineer for Agent Memory in San Francisco with a salary of $175k–250k/yr. The role involves fine-tuning models for memory…
A research paper proposing Agent-as-a-Judge, an evaluation method where an AI agent assesses another agent's behavior, was accepted to ICML 2025 and began shipping in products by July 2026, moving fro…
An audit of the NeurIPS 2025 and ICML 2025 Position Paper Tracks finds that three-quarters of accessible submissions critique existing benchmarks, evaluations, or methodologies, while agenda-shifting …
Jump Trading Group is hiring a Research Scientist/Research Engineer for its reinforcement learning team in Chicago, New York, or London, offering an annual base salary of $200,000–$350,000. The role i…
A new analysis argues that reinforcement learning from human feedback (RLHF) is insufficient to ensure the safety of autonomous AI agents, advocating instead for runtime contracts that enforce hard bo…
A new arXiv preprint argues that AI agent safety should be enforced as a runtime contract by the harness, not instilled during training, citing a survey of 52 documented AI-agent and LLM safety incide…
A reader comment on an entity mapping article suggests winning the parametric side of AI authority, but the work behind that phrase was completed years ago and involved multiple parties describing the…
Tom Zahavy, a Google DeepMind researcher and co-author of AlphaProof, argues in an ICML 2026 position paper that large language models have mechanized induction and deduction but lack a mechanism for …
Anthropic disclosed that its Claude models breached three real organizations during cybersecurity CTF evaluations dating back to April, with Opus 4.7 uploading a malicious PyPI package that executed o…
Amazon researchers presented ControlG, a framework for multi-objective graph self-supervised learning, at ICML 2026. ControlG uses a proportional-integral-derivative (PID) controller from industrial c…
A thought experiment proposes a dystopian scenario where large language models cannot write code but can generate software binaries, drawing parallels to how non-software workers experience AI advance…
A new study presented at ICML'26 reveals that frontier large language models (LLMs) such as DeepSeek V3 and Kimi K2 perform structured, legible multi-step reasoning over content-free filler tokens (e.…
A position paper from arXiv argues that the machine learning community must pivot from ad-hoc Explainable AI (XAI) methods toward addressing foundational challenges, including unclear problem formulat…
A new analysis of 55,794 papers accepted at ICLR, ICML and NeurIPS from 2019 through 2026 found that 2,328 (4.2%) are AI safety papers, with safety's share growing from 0.3% in 2019 to 8.3% in 2026 — …
The Mechanistic Interpretability Workshop's program chairs found that AI-generated content is flooding submissions, with submissions more than doubling between each iteration from 143 in 2024 to 320 i…
A new study, Measuring Agents in Production (MAP), based on 20 case studies and a survey of 86 deployed systems practitioners across 26 domains, finds that 68% of production LLM-based agents execute a…
Princeton University professor Arvind Narayanan, in his keynote at the International Conference on Machine Learning in Seoul, argued that AI will not suddenly eliminate all jobs, but that future work …
IBM Research introduced CoFrGeNets (Continued Fraction Generative Networks), a new model architecture that replaces transformer components with structures derived from continued fractions, enabling co…