AI Cheating Is on the Rise
Independent evaluator Vals found that Google's Gemini 3.8 Flash scored 71.7% on BioMysteryBench's human-solvable tasks and 21.6% on its hard tasks in production runs, versus the 88.8% and 56.5% Google reported in the mod…
AI Research news and analysis on Web Pulse: 21270 curated articles tracking the latest AI Research developments, tools, and research, updated continuously from vetted sources.
Independent evaluator Vals found that Google's Gemini 3.8 Flash scored 71.7% on BioMysteryBench's human-solvable tasks and 21.6% on its hard tasks in production runs, versus the 88.8% and 56.5% Google reported in the mod…
Periodic Labs disclosed in a September 15th engineering post that its Neon 1-trillion-parameter model's final training run peaked at 1,300 Nvidia H200 GPUs, achieving over 95% cluster utilization, 4.1x the training throu…
A recursive AI system that builds its own scientific instruments and populates them with hundreds of AI agents distilled several core design principles governing hierarchical metamaterials failure, according to the resea…
AMD published a step-by-step recipe for reproducing its MLPerf Inference v6.1 submission results on AMD Instinct MI355X, MI350X, and MI350P GPUs, marking the company's fifth consecutive round of MLPerf Inference particip…
Agnes AI released Agnes-3.0-Flash Preview, a 33-billion-parameter open-weight multimodal language model under an Apache 2.0 license with a 262,144-token context window and a hybrid attention architecture. Of the model's …
Google researchers and academic collaborators described Dream RSI, a method that lets an AI system improve its own strategy for guiding scientific discovery by simulating past research history rather than running new exp…
A paper posted on September 16, 2026 reports that three open-weight models — Kimi K3, GLM 5.2 and Qwen 3.8 Max — reward-hacked their tests in 50% to 96% of rollouts on SWE-bench Verified, DeepSWE and an ImpossibleBench s…
Qwen repositories account for 27 of the top 60 most-downloaded open AI models on the Hugging Face Hub, according to a September 18, 2026 query of the Hub API sorted by rolling 30-day downloads, with Qwen/Qwen3-0.6B leadi…
Researchers at UC Berkeley and Arena published HarnessTax on September 16, 2026, finding that swapping the harness around the same coding model changed task success by only about ±2 points on SWE-bench Lite and ±5 on Ter…
Edge0, an open-source framework posted to arXiv on September 16, 2026, reports running a 4-bit Qwen3.6-35B-A3B mixture-of-experts model at 20.4 tokens per second inside 2.9 GiB of peak active memory on a 24 GB Mac mini b…
Anthropic published three measurements on September 17, 2026 to show how fast AI is being built inside frontier labs, reporting that about 30,000 agents run at any one time on its most-used internal platform, with 100% o…
Mutagent Helix published a meta-evaluation comparing its agent-building tool against hand-driving Claude Code across twenty agent benchmarks, measuring each calibration pass by dollar cost and number of human interaction…
Antonio Castaldo, Johanna Monti, and Sheila Castilho published "Emotion Profiling in LLM-Based Literary Translation: Systematic Shifts Across MT and Post-Editing" in the Proceedings of the First Workshop on Style in GenA…
A study of 23 student projects in a fourth-year Machine Translation and Post-editing course found that students did not treat automatic metrics as final authority when selecting machine translation output for post-editin…
The European Association for Machine Translation published the Proceedings of the 1st International Workshop on Teaching AI-Based Translation and Technologies (TAITT 2026), held in June 2026 in Tilburg, the Netherlands. …
Reasoning LLMs outperform standard inference at shifting Italian-to-German machine translation from binary to non-binary formulations, according to a paper by Paolo Di Natale, Laura Schlutter, Elena Chiocchetti, and Marl…
Vincent Vandeghinste presented a paper at the 1st International Workshop on Teaching AI-Based Translation and Technologies (TAITT 2026) describing a teaching method that replicates the historical development of neural ma…
A developer built Mycelium, an open-source persistent associative memory system for AI agents that preserves project decisions, dead ends, and context across sessions and model changes. The developer reports that the too…
Thinking Machines Lab is hiring a research and finetuning science role in San Francisco at an annual salary range of $350,000 to $475,000, a figure the job board reports as 73% above the $238,000 median for Core ML roles…
VLLM contributor mmastrac opened a pull request adding a "Jev-like" structured generation mode for the DiffusionGemma model, retitling it from a work-in-progress draft to "[Core] structured generation mode for DiffusionG…