Teaching AI to Read a Map
Google researchers introduced MapTrace, a synthetic data generation pipeline and dataset of 2 million question-answer pairs, to teach multimodal large language models (MLLMs) to trace routes on maps, …
Google researchers introduced MapTrace, a synthetic data generation pipeline and dataset of 2 million question-answer pairs, to teach multimodal large language models (MLLMs) to trace routes on maps, …
A developer has built a browser-based 3D circuit simulator using Three.js, React 19, TypeScript, and Vite, with an Express backend on Render and Firebase for authentication and data storage. The tool …
RayNeo launched the RayNeo iO Glasses, an AI wearable with no camera and no speaker, featuring a green waveguide display with 97 percent transparency and about 1,300 nits of brightness. The glasses li…
Metric founder and CEO Hrant Davtyan released ArmBench-ASR on August 20, a benchmark evaluating nearly 30 speech recognition systems on 20.7 hours of Armenian audio across five datasets, with Google's…
Google's Gemini CLI free tier now serves Flash models only as of 25 March 2026, with a daily limit of 1,000 model requests per user on a personal Google account and 250 for unpaid API keys, according …
Researchers at MIT have developed an LLM-based system that learns recurring patterns from weeks of a person's real conversations to predict their likely next conversational behavior. In an experiment,…
Researchers at the Max Planck Institute for Informatics have developed a technique to extract hidden reasoning traces from proprietary AI models like Claude, GPT, and Gemini via API. The method, which…
Researchers introduced SKILL, a self-correcting knowledge-guided iterative large language model agent that unifies multi-agent LLM reasoning with reinforcement learning for logic synthesis optimizatio…
A new arXiv preprint (2608.13889v1) presents a consensus-gated multi-agent neural architecture search (NAS) system that uses three large language models (Claude, GPT-5.1, and Gemini 2.5 Pro) to debate…
GLM-5.2 scores 99.2% on AIME 2026 with 40 billion active parameters, while Qwen3.5's 9B model hallucinates on 82% of knowledge questions in Artificial Analysis's Omniscience benchmark, reflecting a de…
Reasoning models are deliberately trading world knowledge for reasoning skill, with GLM-5.2 scoring 99.2% on AIME 2026 using about 40 billion active parameters per token, while factual recall remains …
Researchers introduced RefineBench, a benchmark of 1,000 challenging problems across 11 domains, to evaluate language models' ability to refine their own responses. In self-refinement, even frontier m…
A Carnegie Mellon, MIT, and Cornell research team posted two randomized case studies on August 6 finding that short AI conversations reduced belief in new conspiracies by 6.95 and 7.56 points on a 0-t…
Overlap Research, with support from BlueDot Impact, found that supervised fine-tuning with self-other overlap (SOO SFT) reduced deception in large language models from 96-100% to 30.24% (Qwen2.5-14B-I…
A preprint posted on August 6 found that six frontier language models—Claude Opus 4.7, GPT-5, Gemini 2.5 Pro, DeepSeek-R1, Qwen3.7-Max, and Llama-3.3-70B-Instruct-Turbo—split under steering pressure, …
Google released a stack of AI updates in July 2026, headlined by Gemini 2.5 Pro with expanded context windows and improved chain-of-thought reasoning, alongside Gemini 2.5 Flash for lower latency and …
A new arXiv preprint (2608.02665v1) finds that evaluating large language models on a single canonical prompt surface underestimates unsafe compliance by 3.3 to 12.9 percentage points across five model…
Roboflow published a tutorial on detecting small objects in drone imagery, training RF-DETR on roughly 7,000 aerial images with over 120,000 person annotations to identify nine classes of people and v…
NVIDIA released Alpamayo 2 Super, a 34-billion-parameter reasoning model for autonomous driving, for commercial use on August 4, 2026, under the Linux Foundation's OpenMDW-1.1 license, clearing a lice…
Roboflow published a tutorial demonstrating human-object interaction (HOI) detection without training an interaction model, using RF-DETR to detect workers, forklifts, pallets, and carts (77.5% mAP@50…