Limits of Confidence in Diffusion
Apple researchers Russ Webb, Amitis Shidani, Alice Bizeul and Dan Busbridge published an October 2026 paper showing that discrete diffusion samplers — including remasking and uniform-state samplers — …
Apple researchers Russ Webb, Amitis Shidani, Alice Bizeul and Dan Busbridge published an October 2026 paper showing that discrete diffusion samplers — including remasking and uniform-state samplers — …
A paper published in October 2026 by Kirill Brilliantov, Alejandro Hernández-Cano and Emmanuel Abbé finds that open-source state-of-the-art MLE agent harnesses provide no advantage over a single sessi…
A new method called RLTL;DR lets a Qwen 3.5 9B Thinking policy break through a learning barrier on tool-calling and coding datasets filtered to Pass@128 = 0, reaching a Pass@1 of 14–31% with insights …
A systematic study by Iuri Macocco, Pau Rodríguez Lopez, Arno Blaas, Luca Zappella, Marco Baroni and Xavier Suau Cuadros, published in Transactions on Machine Learning Research (TMLR) in September 202…
A round-trip study of sixteen language models found that serializing tree-structured arithmetic expressions into natural language is a lossy, asymmetric channel, with swapping the generating and extra…
Apple detailed its production use of homomorphic encryption (HE) combined with private information retrieval (PIR) and private nearest neighbor search (PNNS) to power Enhanced Visual Search in Photos,…
Apple researchers compressed an on-device streaming neural audio tokenizer by 2.8× using latent-space distillation, keeping the distilled student within 1.9% relative word error rate of its teacher on…
Researchers Alex Ferrando de las Morenas, Xavier Suau Cuadros, Jordi Gonzàlez Sabaté, and Pau Rodríguez Lopez introduced Dynamically Scaled Activation Steering (DSAS), a method-agnostic framework that…
Researchers Kostia Kudriavtsev, Parvez Rafi, and Sha Sundaram presented Glyph, a production multi-agent LLM system that generates column descriptions and assigns sensitivity-ontology tags for enterpri…
Researchers affiliated with The Ohio State University and Apple proposed DACA-GRPO, a plug-and-play enhancement to GRPO-style trainers for diffusion language models that adds Denoising Progress Scores…
Apple researchers Arnav Arora, Natalie Schluter, Katherine Metcalf and Maartje ter Hoeve found that fine-tuning conversational large language models on curated value subsets of existing preference dat…
A September 2026 paper by Sanjana Pedada, Aditya Dhavala, and Neelraj Patil introduces shared selective persistent memory, a memory architecture for agentic LLM systems that retains task specification…
Apple unveiled its third generation of Apple Foundation Models (AFM), a family of five models built in collaboration with Google, spanning two on-device models and three server-based models running on…
Researchers Vasileios Baltatzis, Mert Inan, Connor Gillis, Raja Kushalnagar, Lorna Quandt, Leah Findlater, and Colin Lea introduced DiscoSign, a modular Large Language Model-based framework for discou…
Researchers Riyaaz Shaik and Chandru Venkataraman introduced REFACTOR-VLA, a system that learns reusable motor skills for vision-language-action models using a wake/sleep architecture with a Behaviora…
Researchers from Stanford University and Apple found that large language models (LLMs) are not consistently Bayesian when updating probabilistic beliefs from evidence, with non-Bayesian heuristic upda…
Researchers propose Luce, a 3D representation that unifies geometry and PBR materials in a voxelized multimodal Gaussian cloud, enabling relightable image-to-3D generation. On Toys4K, Luce improves FI…
Researchers from Amazon and Georgia Tech proposed IDEA Prune, an integrated enlarge-and-prune pipeline for generative language model pretraining that combines enlarged model training, pruning, and rec…
Researchers from UIUC and Apple introduced STARFlow2, a unified multimodal model built on the Pretzel architecture that interleaves a frozen pretrained vision-language model with a TARFlow stream via …
Researchers from Microsoft and academic institutions introduced Internalized Visual Thinking (IVT), a post-training framework that enables multimodal large language models to reason about videos witho…