Autonomous LLM post-training with Tunix on TPUs
The autofinetune project applies autonomous research loops to LLM post-training, using Google's Tunix library, Gemma models, and Cloud TPUs orchestrated with Antigravity CLI and Gemini Flash 3.7. In a…
The autofinetune project applies autonomous research loops to LLM post-training, using Google's Tunix library, Gemma models, and Cloud TPUs orchestrated with Antigravity CLI and Gemini Flash 3.7. In a…
Google's open-source, JAX-native post-training library Tunix now ships an agentic RL trainer built around asynchronous rollouts and a decoupled producer-consumer pipeline, aiming to eliminate idle TPU…
Google's Tunix post-training library introduces a high-throughput framework for agentic reinforcement learning that keeps TPUs fully utilized during multi-turn training. Tunix uses an asynchronous tra…
MarkTechPost published a tutorial on training Gemma-3 for structured mathematical reasoning using Tunix GRPO, LoRA adapters, and GSM8K rewards. The workflow includes environment setup, prompt formatti…
Google hosted a Kaggle hackathon challenging developers to train non-reasoning Gemma-2-2B and Gemma-3-1B models into general reasoning models using Tunix and Kaggle TPUs. Over 11,000 entrants and 300+…
MaxText has introduced new post-training capabilities, specifically Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), now available on single-host TPU configurations like v5p-8 and v6e-8. …