Autonomous LLM post-training with Tunix on TPUs
17 September 2026 at 20:01
The "autofinetune" project introduces an autonomous research loop that fully automates LLM post-training workflows, including Supervised Fine-Tuning (SFT) and Reinforcement Learning via GRPO. By defining boundary conditions and evaluation metrics in a single Markdown specification, developers can deploy an AI agent to iteratively edit training scripts, launch experiments, and automatically commit verified hyperparameter optimizations to Git. Built on Googleβs AI stackβincluding Tunix, Gemma, and Cloud TPUsβthis framework eliminates manual tuning cycles, successfully demonstrating hands-off performance gains in both function calling and math reasoning models.