{"slug": "halo-post-train-llms-3x-faster-than-trl-and-megatron", "title": "Halo: Post-train LLMs 3x faster than TRL and Megatron", "summary": "White Circle launched Halo, a post-training framework for open-source models that delivers up to 2.8x the throughput of stock TRL with lower peak memory while keeping models in native HuggingFace format. On OpenAI's gpt-oss-20b, Halo reached 2.3–2.8x TRL throughput under identical FlashAttention, Liger, fused cross-entropy and grouped-GEMM optimizations, and fine-tuning Z.ai's GLM-4.7-Flash on 177M tokens of agentic traces raised SWE-rebench-V2 from 33% to 42% at up to 1.63x TRL throughput. White Circle said every model it trains now uses Halo, which runs LoRA on a 24 GB GPU, multi-node training on B300s, and async RL, and it is partnering with the SGL project to make Halo the primary engine for rollouts.", "body_md": "White Circle on X: \"Introducing Halo, the best framework for post-training of open-source models.\nHalo delivers up to 2.8x the throughput of stock TRL with less peak memory, while models stay in their native HuggingFace format.\nStar us on GitHub: https://t.co/3mAiUljdrN\" / X\n\nWhite Circle on X: \"Introducing Halo, the best framework for post-training of open-source models.\nHalo delivers up to 2.8x the throughput of stock TRL with less peak memory, while models stay in their native HuggingFace format.\nStar us on GitHub: https://t.co/3mAiUljdrN\"\n\nIntroducing Halo, the best framework for post-training of open-source models.\nHalo delivers up to 2.8x the throughput of stock TRL with less peak memory, while models stay in their native HuggingFace format.\nStar us on GitHub: github.com/whitecircle/ha…\n\nIntroducing Halo, the best framework for post-training of open-source models.\nHalo delivers up to 2.8x the throughput of stock TRL with less peak memory, while models stay in their native HuggingFace format.\nStar us on GitHub: github.com/whitecircle/ha…\n\nEvery model we train at White Circle now uses Halo.\nThe same codebase runs LoRA on a 24 GB GPU, multi-node training on B300s, and async RL.\nTechnical blog:\n\nOther training frameworks often require a separate implementation for each model family.\nIn Halo, adding a new model family takes just 100 lines of connecting a wrapper instead of rewriting the model.\n\nHalo is not limited to supervised fine-tuning. It supports async RL and training with external environments.\nWe are excited to partner with the @sgl_project team to make it the primary engine for rollouts.\n\nOn @OpenAI gpt-oss-20b, Halo delivered 2.3–2.8x the throughput of stock TRL with less peak memory.\nBoth runs used the same FlashAttention, Liger, fused cross-entropy and grouped-GEMM optimizations.\n\nWe also used Halo to fine-tune @Zai_org GLM-4.7-Flash on 177M tokens of agentic traces.\nThe resulting model improved SWE-rebench-V2 by @nebiusai from 33% to 42%. Halo reached up to 1.63× TRL throughput on the same model precision and data.\nModel and write-up:Show more\n\nScaling an AI model shouldn't require rewriting half your codebase\nBut that's exactly what happens when you outgrow your training framework\n@whitecircle just launched Halo:\n- Up to 2.8x faster than TRL\n-- Distributed training\n--- Native Hugging Face models\n---- One stack fromShow more", "url": "https://wpnews.pro/news/halo-post-train-llms-3x-faster-than-trl-and-megatron", "canonical_source": "https://twitter.com/whitecircle/status/2102087563913609534", "published_at": "2026-09-21 17:59:52+00:00", "updated_at": "2026-09-21 18:24:05.217637+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-infrastructure", "mlops", "ai-agents"], "entities": ["White Circle", "Halo", "TRL", "Megatron", "OpenAI", "gpt-oss-20b", "Z.ai", "GLM-4.7-Flash"], "alternates": {"html": "https://wpnews.pro/news/halo-post-train-llms-3x-faster-than-trl-and-megatron", "markdown": "https://wpnews.pro/news/halo-post-train-llms-3x-faster-than-trl-and-megatron.md", "text": "https://wpnews.pro/news/halo-post-train-llms-3x-faster-than-trl-and-megatron.txt", "jsonld": "https://wpnews.pro/news/halo-post-train-llms-3x-faster-than-trl-and-megatron.jsonld"}}