From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix Enterprises facing data-residency constraints must self-host LLMs, but adopting newer models without retiring older ones fragments GPU resources. A technical report describes consolidating traffic from over 200 internal applications onto a single self-hosted model by closing quality gaps through post-training, covering the corporate request mix. Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decommissioning their predecessors expands the serving fleet, fragmenting a finite GPU pool. We consolidate traffic from over 200 internal applications onto a single model by closing quali