I benchmarked 8 LLMs for a niche production app. The flagship cost 5.8x more - and lost.
A developer benchmarked eight LLMs for a niche BaZi birth-chart app and found that the flagship model cost 5.8x more yet lost on accuracy. The developer built a routing layer with ordered fallback chains and error classi…