Ask HN: What are you using for LLM inference in production? A developer on Hacker News asks what tools others use for LLM inference in production, noting they settled on Gemini 2.5 Flash Lite after starting with OpenAI's GPT-3 and trying other major labs. They express interest in open-source options but have not found comparable cost, performance, and speed, and mention Groq's developer access has been unavailable for months. | |||||||||||| 2 points by | I started with OpenAI back in the GPT-3 days, then bounced between the major labs with a brief interlude with Workers AI . I eventually settled on Gemini 2.5 Flash Lite as my workhorse, using it against structured data/vectors as a chatbot. It was cheap, good enough, and most importantly, Groq seems promising but developer access hasn't been available for months. I'd love to use open source, but I never found anything comparable cost/performance/speed . Would love to hear your suggestions. | ||||||||||| Applications are open till July 27. |