OpenAI slashes inference costs by over 50% with Nvidia GPU efficiency: The Information
OpenAI has reduced inference costs by over 50% for some existing models, operating logged-out ChatGPT traffic on just a few hundred Nvidia GPUs, according to The Information. The efficiency gains, ach…