Kavram: testing one async image API over distributed GPU capacity Kavram, a new image inference API built by the author, is testing distributed GPU capacity with a price of $0.0012 per completed image for FLUX.1 Schnell at 1 MP, compared to fal.ai's listed $0.003/MP. The project seeks technical feedback on what minimal reproducible artifact would be expected before testing a new image inference API. Disclosure: I am one of the builders. We built Kavram around a practical infrastructure question: can an image product keep a normal provider-style async API without making the product team operate GPU capacity itself? Concrete artifact: The current price reference we are testing is FLUX.1 Schnell at 1 MP: $0.0012 per completed image on Kavram versus fal.ai’s listed $0.003/MP. We are not treating that as a universal benchmark; output acceptance, settings, latency, retries, and failure handling must be compared on the same workload. I would value technical feedback on one question: What minimal reproducible artifact would you expect before testing a new image inference API—a public request harness, failure traces, latency distributions, or an interactive Space? Playground/docs: Early-access workload form: If you already run a relevant model, share the model, resolution, and one non-sensitive prompt. We will publish the comparison conditions with any result.