This was not supposed to become a product.
When PaddleOCR-VL-1.6 dropped, independent benchmarks put it at the top of document parsing models. I had to try it. I needed a provider, but there simply isn't one ready for production that I would trust.
So i set one up myself. I assumed that even after getting it running, serving a vision-language model would be expensive.
It turns out the opposite is true. Once I had it running properly, the cost was absurdly low. At proper GPU utilization, the cost is only around $1 per 1,000 pages.
The nearest competitors are either much lower quality (Azure Read) or absurdly expensive (Extend or Reducto). Even Mistral OCR 4 which is really good and pretty cheap is still 4x more expensive.
So I had to share it. I made the endpoitns public, and vibe-coded a simple dashboard.
Let me know what you think!
Comments URL: [https://news.ycombinator.com/item?id=49041317](https://news.ycombinator.com/item?id=49041317)
Points: 2