{"slug": "show-hn-500-tb-internet-index-in-clickhouse-with-congestion-pricing", "title": "Show HN: 500 TB internet index in ClickHouse, with congestion pricing", "summary": "A developer launched Scry, a search service built on a 500 TB NVMe internet index stored in ClickHouse that lets users run arbitrary read-only SQL and some Datalog queries, with congestion-based micro-auction pricing to manage resource contention. The service is free for non-commercial use when capacity is available, and the developer said he intends to scale the paradigm on differentiated hardware over more data. The developer argued that Google Search, Tavily, and Exa map agent context to a small subset of their indexes at a fixed cost, leaving users to absorb the downsides when budgeted compute runs out, and that search companies abandoned text-to-SQL after the \"traumatic 2024 text-to-sql days.", "body_md": "Hello. It's 2026, we're training simulated fruit fly brains to play Beat Saber, do we still have to be stuck with internet (re)search as fn: natural language -> black box we can't do anything about -> ranked_list/summary?\n\nThere is a long history of people trying to do very fancy things that end up being done in relational databases and a little SQL. There is a gravity to them, a bitter lesson, just like scaling of generalized ml training methods. I mean many, many information products can be built off essentially giant real-time OLAP databases and frontier LLMs writing brilliant SQL+Datalog+vector+Jev etc. queries.\n\nGoogle Search, Tavily, Exa essentially have the problem of *mapping* your agents' context you are willing to provide, to a tiny subset of their index. You pay a fixed cost to an extremely hard problem that has a distribution of hardness, which means YOU eat the downsides when they are running out of budgeted compute to help you out.\n\nTheir algorithms are opaque to the caller, there's really not much user control, and there's not a serious opportunity to communally improve search recipes, like the lexical+Jev recipes you trust to select bleeding edge AI builders.\n\nFurthermore, search companies aren't even pursuing text-to-SQL anymore (several have talked to me)... they made up their minds during the traumatic 2024 text-to-sql days. They were just too early.\n\nMeet Scry, where I have a 500 TB NVMe internet index (I'm doing my best indexing and normalizing all the intelligence explosion alpha) that you can run ~arbitrary readonly SQL and some of Datalog over, and I handle the problem of resource-contention with congestion-based micro-auction pricing. When there's capacity, the service is free for non-commercial use.\n\nI hope you enjoy. I'm intent on scaling this paradigm on differentiated hardware over much more data, so any compelling use cases or queries I could show off, would be much appreciated!\n\nComments URL: [https://news.ycombinator.com/item?id=49748041](https://news.ycombinator.com/item?id=49748041)\n\nPoints: 1\n\n# Comments: 0", "url": "https://wpnews.pro/news/show-hn-500-tb-internet-index-in-clickhouse-with-congestion-pricing", "canonical_source": "https://scry.io/", "published_at": "2026-09-17 23:15:57+00:00", "updated_at": "2026-09-17 23:25:06.905649+00:00", "lang": "en", "topics": ["ai-search", "structured-data", "ai-agents", "ai-infrastructure"], "entities": ["Scry", "ClickHouse", "Google Search", "Tavily", "Exa"], "alternates": {"html": "https://wpnews.pro/news/show-hn-500-tb-internet-index-in-clickhouse-with-congestion-pricing", "markdown": "https://wpnews.pro/news/show-hn-500-tb-internet-index-in-clickhouse-with-congestion-pricing.md", "text": "https://wpnews.pro/news/show-hn-500-tb-internet-index-in-clickhouse-with-congestion-pricing.txt", "jsonld": "https://wpnews.pro/news/show-hn-500-tb-internet-index-in-clickhouse-with-congestion-pricing.jsonld"}}