cd /news/artificial-intelligence/onetriangle-launches-deepseek-v4-fla… · home topics artificial-intelligence article
[ARTICLE · art-109341] src=runtimewire.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

OneTriangle launches DeepSeek V4 Flash hosting at $0.15 per million input tokens

OneTriangle, a Y Combinator Summer 2026 company, launched hosted DeepSeek V4 Flash inference on August 24 at $0.15 per million input tokens and $0.35 per million output tokens, with CEO Hannah Chung claiming it is the "fastest lightweight, cheap inference," though the claim lacks independent verification. The pricing is about 7% higher for input and 25% higher for output than DeepSeek's direct API rates of $0.14 and $0.28 per million tokens, positioning OneTriangle's offering on serving performance and managed infrastructure rather than lowest price.

read4 min views2 publishedAug 24, 2026
OneTriangle launches DeepSeek V4 Flash hosting at $0.15 per million input tokens
Image: Runtimewire (auto-discovered)

The YC S26 company says its eight-H100 stack is tuned for low-latency inference, though its rates exceed DeepSeek's direct API.

By Ryan Merket · Published

Primary source: X

Why it matters #

OneTriangle is testing whether inference startups can charge near commodity model rates while differentiating on serving speed and long-context economics.

Hannah Chung (@hannah_chuu) launched OneTriangle's hosted DeepSeek V4 Flash service on August 24th, pricing it at $0.15 per million input tokens and $0.35 per million output tokens. Chung described the offering as the "fastest lightweight, cheap inference" in an eight-post thread on X, a performance claim that has not been established by an independent comparison. (onetriangle.ai)

https://x.com/hannah_chuu/status/2091949930919465012 Chung, OneTriangle's CEO, studies computer science and business at MIT. CTO Medha Venkatapathy (@medha_rv) studies physics and computer science there and has researched large-model post-training under MIT professor Jacob Andreas. OneTriangle's team page describes a five-person group whose collective experience spans Google DeepMind, Jane Street, SpaceX, MIT Lincoln Laboratory and MIT's Computer Science and Artificial Intelligence Laboratory. OneTriangle is part of Y Combinator's Summer 2026 batch, according to Chung's announcement. (onetriangle.ai)

The founders are entering a crowded market around an unusually deployable frontier model. DeepSeek released the V4 preview in April, then published the official DeepSeek V4 Flash 0731 weights in late July. The official model repository lists an MIT license, a one-million-token context window and 304 billion total parameters. DeepSeek says only a fraction of those parameters are active for each token because the model uses a mixture-of-experts architecture. (api-docs.deepseek.com)

The price is not the floor

OneTriangle's immediate-processing rate sits slightly above DeepSeek's direct API pricing. DeepSeek currently charges $0.14 per million uncached input tokens and $0.28 per million output tokens. OneTriangle's listed prices are about 7% higher for input and 25% higher for output. That makes OneTriangle's pitch a serving-performance and managed-infrastructure argument rather than a claim to the lowest sticker price. (api-docs.deepseek.com)

OneTriangle also offers a delayed tier at 40% below its base rate, with delivery within 10 hours. That reduces the effective price to $0.09 per million input tokens and $0.21 per million output tokens, a potentially useful option for evaluations, data processing and other workloads that do not require an immediate response. The company says customers pay only for tokens used, without a subscription. (onetriangle.ai)

OneTriangle wants to move the model's memory

The hosted DeepSeek release gives Chung and Venkatapathy a product to sell while they develop a more ambitious inference technique: transferring the key-value cache generated while one model reads a prompt into another model that produces the answer.

Ordinary inference makes the answering model process the entire prompt before generating its first token. OneTriangle wants a smaller model to perform that prefill work, translate the resulting cache into a form the larger model can use, and let the larger model begin decoding without rereading the full context. The economic benefit should increase as prompts and agent histories grow longer, since prefill consumes more compute before the user sees an answer. (onetriangle.ai)

OneTriangle's strongest published result so far covers a Minitron 4B-to-Llama 3.1 8B handoff rather than DeepSeek V4 Flash. In an August 18th research note, OneTriangle reported that an 8,192-token prompt reached the Llama model's first token in 38.25 milliseconds when the smaller model's cache was already resident, compared with 302.69 milliseconds for native prefill. Including the smaller model's initial read reduced the speedup from 7.91 times to 1.25 times. The transferred path also reached 82.52% top-token agreement with native Llama behavior over 32 generated steps, a bounded continuation test rather than proof of equivalent answer quality. (onetriangle.ai)

The DeepSeek service currently provides a more conventional demonstration of OneTriangle's systems work. In a separate August 24th production note, OneTriangle said it tested DeepSeek V4 Flash on eight Nvidia H100 GPUs. Its retained configuration processed about 10,039 output tokens per second across 128 concurrent decode-heavy requests, with median time to first token of roughly 0.35 seconds. A configuration that was faster for a single request collapsed under the same load, pushing median time to first token above 22 seconds. These figures compare OneTriangle's own server configurations, so they do not substantiate the broader "fastest" label used in Chung's launch thread. (onetriangle.ai)

That distinction defines OneTriangle's immediate challenge. Open model weights have created a market where many providers can sell the same underlying intelligence. Chung and Venkatapathy are betting that model-specific systems work, long-context optimization and eventually transferable caches can turn commodity access into a differentiated inference product. The DeepSeek launch puts that bet in front of customers before the founders' cache-transfer research has matured into the service's main selling point.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @onetriangle 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/onetriangle-launches…] indexed:0 read:4min 2026-08-24 ·