# Parallel Constrained Decoding

> Source: <https://twitter.com/harshagundal/status/2100044305536889015>
> Published: 2026-09-16 15:37:08+00:00

They were building in stealth for 2 years, I was building in stealth for 2 hours…
Happy to open source Qwen-2.5-1B-RLCD, 5x faster on-device inference for JSON workloads that need to be type-safe. 
⚡️Demo below on a M4 MacBook⚡️
every LLM has the ability to efficiently batch inference every key of a JSON at the same time and generate probabilities from a set of possible categories. No new training required, but it’s easy to optimize if you need!
On hugging face now!

00:00

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x 

00:00
