cd /news/artificial-intelligence/reka-rho-1-collapsing-the-multimodal… · home › topics › artificial-intelligence › article
[ARTICLE · art-145527] src=reka.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Reka Rho-1: Collapsing the multimodal stack

Reka released a research preview of Rho-1, a 19B omni-reasoning model trained from scratch that unifies text, images, video, and robotic actions as tokens in a single context window. The base Rho-1 model generates video at 0.79× real-time (median) with a watchable stream starting in roughly 6 seconds, and a distilled variant cuts the denoising trajectory from 99 steps to 8, returning a 5.3-second video clip in about a second. Reka says collapsing the modality-specific pipeline stack into one network enables end-to-end optimization and is a prerequisite for physical AGI.

read10 min views2 publishedOct 5, 2026
Reka Rho-1: Collapsing the multimodal stack
Image: source

An omni-reasoning model that understands, simulates, and acts.

Today, we are releasing a research preview of Rho-1, our 19B omni-reasoning model trained from scratch. Within a single neural network, it understands and generates text, images, and video, reasons over them, and takes actions.

Most AI systems today are agentic pipelines: a central model plans and delegates, handing jobs off to modality-specific specialists. Each handoff adds latency, and each specialist sees only a narrow request, not the full context. Rho-1 collapses that stack. Text, vision, and robotic actions are unified as tokens within a single context window.

We believe this architecture points toward the future of intelligence: not a patchwork of narrow models wired together, but a single foundation that handles the entire loop. Rho-1 can imagine an environment, track its state, answer questions about it, and set it in motion without ever dropping the thread. With no seams between components, the whole system can be optimized end-to-end, making it a stronger foundation for any application that spans modalities. We see this kind of unification as the core prerequisite for physical AGI.

What Rho-1 produces is not a one-off output but an evolving world state. Steer it with dialogue, and you get fluid conversation. Steer it with real-time controls, and you get a persistent, interactive simulation. Steer it with motor policies, and you get embodied robotic control. This post explores each of these three frontiers in turn.

Rho-1 as a multimodal assistant #

Figure 1 shows a single, unedited session with Rho-1. Across five turns, it draws a static scene, locates an object within it, animates the scene into a video, edits the video to change the weather, and explains what changed, reasoning through each step along the way. Work that typically spans several specialized systems is handled within a single model.

Watch the state in the model view grow with each turn. A request to generate a picture routes to the diffusion tower, while a request for a bounding box stays on the understanding side. Either way, every turn reads from and writes to the same KV cache, which acts as a persistent world state:

  • Reasoning: The block above each reply shows the model’s chain of thought. There is no bolt-on prompt enhancer: the same model plans the shot and renders it.
  • Grounding: Rho-1 generated the lighthouse, so the scene is already in its context. The bounding box is emitted as coordinate tokens attending to that state, with no secondary detection model.
  • Animating: The first video frame isn’t a re-encoded image. It’s the original representation held in context, so lighting, geometry, and object identity carry through.
  • Understanding: When asked what changed between clips, Rho-1 isn’t running post-hoc captioning on exported frames. It reads the latent state that produced the change, so its answer is grounded in what it actually generated.

Rho-1 is a fast model. The base Rho-1 model generates video at 0.79× real-time (median), with a watchable stream starting in roughly 6 seconds.

A distilled variant cuts the denoising trajectory from 99 steps to 8 with minimal quality loss, pushing the system into interactive territory. In our internal testing, it is among the fastest models in the world in every modality we measured. It returned a 5.3-second video clip in about a second, faster than any other video model we timed. It made images as quickly as the fastest dedicated image models. And it was the quickest of the models we tested to the first token of a text response.

A world can now be imagined, interrogated, edited, and carried forward within the window of human conversational patience.

Rho-1 as a steerable real-time world model #

Fast generation alone just yields a faster render farm. But generation that is both real-time and controllable unlocks something fundamentally different: a world model you can simulate and steer.

Rho-1 streams continuously, clip after clip, each picking up exactly where the last left off, so time inside the simulation never stops. New instructions can arrive at any moment. A command enters through the understanding stream, updates the underlying state memory, and moments later the generation stream renders the new trajectory. The world shifts without a single cut.

Because generation is a continuous rollout, a scene can be monitored live and steered at key moments:

Real-time simulation enables new applications.

Generative environments on demand: An environment that advances under live input is no longer a static asset pipeline, it becomes an interactive simulation. Instead of manually engineering edge-case 3D worlds to benchmark autonomous driving or train embodied robotic policies, developers can spin up reactive environments and perturb them on the fly.

Infinite, steerable livestreams: Running real-time understanding and generation within the same attention context dissolves the boundary between viewer and director. You can watch an uninterrupted stream, adjust the physics or weather mid-frame, inspect object state, and extend the horizon indefinitely, all without a restart.

Rho-1 as a world-language-action model for robotics #

Because physical actions and future video frames decode from the same latent state, Rho-1 allows a policy to mentally simulate the next few seconds of an environment and output the trajectory to enact it.

Planning is no longer an external planner querying an auxiliary world model. The network predicting future camera observations is the exact network dictating joint actuation. With actions treated natively as continuous tokens alongside text and pixels, the model does not require an ad-hoc robotics wrapper; the action channels share the core attention space.

Even under pure text conditioning, Rho-1 exhibits an intuitive grasp of physical contact and rigid-body mechanics. Grippers secure objects that remain solid rather than morphing or clipping into surfaces:

In practice, to scale beyond scarce teleoperated robot logs, Rho-1 can be paired with our Inverse Dynamics Model. While most internet video lacks logged steering angles or joint torques, an IDM observes raw video and infers the underlying control signals. This unlocks web-scale video as action-labeled pretraining data: Reka IDM reconstructs the implicit actions, and Rho-1 ingests those actions as native tokens within its multi-modal state space.

Limitations #

The demos above show what Rho-1 can do today, but real-world deployment requires being candid about the failure modes:

  • Long-Horizon Drift: Extended rollouts degrade structurally before they degrade visually. A 30-second stream may keep photorealistic texture and fine detail while drifting into a structurally incompatible room layout.
  • Grounding Across Time: Object detection and coordinate grounding work reliably on static images, but not yet across video.
  • Editing Stability: Targeted visual editing remains nascent. Our conversational demo shows how an edit works, but consistency across diverse prompts is brittle.
  • Resolution: Native video rollouts are currently capped at 672×384.

We believe most of these failures can be attributed to the modest scale of our training compute and data, and we expect them to narrow as we scale both.

One model: a symmetric architecture #

Most multimodal systems are asymmetric. Vision-Language Models (VLMs) ingest text, images, and video, but only emit text. Video diffusion models take in multimodal conditioning, but only emit pixels. This input-output imbalance confines each model to half the interactive loop.

Rho-1 restores full symmetry. Every signal entering or leaving the network belongs to one of two native formats:

  • Discrete tokens encode text, symbolic reasoning, and high-level commands.
  • Continuous tokens encode image latents, video frames, robotic actions, and proprioception.

Because outputs share the same formats as inputs, anything the model generates can be fed straight back into its context. And because each modality keeps its natural form, Rho-1 never has to quantize rich continuous dynamics into lossy discrete tokens, or force language into continuous approximations.

Within each transformer block, compute is split across two expert weight streams: an understanding stream for language and visual parsing, and a generation stream that denoises latents into images and video. Crucially, these streams do not run in isolation:

  • Shared Attention and State: Tokens from both streams attend to one another, operating over the same KV cache. User instructions, chain-of-thought traces, prior video frames, and emitted motor policies all accumulate in one context.
  • Zero-Latency Handoffs: The understanding stream decides the output modality. When a response requires visual synthesis, it emits a discrete handoff token, and the generation stream immediately begins rendering from the full accumulated state.

Rho-1 is optimized end-to-end under two concurrent objectives:

  • Next-token prediction for discrete sequences.
  • Flow matching for continuous image, video, and action generation.

Because both streams share attention and a common representation space, improvements do not stay siloed. We hypothesize that when the model sharpens its spatial reasoning during text pretraining, its video consistency improves. When it learns richer physical dynamics from embodied trajectories, its real-time steering becomes more grounded. Progress compounds across the entire network at once.

Early on the compute curve #

Every result in this post comes from a checkpoint trained from scratch on just 320 H100 GPUs for three months, a tiny fraction of the compute behind modern frontier language and video models. Rho-1 is a functional proof-of-concept and an architectural direction, not a finished product.

To us, that modest compute budget is the most compelling takeaway. The failure modes highlighted earlier, from structural drift and limited temporal grounding to brittle edits and lower native resolution, are the kinds of problems that have historically improved with scale.

More importantly, scaling a unified architecture yields compounding returns. In a fragmented pipeline, additional compute must be divided among an LLM, an image generator, a video backbone, and an action policy. In Rho-1, every added FLOP goes into one model, so gains in reasoning, visual generation, and physical control can reinforce one another rather than being trained in isolation.

Infrastructure for serving omni models #

A unified architecture fundamentally shifts the systems economics of inference. When reasoning, multimodal generation, visual grounding, and motor control collapse into a single set of weights, conventional LLM serving frameworks, designed for autoregressive text generation over a growing KV cache, are no longer enough.

We pursue this research not only to explore frontier model designs, but to build the runtime systems required to serve them at scale. Interleaving continuous diffusion trajectories with discrete autoregressive decoding exposes unprecedented runtime bottlenecks:

  • Heterogeneous Memory Layouts: High-throughput discrete KV caches must coexist in HBM alongside transient continuous latent states without causing fragmentation.
  • Shifting Compute Intensity: Workloads swing between compute-bound dense attention during visual synthesis and memory-bandwidth-bound autoregressive decoding during text generation.
  • Hard Real-Time Latency: Interactive steering leaves little room for pipeline bubbles. The system must ingest new instructions and update trajectories within fractions of a second.

Every finding from Rho-1 feeds directly into our production serving engine on Reka Cloud. The choices behind Rho-1, from token representation to kernel scheduling, shape how we optimize memory bandwidth, schedule execution, and co-locate multimodal workloads in production. Scaling unified models is only half the equation. The other half is building the runtime infrastructure that makes them fast, reliable, and economically viable to deploy.

Unedited, single-pass outputs across diverse domains, presented at Rho-1’s native runtime resolution. 60 clips, 10 domains, one checkpoint with no per-domain fine-tuning. Each clip represents an active latent state, an entry point into an environment the model can simulate, steer, and carry forward indefinitely.

Get in touch #

Rho-1 is currently available as a research preview.

If you are building frontier applications—embodied robotics, interactive simulation environments, or closed-loop vision-action systems—or are looking to deploy optimized omni architectures at scale, we’d love to collaborate: contact@reka.ai. Reka Cloud

Interested in partnering?

What we learn building Rho-1 goes straight into our serving stack on Reka Cloud. If you want to train, customize or deploy multimodal models on infrastructure engineered for performance, efficiency and scale, tell us about your project.

Further reading #

More research from Reka Labs

More research from Reka Labs

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @reka 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/reka-rho-1-collapsin…] indexed:0 read:10min 2026-10-05 · —