Reka Rho-1: 19B Model Cuts 13.8s Pipelines to 7.0s and Emits Robot Actions

October 6, 2026 • news
Robotics

Reka has released a research preview of Rho-1, a 19B-parameter omni-reasoning model that processes text, imagery, video, and robot actions inside a single neural network context window. Instead of handing work between specialist models, Rho-1 maintains one shared state across modalities, trading pipeline modularity for lower latency and preserved visual context.

Two streams, one KV cache

Every transformer block in Rho-1 holds two expert weight streams. The understanding stream parses language and vision; the generation stream denoises latents into images and video. Both share attention and a single KV cache. When a reply needs pixels, the understanding stream emits a discrete handoff token, and the generation stream renders from the full accumulated state.

Inputs and outputs use two native token formats. Text, high-level commands, and symbolic reasoning are discrete tokens trained with next-token prediction. Image latents, video frames, robot actions, and proprioception are continuous tokens trained with flow matching. Because outputs stay in context, Reka's five-turn demo has the model draw a lighthouse, emit bounding-box coordinate tokens without a separate detector, animate the scene, edit it into a snowstorm, and explain what changed. The animation reuses the in-context image representation rather than a re-encoded copy, and the whole session makes no tool call and no second model.

Speed, distillation, and robotics

Reka measured Rho-1's first clip at 7.0 seconds, against an illustrative 13.8 seconds for a multi-agent pipeline. The base model uses 99 denoising passes per clip and generates video at 0.79x real-time median; Reka reports a watchable stream starts in roughly six seconds. A distilled variant cuts denoising to 8 steps. Reka reports minimal quality loss and returns a 5.3-second clip in about one second. In Reka's internal tests, the distilled model matched the fastest dedicated image models and was the quickest tested to the first text token, but these are vendor-run tests, not independent benchmarks. Native video rollouts are capped at 672x384.

Rho-1 also streams continuously: new instructions enter through the understanding stream and update state mid-rollout. Reka shows one environment opening forked into bank left and bank right continuations. For robotics, actions and future video frames decode from the same latent state. Proprioception is a continuous channel, so the policy can mentally simulate the next few seconds before acting. In a LIBERO simulation episode, Rho-1 emits 7 action channels. To compensate for scarce teleoperation logs, Reka pairs the model with its Inverse Dynamics Model, which infers control signals from raw video.

Feature Reka Rho-1 ByteDance BAGEL BAAI Emu3.5 Google Genie 3
Parameters 19B 14B total, 7B active (MoT) 34B Not disclosed
Inputs Text, image, video, actions, proprioception Text, image Interleaved text and image Text prompt, navigation inputs
Outputs Text, image, video, actions, proprioception Text, image Interleaved text and image Interactive video world
Native video generation Yes, capped at 672×384 No No (image frames, not native clips) Yes, 720p at 24 fps
Real-time steering Yes, continuous rollouts No No Yes, promptable world events
Robot actions Native continuous action tokens No Embodied manipulation demos Takes navigation actions, does not emit them
Open weights No Yes, Apache 2.0 Yes, Apache 2.0 No
Access today Research preview via contact@reka.ai Hugging Face, GitHub Hugging Face, emu.world app Project Genie for Google AI Ultra subscribers

AI Mastery analysis

Reka's move to collapse the multimodal stack into a 19B model departs from modular pipelines. The main gain is removing inter-model communication overhead: separate vision, detection, and reasoning models pay serialization and round-trip costs that a shared pipeline architecture avoids. But the monolithic design also limits component-level scaling. Teams cannot swap in a small text router or scale video generation independently; every call runs through the full 19B parameter set. Rho-1 is a research preview trained on 320 H100 GPUs for three months, with no public API, pricing, or open weights, so infrastructure teams cannot yet test Reka's latency claims against their own stacks. The release points toward unified token spaces for text, imagery, and physical simulation, but production reliability and inference economics will decide whether this replaces optimized multi-model pipelines.

Sources

Frequently asked questions

What is Reka Rho-1?

Reka Rho-1 is a 19B-parameter omni-reasoning model released as a research preview. It processes text, images, video, and robot actions in one context window using two expert weight streams and one shared KV cache.

How fast does Reka Rho-1 generate video?

Reka measured a first clip at 7.0 seconds. The base model generates at 0.79x real-time median and uses 99 denoising passes per clip; a distilled 8-step version returns a 5.3-second clip in about one second, according to Reka's internal tests.

Does Reka Rho-1 have open weights or an API?

No. Rho-1 is currently a research preview with no public API, pricing, or open weights. Reka trained it on 320 H100 GPUs over three months, but has not released the weights for external evaluation.

How does Rho-1 handle robot actions?

Rho-1 treats robot actions and proprioception as continuous tokens in the same attention space as images and video. In a LIBERO simulation episode, it emits 7 action channels, and Reka pairs it with an Inverse Dynamics Model that infers control signals from raw video.

Free interactive tools for the decisions this piece raises.

Related Reading