Reka Rho-1: 19B Model Cuts 13.8s Pipelines to 7.0s and Emits Robot Actions
Reka has released a research preview of Rho-1, a 19B-parameter omni-reasoning model that processes text, imagery, video, and robot actions inside a single neural network context window. Instead of handing work between specialist models, Rho-1 maintains one shared state across modalities, trading pipeline modularity for lower latency and preserved visual context.
Two streams, one KV cache
Every transformer block in Rho-1 holds two expert weight streams. The understanding stream parses language and vision; the generation stream denoises latents into images and video. Both share attention and a single KV cache. When a reply needs pixels, the understanding stream emits a discrete handoff token, and the generation stream renders from the full accumulated state.
Inputs and outputs use two native token formats. Text, high-level commands, and symbolic reasoning are discrete tokens trained with next-token prediction. Image latents, video frames, robot actions, and proprioception are continuous tokens trained with flow matching. Because outputs stay in context, Reka's five-turn demo has the model draw a lighthouse, emit bounding-box coordinate tokens without a separate detector, animate the scene, edit it into a snowstorm, and explain what changed. The animation reuses the in-context image representation rather than a re-encoded copy, and the whole session makes no tool call and no second model.
Speed, distillation, and robotics
Reka measured Rho-1's first clip at 7.0 seconds, against an illustrative 13.8 seconds for a multi-agent pipeline. The base model uses 99 denoising passes per clip and generates video at 0.79x real-time median; Reka reports a watchable stream starts in roughly six seconds. A distilled variant cuts denoising to 8 steps. Reka reports minimal quality loss and returns a 5.3-second clip in about one second. In Reka's internal tests, the distilled model matched the fastest dedicated image models and was the quickest tested to the first text token, but these are vendor-run tests, not independent benchmarks. Native video rollouts are capped at 672x384.
Rho-1 also streams continuously: new instructions enter through the understanding stream and update state mid-rollout. Reka shows one environment opening forked into bank left and bank right continuations. For robotics, actions and future video frames decode from the same latent state. Proprioception is a continuous channel, so the policy can mentally simulate the next few seconds before acting. In a LIBERO simulation episode, Rho-1 emits 7 action channels. To compensate for scarce teleoperation logs, Reka pairs the model with its Inverse Dynamics Model, which infers control signals from raw video.
| Feature | Reka Rho-1 | ByteDance BAGEL | BAAI Emu3.5 | Google Genie 3 |
|---|---|---|---|---|
| Parameters | 19B | 14B total, 7B active (MoT) | 34B | Not disclosed |
| Inputs | Text, image, video, actions, proprioception | Text, image | Interleaved text and image | Text prompt, navigation inputs |
| Outputs | Text, image, video, actions, proprioception | Text, image | Interleaved text and image | Interactive video world |
| Native video generation | Yes, capped at 672×384 | No | No (image frames, not native clips) | Yes, 720p at 24 fps |
| Real-time steering | Yes, continuous rollouts | No | No | Yes, promptable world events |
| Robot actions | Native continuous action tokens | No | Embodied manipulation demos | Takes navigation actions, does not emit them |
| Open weights | No | Yes, Apache 2.0 | Yes, Apache 2.0 | No |
| Access today | Research preview via contact@reka.ai | Hugging Face, GitHub | Hugging Face, emu.world app | Project Genie for Google AI Ultra subscribers |
AI Mastery analysis
Reka's move to collapse the multimodal stack into a 19B model departs from modular pipelines. The main gain is removing inter-model communication overhead: separate vision, detection, and reasoning models pay serialization and round-trip costs that a shared pipeline architecture avoids. But the monolithic design also limits component-level scaling. Teams cannot swap in a small text router or scale video generation independently; every call runs through the full 19B parameter set. Rho-1 is a research preview trained on 320 H100 GPUs for three months, with no public API, pricing, or open weights, so infrastructure teams cannot yet test Reka's latency claims against their own stacks. The release points toward unified token spaces for text, imagery, and physical simulation, but production reliability and inference economics will decide whether this replaces optimized multi-model pipelines.
Sources
Frequently asked questions
What is Reka Rho-1?
Reka Rho-1 is a 19B-parameter omni-reasoning model released as a research preview. It processes text, images, video, and robot actions in one context window using two expert weight streams and one shared KV cache.
How fast does Reka Rho-1 generate video?
Reka measured a first clip at 7.0 seconds. The base model generates at 0.79x real-time median and uses 99 denoising passes per clip; a distilled 8-step version returns a 5.3-second clip in about one second, according to Reka's internal tests.
Does Reka Rho-1 have open weights or an API?
No. Rho-1 is currently a research preview with no public API, pricing, or open weights. Reka trained it on 320 H100 GPUs over three months, but has not released the weights for external evaluation.
How does Rho-1 handle robot actions?
Rho-1 treats robot actions and proprioception as continuous tokens in the same attention space as images and video. In a LIBERO simulation episode, it emits 7 action channels, and Reka pairs it with an Inverse Dynamics Model that infers control signals from raw video.
Related Reading
AXIS Lifts π0.5 to 88.8 on LIBERO-Plus With 50,129 Browser-Collected Trajectories
Axis Robotics releases AXIS: 207 manipulation tasks, 50,129 trajectories collected via browser teleoperation, lifting π0.5 from 83.9 to 88.8 on LIBERO-Plus.
Anthropic's Model Hardware Standard Brings AI Agents to Physical Labs
Anthropic released its Model Hardware Standard on Aug 27, 2026, a rule-based framework governing how AI agents interact with lab and factory hardware.
Gemini Omni 1.1 Flash: 40s Extension, First/Last Frame, and 4K Upscaling
Google ships gemini-omni-1.1-flash with 10s context reads, first/last frame interpolation, 360p draft mode at ⅓ cost, and upscaled 4K output.