Reka Releases Rho-1, a 19-Billion-Parameter Model That Handles Text, Images, Video and Robot Control in One Network
Reka released a research preview of Rho-1, a 19-billion-parameter model that handles text, image, video and robot-action generation in one network, trained on 320 H100 GPUs.
Editor's Note ·
- Clarification:
- The title and summary say Rho-1 handles robot control and robot-action generation in one network. Reka's post demonstrates this through one episode on a LIBERO simulation task and generated gripper video clips; it reports no real-robot results or task success rates, and The Decoder's report relays Reka's description.
- Clarification:
- The article says The Decoder describes Rho-1 as processing 'robot control instructions'. The Decoder says the model 'processes and generates text, images, video, and robot control actions'; the actions are outputs the model generates, not instructions it receives.
Overview
Reka said on October 5 that it is releasing a research preview of Rho-1, a 19-billion-parameter omni-reasoning model that unifies text, image, video and robotic action generation in a single neural network. The Decoder describes it as an omni-model that processes text, images, video and robot control instructions without routing tasks to specialized models.
What We Know
- Architecture. According to Reka, Rho-1 treats all modalities as tokens within one context window rather than chaining modality-specific systems through an agentic pipeline. It is trained end-to-end under two concurrent objectives: next-token prediction for discrete sequences and flow matching for continuous image, video and action generation.
- Robot control. The Decoder reports that the same neural weights are used to predict camera images and to control robot movements, and that Reka used an inverse dynamics model to extract control signals from standard internet videos to address the shortage of robot training data.
- Speed. Reka reports that the base model generates video at 0.79x real-time (median), and that a distilled variant cuts the denoising trajectory from 99 steps to 8 with minimal quality loss, per Reka’s post.
- Training cost. Reka says every result in its post comes from a checkpoint trained from scratch on 320 H100 GPUs for three months. The Decoder reports the same GPU count and duration.
- Background. The Decoder notes that Reka previously released Reka Core in April 2024, a multimodal model that competed with leading AI systems on benchmarks.
What We Don’t Know
- Limitations Reka lists. The company’s post acknowledges long-horizon drift, weak grounding across time, unstable editing, and resolution limits: native video rollouts are currently capped at 672x384.
- Access. Rho-1 is a research preview. Reka directs inquiries about embodied robotics, interactive simulation and vision-action systems to [email protected]. The sources reviewed do not describe public API access, pricing or licensing terms.
- Independent evaluation. The results come from Reka’s own post; the sources reviewed do not include third-party benchmarks of the model.
Analysis
The release is notable less for scale than for design: a single set of weights handling language, video generation and robot actions, built on the 320-GPU, three-month training run Reka’s post describes. Whether the approach holds up outside curated demos depends on the limitations the company itself lists and on independent testing, which has not yet been reported in the sources reviewed.