RLDX-1 Technical Report

Dongyoung Kim; Huiwon Jang; Myungkyu Koo; Suhyeok Jang; Taeyoung Kim; Beomjun Kim; Byungjun Yoon; Changsung Jang; Daewon Choi; Dongsu Han; Donguk Lee; Heeseung Kwon; Hojin Jeon; Jaehyun Kang; Jaekyoung Bae; Jihyuk Lee; Jimin Lee; John Won; Joonwoo Ahn; Junhyeong Park; Junyoung Sung; Kyungmin Lee; Minseong Han; Minsung Yoon; Sejune Joo; Seonil Son; Seungcheol Park; Seunggeun Cho; Seungjun Moon; Seungku Kim; Yonghoon Dong; Yongjin Cho; Youngchan Kim; Chang Hwan Kim; Dohyeon Kim; Heecheol Kim; Heewon Lee; Hensen Ahn; Hyungkyu Ryu; Hyunsoo Choi; Hyunsoo Shin; Jaeheon Jung; Jaewoo Kim; Jinwook Kim; Joochul Chang; Joonsoo Kim; Junghun Park; Jungwoo Park; Junho Cho; Junhyeok Park; Junwon Lee; Kangwook Lee; Kwanghoon Kim; Kyoungwhan Choe; Manoj Bhadu; Nayoung Oh; Sangjun Kim; Sangwoo Kim; Seunghoon Shim; Seunghyun Kim; Seungjun Lee; Seungyup Ka; Sungryol Yang; Wook Jung; Yashu Shukla; Yeonjae Lee; Yeonwoo Bae; Jinwoo Shin

arXiv:2605.03269·cs.RO·May 7, 2026

RLDX-1 Technical Report

Dongyoung Kim, Huiwon Jang, Myungkyu Koo, Suhyeok Jang, Taeyoung Kim, Beomjun Kim, Byungjun Yoon, Changsung Jang, Daewon Choi, Dongsu Han, Donguk Lee, Heeseung Kwon, Hojin Jeon, Jaehyun Kang, Jaekyoung Bae, Jihyuk Lee, Jimin Lee, John Won, Joonwoo Ahn, Junhyeong Park

PDF

1 Repo 10 Models

TL;DR

RLDX-1 is a versatile robotic policy that integrates multiple modalities to outperform recent vision-language-action models in complex real-world manipulation tasks, including humanoid control.

Contribution

The paper introduces RLDX-1, a novel multi-stream transformer architecture with system-level optimizations for dexterous manipulation, surpassing existing models in simulation and real-world tests.

Findings

01

RLDX-1 achieves 86.8% success in ALLEX humanoid tasks.

02

It outperforms recent models like π_{0.5} and GR00T N1.6 in benchmarks.

03

RLDX-1 demonstrates reliable control of high-DoF humanoid robots.

Abstract

While Vision-Language-Action models (VLAs) have shown remarkable progress toward human-like generalist robotic policies through the versatile intelligence (i.e. broad scene understanding and language-conditioned generalization) inherited from pre-trained Vision-Language Models, they still struggle with complex real-world tasks requiring broader functional capabilities (e.g. motion awareness, long-term memory, and physical sensing). To address this, we introduce RLDX-1, a general-purpose robotic policy for dexterous manipulation built on the Multi-Stream Action Transformer (MSAT), an architecture that unifies these capabilities by integrating heterogeneous modalities through modality-specific streams with cross-modal joint self-attention. RLDX-1 further combines this architecture with system-level design choices, including data synthesis for rare manipulation scenarios, learning…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

rlwrld/RLDX-1
github

Models

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.