Jaden Clark

63 posts

Jaden Clark banner
Jaden Clark

Jaden Clark

@jadenvclark

PhD @Stanford @KnightHennessy. AI, robotics, and conservation

Los Angeles, CA Katılım Aralık 2019
403 Takip Edilen394 Takipçiler
Sabitlenmiş Tweet
Jaden Clark
Jaden Clark@jadenvclark·
Can we enable robots to develop a sense of touch without forgetting what they learned from large-scale vision-only pretraining? Introducing MultiSensory World Model (MuSe) 🌍: A new approach for finetuning visuomotor policies on minimal data from new sensor modalities, such as force/torque (F/T) With Muse, touch learned later improves skills learned earlier — a small amount of F/T data on new tasks improves zero-shot on diverse pretraining tasks that were never supervised with F/T We believe MuSe provides a practical pathway towards training multisensory foundation models that leverage both abundant vision data, and smaller multisensory datasets 🧵👇
English
8
46
237
87.8K
Jaden Clark retweetledi
Josh Citron
Josh Citron@josh__citron·
Teleop systems are usually designed for a single embodiment, but do they have to be? Introducing ModPack 🎒: a modular teleoperation interface for bimanual mobile robots. A wearable backpack provides shared infrastructure, robot-specific leader arms adapt to different embodiments, and plug-and-play modules add capabilities like haptic feedback, active perception, and mobile manipulation. 🧵(1/n)
English
11
25
87
19.8K
Jaden Clark retweetledi
Haoyu Xiong
Haoyu Xiong@Haoyu_Xiong_·
Success rate has long been the primary metric for evaluating robot manipulation. What about speed? Today, we introduce ⚡️B-spline Policy (BSP). Instead of predicting discrete fixed-rate action chunks, we parameterize actions as continuous B-spline curves. Together with our system design, BSP enables fast manipulation on low-cost robot arms. This project is co-led by @xshenhan, check out his following threads for more details. 🧵 PS: one of my favorite parts of this project was the first time we saw the robots move significantly faster and smoother than the baselines. The videos below are all real time. 👇
English
22
64
507
66.6K
Jaden Clark retweetledi
Dian Wang
Dian Wang@Dian_Wang_·
Long horizon bimanual mobile manipulation requires reasoning in many coordinate frames: base, L/R hands, etc. In which frame would policy work best? It really depends, so don’t pick. Mixture of Frames Policy: denoise in multiple frame in parallel. 🌐mofpo.github.io (1/9)
English
4
47
210
40.4K
Jaden Clark retweetledi
Jaden Clark retweetledi
Shuran Song
Shuran Song@SongShuran·
I'm increasingly interested in the problem of *multisensory continual learning*, since it feels inevitable for robotics. Unlike vision, many robot sensors (e.g., force/torque, tactile, audio) are highly task- and system- specific. It's unrealistic to expect a single pretraining dataset to contain every future sensor. And as robotics evolves, we'll keep building new sensors. So the question is: Can we plug a new sensor into a pretrained vision-only foundation model without forgetting everything it already knows? Better yet, can the new sensor actually improve the model's existing vision-based skills? That's exactly the question that motivated MuSe 👇
Jaden Clark@jadenvclark

Can we enable robots to develop a sense of touch without forgetting what they learned from large-scale vision-only pretraining? Introducing MultiSensory World Model (MuSe) 🌍: A new approach for finetuning visuomotor policies on minimal data from new sensor modalities, such as force/torque (F/T) With Muse, touch learned later improves skills learned earlier — a small amount of F/T data on new tasks improves zero-shot on diverse pretraining tasks that were never supervised with F/T We believe MuSe provides a practical pathway towards training multisensory foundation models that leverage both abundant vision data, and smaller multisensory datasets 🧵👇

English
6
20
188
23.3K
Jaden Clark retweetledi
Jaden Clark retweetledi
Jaden Clark
Jaden Clark@jadenvclark·
While naive finetuning leads to catastrophic forgetting, MuSe exhibits backward transfer: improving in pretraining tasks that never seen F/T conditioning in training
English
1
0
1
1.2K
Jaden Clark
Jaden Clark@jadenvclark·
MuSe leverages broad pretraining data collected with UMI grippers to improve generalization on contact-rich tasks from finetuning (forward transfer)
English
1
0
1
1K
Jaden Clark
Jaden Clark@jadenvclark·
MuSe outputs force/torque predictions enabling policy to adapt its compliance for forceful yet precise contact rich interaction 💪
English
1
0
2
1K
Jaden Clark
Jaden Clark@jadenvclark·
Can we enable robots to develop a sense of touch without forgetting what they learned from large-scale vision-only pretraining? Introducing MultiSensory World Model (MuSe) 🌍: A new approach for finetuning visuomotor policies on minimal data from new sensor modalities, such as force/torque (F/T) With Muse, touch learned later improves skills learned earlier — a small amount of F/T data on new tasks improves zero-shot on diverse pretraining tasks that were never supervised with F/T We believe MuSe provides a practical pathway towards training multisensory foundation models that leverage both abundant vision data, and smaller multisensory datasets 🧵👇
English
8
46
237
87.8K
Jaden Clark
Jaden Clark@jadenvclark·
Naive finetuning forgets old tasks. We finetune with 2⃣ ER to preserve visual and task generalization and 3⃣ multistage fusion to amplify the new modality
Jaden Clark tweet media
English
1
1
6
1.3K
Jaden Clark
Jaden Clark@jadenvclark·
How does MuSe work? We outline 3 key components: 1⃣ Multisensory future prediction (world modeling): Training the model to predict future actions AND multisensory observations encourages it to learn shared representations
English
1
1
8
1.7K
Jaden Clark retweetledi
Stanford MSL
Stanford MSL@StanfordMSL·
π, But Make It Fly ✈️ We fine-tuned π0, a VLA model pretrained entirely on manipulators, to fly a drone that picks up objects, navigates through gates, and composes both skills from language commands.
English
14
43
366
104.5K
Jaden Clark retweetledi
Xiaomeng Xu
Xiaomeng Xu@XiaomengXu11·
Can we learn whole-body mobile manipulation directly from human demonstrations? Introducing Whole-Body Mobile Manipulation Interface (HoMMI) Egocentric + UMI, 0 teleop -> bimanual & whole-body manipulation, long-horizon navigation, active perception hommi-robot.github.io
English
12
71
334
82.3K
Jaden Clark retweetledi
Zhanyi Sun
Zhanyi Sun@s_zhanyi·
We find that RL post-training can substantially improve BC policies without teaching them anything fundamentally new. So what is RL doing? In DICE-RL, it contracts a broad behavior prior toward high-value modes. (1/n) zhanyisun.github.io/dice.rl.2026/
English
6
42
270
28.1K
Jaden Clark retweetledi
Zeyi Liu
Zeyi Liu@Liu_Zeyi_·
For video generation in robotic applications, looking pretty is usually not enough. Robot manipulation requires understanding how visual observations and 3D geometry evolve over time under agent actions, with temporal coherence and geometric consistency across camera views. We study this challenge in our work (recently accepted by @iclr_conf ), 4D Video Generation for Robot Manipulation, which enforces multi-view 3D consistency via geometric supervision to generate spatio-temporally aligned videos.
English
8
39
312
54.4K
Jaden Clark retweetledi
Moo Jin Kim
Moo Jin Kim@moo_jin_kim·
We release Cosmos Policy 💫: a state-of-the-art robot policy built on a video diffusion model backbone. - policy + world model + value function — in 1 model - no architectural changes to the base video model - SOTA in LIBERO (98.5%), RoboCasa (67.1%), & ALOHA tasks (93.6%) 🧵👇
English
18
107
866
149.5K