
Jaden Clark
63 posts

Jaden Clark
@jadenvclark
PhD @Stanford @KnightHennessy. AI, robotics, and conservation





We are also excited to release iPhUMI ("eye-foo-me”)! It solves the localization challenges with the GoPro UMI, enabling rapid data collection across diverse environments and tasks. During deployment, iPhUMI lets you command your robot via demonstration. github.com/real-stanford/… (7/8)

Can we enable robots to develop a sense of touch without forgetting what they learned from large-scale vision-only pretraining? Introducing MultiSensory World Model (MuSe) 🌍: A new approach for finetuning visuomotor policies on minimal data from new sensor modalities, such as force/torque (F/T) With Muse, touch learned later improves skills learned earlier — a small amount of F/T data on new tasks improves zero-shot on diverse pretraining tasks that were never supervised with F/T We believe MuSe provides a practical pathway towards training multisensory foundation models that leverage both abundant vision data, and smaller multisensory datasets 🧵👇

Can we enable robots to develop a sense of touch without forgetting what they learned from large-scale vision-only pretraining? Introducing MultiSensory World Model (MuSe) 🌍: A new approach for finetuning visuomotor policies on minimal data from new sensor modalities, such as force/torque (F/T) With Muse, touch learned later improves skills learned earlier — a small amount of F/T data on new tasks improves zero-shot on diverse pretraining tasks that were never supervised with F/T We believe MuSe provides a practical pathway towards training multisensory foundation models that leverage both abundant vision data, and smaller multisensory datasets 🧵👇

Can we enable robots to develop a sense of touch without forgetting what they learned from large-scale vision-only pretraining? Introducing MultiSensory World Model (MuSe) 🌍: A new approach for finetuning visuomotor policies on minimal data from new sensor modalities, such as force/torque (F/T) With Muse, touch learned later improves skills learned earlier — a small amount of F/T data on new tasks improves zero-shot on diverse pretraining tasks that were never supervised with F/T We believe MuSe provides a practical pathway towards training multisensory foundation models that leverage both abundant vision data, and smaller multisensory datasets 🧵👇

Can we enable robots to develop a sense of touch without forgetting what they learned from large-scale vision-only pretraining? Introducing MultiSensory World Model (MuSe) 🌍: A new approach for finetuning visuomotor policies on minimal data from new sensor modalities, such as force/torque (F/T) With Muse, touch learned later improves skills learned earlier — a small amount of F/T data on new tasks improves zero-shot on diverse pretraining tasks that were never supervised with F/T We believe MuSe provides a practical pathway towards training multisensory foundation models that leverage both abundant vision data, and smaller multisensory datasets 🧵👇







