
Shubhanshu Khatana
159 posts

Shubhanshu Khatana
@hatebunnyplzzz
building @fpv_labs, ex-Builder @Lossfunk, ex-LCS2 @IITD, ex-@SHLglobal Robots--Vision--RL




1/ Announcing Vidur, an end-to-end system to generate fine-grained action labels on raw robot and human videos. Vidur leads across all metrics on WGO and EgoExoLearn Bench across temporal segmentation, semantic precision, semantic recall, and end-to-end action-label quality 🧵




High-quality motion reference data is key for humanoid skill learning 🤖🕺💃 A natural idea is to leverage human motions and “translate” them to humanoid motions, a process known as retargeting. For interaction-rich tasks such as scene interaction and loco-manipulation, retargeting is challenging: it must ensure motion consistency, smoothness, kinematic feasibility (no artifacts like penetration or foot skating), and scalability (one framework can handle thousands of motions). Excited to release OmniRetarget — a scalable retargeting method with a 4-hour high-quality humanoid motion dataset for interaction-rich tasks. OmniRetarget takes an interaction-preserving perspective: we optimize Laplacian deformation between source and target interaction meshes while enforcing kinematic constraints, producing consistent, smooth, and feasible trajectories at scale. Even better, OmniRetarget can efficiently augment motions by varying terrains, objects, and initial poses. This high-quality interaction-preserving retargeting enables a minimal RL setup to execute long-horizon (up to 30s) agile, interaction-rich skills. All tasks in the video share just 5 rewards, 4 domain randomization terms, and rely only on proprioception. More details: omniretarget.github.io

Can a commodity smartphone replace a $10,000 data rig? We're presenting MobileEgo Anywhere at ICRA 2026 - open infrastructure for collecting hour-long egocentric trajectories on commodity hardware for downstream robot policies. 📍 Strauss 3 | 🕒 3:00 - 4:00 PM today Come say hi 👋

open sourcing Marlin-2B 🐟 a tiny VLM to extract structured information from videos Marlin is finetuned for two questions devs want to ask in their videos: what is happening, and when? Best open model in its weight class, competitive with Gemini-2.5-flash at only 2B params 🧵



🚨 New Paper (ICRA Workshop, 2026) MobileEgo Anywhere: Open infrastructure for long-horizon egocentric data on commodity hardware. Existing robotics egocentric datasets are limited by short episode durations and rely on gated hardware. We release a framework designed to facilitate the collection of robust, hour-plus egocentric trajectories on a commodity iPhone.


🚨 New Paper (ICRA Workshop, 2026) MobileEgo Anywhere: Open infrastructure for long-horizon egocentric data on commodity hardware. Existing robotics egocentric datasets are limited by short episode durations and rely on gated hardware. We release a framework designed to facilitate the collection of robust, hour-plus egocentric trajectories on a commodity iPhone.

Introducing Project Stera by FPV Labs, an open data infra for embodied AI research. Project Stera includes Stera-10M, with 10M+ frames of long-horizon data with persistent state tracking, and an open-source pipeline that converts raw data into training-ready formats.








