Yang Fu

46 posts

Yang Fu

Yang Fu

@yangfu21

Research Scientist @NVIDIA | PhD @UCSanDiego

United States Katılım Ağustos 2017
307 Takip Edilen351 Takipçiler
Sabitlenmiş Tweet
Yang Fu
Yang Fu@yangfu21·
Check out our recent work 𝐒𝐑-𝟑𝐃: a unified framework that brings 𝟑𝐃 𝐬𝐩𝐚𝐭𝐢𝐚𝐥 𝐫𝐞𝐚𝐬𝐨𝐧𝐢𝐧𝐠 and 𝐫𝐞𝐠𝐢𝐨𝐧-𝐥𝐞𝐯𝐞𝐥 𝐩𝐫𝐨𝐦𝐩𝐭𝐢𝐧𝐠 into VLMs for powerful cross-view understanding!
An-Chieh Cheng@anjjei

Let your robot peek around corners, size up the gap between chairs, and know exactly where everything sits. 🤖 Our new work SR-3D masters this kind of spatial reasoning. It learns from multiple views to understand distances, layouts, and how objects relate in 3D space.

English
0
3
3
3.2K
Yang Fu retweetledi
Jianglong Ye
Jianglong Ye@jianglong_ye·
How do we make dexterous hands handle both power and precision tasks with ease? 🫳👌🫰 We introduce Power to Precision (💪➡️🎯), our new paper that optimizes both control and fingertip geometry to unlock robust manipulation from power grasp to fine-grained manipulations. With simplified finger motions and augmented fingertips, the hand can perform diverse motions from pinching a nut🔩 to handling a pan🍳. Check the demos below🎥.
English
13
76
390
110K
Yang Fu retweetledi
Jiarui Xu
Jiarui Xu@Jerry_XU_Jiarui·
Test-Time Training (TTT) is available in video generation now! We can directly generate complete one-minute video, with great temporal and spatial coherence. We created more episodes of Tom and Jerry (my favorite cartoon in childhood) with our model. test-time-training.github.io/video-dit/ Enjoy more in 🧵
English
7
27
199
15.5K
Yang Fu retweetledi
Xueyan Zou
Xueyan Zou@xyz2maureen·
[1/n] We are releasing M3 (#ICLR2025): a Gaussian Splatting method that builds LMM memories for arbitrary scenes. 🔥 [Efficient] 16 degrees in each Gaussian primitive for one LMM. 🔥 [Alignment] The rendered features are directly in the source LMM embedding space.
English
2
46
322
53.5K
Yang Fu retweetledi
Yuzhe Qin
Yuzhe Qin@QinYuzhe·
Meet our first general-purpose robot at @DexmateAI dexmate.ai/vega Adjustable height from 0.66m to 2.2m: compact enough for an SUV, tall enough to reach those impossible high shelves. Powerful dual arms (15lbs payload each) and omni-directional mobility for ultimate versatility. More videos showcasing Vega in action coming this week! Stay tuned to see this robot partner.
Dexmate@DexmateAI

Introducing Vega: 🤖 @DexmateAI's newest robot that makes complex manipulation tasks simple. ✨ A step closer to intelligence and automation. 🚀 🎥 Watch now: youtu.be/PecqfiJNwQI #Robotics #AI #Automation

English
13
34
215
32.1K
Yang Fu retweetledi
Binghao Huang
Binghao Huang@binghao_huang·
Can we bring human-like Touch to robots🤖? Introducing our CoRL work on 3D-ViTac. Humans rely on both vision 👁️ and touch 🫳 for complex tasks. With combined visual-tactile sensing, robots can now tackle challenging tasks, like precise in-hand reorientation, fragile objects grasping. Website: binghao-huang.github.io/3D-ViTac/ #Robotics #CoRL2024 #Touch #tactile #AI #ML
English
21
33
189
49.6K
Yang Fu retweetledi
Jiawei Yang
Jiawei Yang@JiaweiYang118·
Very excited to get this out: “DVT: Denoising Vision Transformers”. We've identified and combated those annoying positional patterns in many ViTs. Our approach denoises them, achieving SOTA results and stunning visualizations! Learn more on our website: jiawei-yang.github.io/DenoisingViT/
English
8
82
408
48.5K
Yang Fu retweetledi
Yang Fu retweetledi
An-Chieh Cheng
An-Chieh Cheng@anjjei·
🌟Introducing "🤖SpatialRGPT: Grounded Spatial Reasoning in Vision Language Model" anjiecheng.me/SpatialRGPT SpatialRGPT is a powerful region-level VLM that can understand both 2D and 3D spatial arrangements. It can process any region proposal (e.g., boxes or masks) and provide answers to complex spatial reasoning questions. 🧵(1/n)
English
10
108
464
102.5K
Yang Fu retweetledi
Jiteng Mu
Jiteng Mu@JitengMu·
We introduce🌟Editable Image Elements🥳, a new disentangled and controllable latent space for diffusion models, that allows for various image editing operations (e.g., move, resize,  de-occlusion, object removal, variations, composition) jitengmu.github.io/Editable_Image… More details🧵👇
English
5
32
206
49K
Yang Fu retweetledi
Isabella Liu
Isabella Liu@Isabella__Liu·
Want to obtain time-consistent dynamic mesh from monocular videos? Introducing: Dynamic Gaussians Mesh: Consistent Mesh Reconstruction from Monocular Videos liuisabella.com/DG-Mesh/ We reconstruct meshes with flexible topology change and build the corresp. across meshes. 🧵(1/n)
English
9
49
196
71.7K
Yang Fu retweetledi
Xiaolong Wang
Xiaolong Wang@xiaolonw·
I have been cleaning my daughter's mess for more than two years now. Last weekend our robot came to home to do the job for me. 🤖 wholebody-b1.github.io Our new work on visual whole-body control learns a policy to coordinate the robot legs and arms for mobile manipulation. See how the legs bend on grasping objects on the ground based on visual inputs autonomously. (It is NOT teleoperation!) We again adopt a Sim2Real approach and our policy generalizes to beaches, forests, and streets in San Diego. 👇🧵
English
18
109
641
159.3K
Yang Fu retweetledi
Xiaolong Wang
Xiaolong Wang@xiaolonw·
If you look at the best image/video diffusion models, they are still not able to get the hands quite right, especially when interacting with objects ✍️ We present HOIDiffusion #CVPR2024, a way to generate accurate and realistic hand-object interaction images, in diverse poses, scenes, and styles. mq-zhang1.github.io/HOIDiffusion/ This is not just for synthesizing beautiful pixels, but also a way to generate infinite pairs of hand-object images and 3D GTs data.
English
2
23
142
33.7K
Yang Fu retweetledi
Xiaolong Wang
Xiaolong Wang@xiaolonw·
We have seen a lot of legged robots doing navigation in the wild. But how about mobile manipulation in the wild? I have been pushing the direction of learning a unified, efficient, and dynamic 3D representation of scenes (for navigation) and objects (for manipulation) for the past two years. And now we have GeFF --- our large-scale, generalizable feature field, that combines the speed of a feed-forward neural network with the rich semantics from Foundation Models, to handle dynamically changing scenes, and enable open-ended, language-grounded scene and object understanding. geff-b1.github.io
English
5
45
225
43K
Bowen Cheng
Bowen Cheng@bowenc0221·
Today is my last day @Tesla autopilot. It was a pleasant one and a half years: working with talented people, learning how to build good product, etc. But my journey to AGI will not stop, I will soon join @OpenAI post-training team to build multimodal models.
English
59
44
1.2K
243.6K
Yang Fu retweetledi
Xiaolong Wang
Xiaolong Wang@xiaolonw·
Let’s think about humanoid robots outside carrying the box. How about having the humanoid come out the door, interact with humans, and even dance? Introducing Expressive Whole-Body Control for Humanoid Robots: expressive-humanoid.github.io See how our robot performs rich, diverse, and expressive motions in the real world 👇🧵
English
87
179
1.3K
309.2K
Yang Fu
Yang Fu@yangfu21·
Check our COLMAP-Free 3D Gaussian Splatting results on Sora videos. It seems we have an opportunity to get infinite 3D data.
Xiaolong Wang@xiaolonw

Is 3D scene generation much closer to being solved all of a sudden? It has been a few days since the release of @OpenAI Sora. We run our COLMAP-Free 3D Gaussian Splatting on the released videos. Our method does not need to pre-process cameras and it seems we can directly just get 3D from the videos. Check out our results here. 🧵👇 (1/n

English
0
3
41
6.4K
Yang Fu retweetledi
Xiaolong Wang
Xiaolong Wang@xiaolonw·
Tired of seeing synthetic objects go round and round, round and round, round and round? Introducing the WildRGB-D dataset! We collect a dataset of real-world RGB-D objects in 360 under cluttered scenes. To download: github.com/wildrgbd/wildr… Website: wildrgbd.github.io ✅ It comes with GT depth, GT camera pose, and 360 views. ✅ A lot of objects : 20k videos, 8.5k objects, 46 classes. ✅ It is cluttered, there is no clean background. ✅ We have single-object, multi-object, and HAND-Object videos. ✅ Supporting multiple downstream tasks: camera pose estimation, view synthesis, 3D reconstruction, 6D pose estimation.
English
4
22
216
19.9K