Hanming Ye

111 posts

Hanming Ye

Hanming Ye

@DozenDucc

@theWaddleLabs | math @ harvard | rsi '23

Cambridge, MA Katılım Temmuz 2024
112 Takip Edilen2K Takipçiler
Sabitlenmiş Tweet
Hanming Ye
Hanming Ye@DozenDucc·
Introducing Waddle Labs: Claude Code for robots. Connect our API to your robot and enter a prompt, then our agents write code to achieve the task in 20 minutes. @yiding_song @theWaddleLabs
English
199
232
2.4K
673.5K
Hanming Ye retweetledi
Karan Brar
Karan Brar@deepmatmul·
Every AI company should own its data, its models, and the research lab that keeps making them better. Today, @thomasboser and I are launching @hiloopai to make that possible. The frontier advantage isn’t access to a model. It’s the research organization continuously improving it. Your product already generates the raw material for better intelligence: proprietary data, production feedback, evaluations, and domain expertise. Very few teams have the research capacity to turn those assets into better models. Hiloop builds and operates that capability with you. Bring us the model your product depends on, the data that makes it different, and an evaluation that defines success. We reproduce your baseline and run an autonomous research campaign against it. Agents pursue competing hypotheses in parallel, build on previous results, and promote only improvements that survive verification. Our researchers validate the winners and bring them into production. The work can run hosted or inside your environment. You retain control of your data, and the resulting models, evaluations, and research artifacts are yours. We’re starting with model training, post-training, and inference optimization, where progress is measurable and the value is immediate. But this isn’t a one-off model improvement. It’s a persistent research capability that begins each campaign with everything learned from the last. Over time, the lab accumulates research memory and improves its own tools, evaluations, and agents. The process used to create better intelligence gets better itself. That’s infrastructure for recursive self-improvement. If your product depends on a model and an important metric has stopped moving, bring us the model you can’t make better. We're much better at research than making videos! Reach out to us: founders@hiloop.ai
English
27
29
184
132.3K
Hanming Ye retweetledi
Matija Milenovic
Matija Milenovic@milenovic925501·
Introducing Rook by @Hop_Aero (YC S26): an autonomous hypersonic cargo rocket that delivers 550 lbs up to 450 miles in ~15 minutes. It launches from a standard 40-foot shipping container and lands on unprepared surfaces—no runways or fixed infrastructure required.
English
302
268
3.9K
1.2M
Hanming Ye
Hanming Ye@DozenDucc·
@NoahHWeiss I saw your work on LinkedIn back then! Was it called "farm" or smth like that?
English
1
0
0
152
Noah Weiss
Noah Weiss@NoahHWeiss·
This is really cool. These guys did something special with their harness! I tried something similar in May, and the LLMs then couldn't figure out how to coordinate basic picking and placing actions. This new sol/Fable gen has massively improved in this area. I did some testing within codex for 5.6sol xHigh, which spent about 4 hours building its own harness and roughly 1.5 hours executing for the clip below. I also tested Fable xHigh in claude code, which took about 2 hours to set up but eventually got stuck in a loop after 3 hours during its attempt. Both seemed to improve dramatically over the course of the run; I believe the chat log drove RL-like self-improvement. Raw model capability is finally crossing the threshold where spatial reasoning works, but it's still some way from being competitive with VLAs on inference speed + cost. GPT 5.6sol xhigh "👇
Hanming Ye@DozenDucc

Introducing Waddle Labs: Claude Code for robots. Connect our API to your robot and enter a prompt, then our agents write code to achieve the task in 20 minutes. @yiding_song @theWaddleLabs

English
1
0
3
1.2K
Hanming Ye
Hanming Ye@DozenDucc·
@myanvoos Curious what challenges you ran into? We prob faced the same things initially
English
1
0
0
237
myanvoos
myanvoos@myanvoos·
This is really really cool and reminds me of a personal project back in June when I had Fable pilot my SO-101 arm to pick up cubes, with explicit instructions for it to write policies-as-code to do the task. There were some challenges with this ofc which makes me curious how these guys are approaching it
Hanming Ye@DozenDucc

Introducing Waddle Labs: Claude Code for robots. Connect our API to your robot and enter a prompt, then our agents write code to achieve the task in 20 minutes. @yiding_song @theWaddleLabs

English
2
0
0
466
Hanming Ye
Hanming Ye@DozenDucc·
Introducing Waddle Labs: Claude Code for robots. Connect our API to your robot and enter a prompt, then our agents write code to achieve the task in 20 minutes. @yiding_song @theWaddleLabs
English
199
232
2.4K
673.5K
Hanming Ye
Hanming Ye@DozenDucc·
Sometimes code-as-policy works best (eg. repetitive tasks in factories), sometimes robot models work best (eg. walking a humanoid, controlling 5-fingered hand). People over-index on models, but we still need em. What's surprising is that code-as-policy and LLM controlling robots is far more capable than expected. Ths should unlock faster deployments and automated data collection.
English
0
0
3
417
Shubhanshu Khatana
Shubhanshu Khatana@hatebunnyplzzz·
@DozenDucc @yiding_song @theWaddleLabs is the idea inspired by Cap-X study from NVIDIA? looks impressive Interested to see what is the performance ceiling on some close loop evaluations including approaches involving world models as well. or even simulation like RoboCasa365
English
1
0
0
184
Hanming Ye
Hanming Ye@DozenDucc·
@DJiafei Also using TopReward and Robometer to judge rollouts :)
English
0
0
4
356
Hanming Ye
Hanming Ye@DozenDucc·
@DJiafei Yes! Models beget users beget data begets models
English
0
0
3
307
Hanming Ye
Hanming Ye@DozenDucc·
@tomcocobrico @NickADobos Nah thanks Jeffrey you raise a very valid point (also its X). We've been using code-as-policy to collect tons of data internally - let's see where this goes
English
1
0
1
47
Jeffrey 杰弗瑞
Jeffrey 杰弗瑞@tomcocobrico·
@DozenDucc @NickADobos Sorry didn’t wanna come across as an asshole here. I am sure it’s really exciting work. I wonder if this could be used scale the generation of training data instead of direct application (as it does sound quite expensive to run) Also robotics is a blind spot for me 😬
English
1
0
1
50
Hanming Ye
Hanming Ye@DozenDucc·
It sounds inelegant to burn a million tokens to fold a shirt. But isn't it cool that you can do this at all with a big next-token predictor?
English
1
0
1
38
Hanming Ye
Hanming Ye@DozenDucc·
@NickADobos What's surprising is that code-as-policy and LLM controlling robots is far more capable than expected. I expect this to unlock faster deployments and automated data collection.
English
0
0
3
234
Hanming Ye
Hanming Ye@DozenDucc·
@NickADobos I don't think so. There’s no one path to robot intelligence. Sometimes code-as-policy works best (eg. repetitive tasks in factories), sometimes robot models work best (eg. walking a humanoid, controlling 5-fingered hand). People over-index on models, but we still need em.
English
0
2
11
756
Sujeet Soni
Sujeet Soni@Sujeetsoni123·
@DozenDucc @yiding_song @theWaddleLabs Got it, in that perspective, yes it's better than turn based model but can't we hack around by forcing the frequency to much higher or using multiple async agent? Also do you think current llm understand the physics better or there's still gap or you didn't observe yet ?
English
2
0
0
51
Hanming Ye
Hanming Ye@DozenDucc·
@Sujeetsoni123 @yiding_song @theWaddleLabs LLMs definitely suck as physics. Maybe the root cause is that they suck at perception. We tried many models (every gemini, Nvidia Cosmos, you name it) and they all hallucinate an unbearable amount on videos.
English
1
0
2
53
Hanming Ye
Hanming Ye@DozenDucc·
@BrutalCaeser @yiding_song @theWaddleLabs In all the demos in the video there were no robot models (eg. ACT, VLAs). This goes to show that code-as-policy works in diverse situations. I believe that robot models will be crucial for dexterous tasks, it's just that they're not the only way towards physical intelligence.
English
1
0
3
629
Yashvardhan Gupta
Yashvardhan Gupta@BrutalCaeser·
@DozenDucc @yiding_song @theWaddleLabs Love the legibility angle @DozenDucc . But code-as-policy is only as good as its primitives, and the hard part (contact-rich dexterity) lives inside them. You still train ACT and call VLAs as tools. Isn't that moving the robot-data problem down a layer, not removing it?
English
1
0
2
695
Hanming Ye
Hanming Ye@DozenDucc·
@BrainsAndTennis The voyager paper was a substantial inspiration for our skill library design!
English
1
0
4
309
Hanming Ye
Hanming Ye@DozenDucc·
@Sujeetsoni123 @yiding_song @theWaddleLabs I thought about duplex models as a way to improve accuracy actually. Because turn-based models miss changes in the environment unless they call perception tools intentionally. So full-duplex would make it easier to catch things like wind blowing over a water bottle.
English
1
0
1
57
Sujeet Soni
Sujeet Soni@Sujeetsoni123·
@DozenDucc @yiding_song @theWaddleLabs Interesting, yes duplex vision model can reduce latency a lot. I'd have focus on accuracy first instead of latency because there are several ways to reduce latency later but accuracy will be the moat. Combination or world model and LLM can give more physics aware output I think.
English
1
0
0
54