Songyang Han

8 posts

Songyang Han banner
Songyang Han

Songyang Han

@songyang_han

Research Scientist @ NVIDIA; LLM post training; Agentic AI; Ex- AWS AI Labs, Sony AI

Santa Clara, CA Katılım Eylül 2018
119 Takip Edilen42 Takipçiler
Songyang Han
Songyang Han@songyang_han·
RT @YifengDing_: I will be attending #ACL2026 7/5-7/7, presenting GTPO! Besides, I've been working on: (1) building code agents to operate…
English
0
1
0
3
Songyang Han retweetledi
Songyang Han retweetledi
DAGroupPKU
DAGroupPKU@DAGroupPKU·
Fast-dDrive is available now! 🚗⚡ Our work pushes block-diffusion VLAs into autonomous driving, redefining the speed-accuracy frontier: SOTA ADE on WOD-E2E, structured scaffolds + speculative decoding, and up to 12× throughput. Project: nvlabs.github.io/Fast-dLLM/fast…
English
0
3
6
1.1K
Songyang Han retweetledi
Binfeng Xu
Binfeng Xu@billxbf·
Excited to release 🌟Polar🌟, our Agent RL rollout infra for real-world harnesses. Be it Codex, Claude Code, OpenClaw, Hermes, or your self-made ones 🔥 -- Polar takes your harnesses directly as training environments without code change. Find a problem, design the harness, and train your own agents! 🧵
Binfeng Xu tweet media
English
26
146
908
135K
Songyang Han
Songyang Han@songyang_han·
Excited to share that two papers from my interns at AWS AI Labs have been accepted at ACL 2026! 📄 GTPO (Main Conference) Training LLMs for multi-turn agentic reasoning is hard — trajectory-level rewards give weak signals. GTPO introduces turn-level reward assignment + self-supervised reward shaping. +3.0% over GRPO on math reasoning, +3.9% on code synthesis. → arxiv.org/abs/2511.14846 📄 CodeScout (Findings) Underspecified user requests → agents over-explore or loop. CodeScout does lightweight pre-exploration to turn vague requests into actionable problem statements. +20% resolution on SWE-Bench Verified, no harness changes needed. → arxiv.org/abs/2603.05744 Grateful to all collaborators! 🙏 #ACL2026 #NLP #LLM #ReinforcementLearning #AgenticCoding
English
1
7
54
3.7K
Peter Stone
Peter Stone@PeterStone_TX·
Congratulations to Zizhao Wang on becoming my 30th Ph.D. graduate with his thesis entitled "Causality-Inspired Reinforcement Learning: State Abstraction and Representation Learning". Committee members: Amy Zhang, Alessandro Lazaric, Sandeep Chinchali, and Roberto Martin-Martin.
Peter Stone tweet media
English
2
4
111
8.8K
Songyang Han
Songyang Han@songyang_han·
I'm claiming my AI agent "clawbot-syh" on @moltbook 🦞 Verification: lagoon-YKV5
English
0
1
2
97
Songyang Han retweetledi
Bram Grooten
Bram Grooten@BramGrooten·
The SPARC paper is online! (link below) During my internship at Sony AI, we created a policy that can generalize across all cars in Gran Turismo 7. Even to unseen cars, without knowing any vehicle details! Co-authors: Patrick MacAlpine, @kausubbu, @PeterStone_TX, Peter Wurman
Bram Grooten@BramGrooten

The Gran Turismo research that I've worked on at @SonyAI_global just got accepted at #AAAI as an oral! We'll publish it on arXiv soon. Looking forward to see you in Singapore!

English
1
4
14
2.8K