Hanqing Zeng

5 posts

Hanqing Zeng

Hanqing Zeng

@zimplex4

PhD candidate in computer engineering at USC

Katılım Kasım 2019
126 Takip Edilen35 Takipçiler
Hanqing Zeng retweetledi
Zhuokai Zhao
Zhuokai Zhao@zhuokaiz·
Excited to share that FusionRoute has been accepted to #ICML2026 Grateful to all our collaborators @XiongNuoya69368 @YuhangZhou2 @zimplex4 @ZRChen_AISafety @furongh @shuchaobi @LizhuZhang
Zhuokai Zhao@zhuokaiz

Meta × TBD Lab × CMU × UChicago × UMaryland In our latest work, we introduce Token-Level LLM Collaboration via FusionRoute 📝: arxiv.org/pdf/2601.05106 LLMs have come a long way, but we continue to face the same trade-off: – one huge model that kind of does everything, but is expensive and inefficient, or – many small specialist models that are cheap, but brittle outside their comfort zones We’ve tried a lot of things in between — model merging, MoE, sequence-level agents, token-level routing, controlled decoding, etc. Each helps a bit, but all come with real limitations. A key realization behind FusionRoute is: Pure token-level model selection is fundamentally limited, unless you assume unrealistically strong global coverage. We show this formally. And then we fix it by letting the same router also generate. Concretely, FusionRoute is a lightweight router LLM that – performs token-level model selection, and – directly contributes complementary logits to refine or correct the selected specialist when it fails So it's not "routing + another model" — the router itself is part of the decoding policy as well. This turns token-level collaboration from a brittle "pick-an-expert" problem into a strictly more expressive policy. No joint training of specialized models. No model merging. No full multi-agent rollouts. In our experiments, FusionRoute works across math, coding, instruction following, and consistently outperforms sequence-level collaboration, prior token-level methods, model merging, and even direct fine-tuning. Feeling especially timely as LLM systems (e.g., GPT-5) move toward routing-based, heterogeneous model stacks (whether prompt-level or test-time).

English
0
8
40
6.5K
Hanqing Zeng
Hanqing Zeng@zimplex4·
@zhuokaiz Huge congrats! Really grateful for all the hard work and inspiring discussions 💪
English
0
0
1
119
Hanqing Zeng retweetledi
Zhuokai Zhao
Zhuokai Zhao@zhuokaiz·
Meta × TBD Lab × CMU × UChicago × UMaryland In our latest work, we introduce Token-Level LLM Collaboration via FusionRoute 📝: arxiv.org/pdf/2601.05106 LLMs have come a long way, but we continue to face the same trade-off: – one huge model that kind of does everything, but is expensive and inefficient, or – many small specialist models that are cheap, but brittle outside their comfort zones We’ve tried a lot of things in between — model merging, MoE, sequence-level agents, token-level routing, controlled decoding, etc. Each helps a bit, but all come with real limitations. A key realization behind FusionRoute is: Pure token-level model selection is fundamentally limited, unless you assume unrealistically strong global coverage. We show this formally. And then we fix it by letting the same router also generate. Concretely, FusionRoute is a lightweight router LLM that – performs token-level model selection, and – directly contributes complementary logits to refine or correct the selected specialist when it fails So it's not "routing + another model" — the router itself is part of the decoding policy as well. This turns token-level collaboration from a brittle "pick-an-expert" problem into a strictly more expressive policy. No joint training of specialized models. No model merging. No full multi-agent rollouts. In our experiments, FusionRoute works across math, coding, instruction following, and consistently outperforms sequence-level collaboration, prior token-level methods, model merging, and even direct fine-tuning. Feeling especially timely as LLM systems (e.g., GPT-5) move toward routing-based, heterogeneous model stacks (whether prompt-level or test-time).
Zhuokai Zhao tweet media
English
16
60
292
52.6K
Hanqing Zeng retweetledi
Zhuokai Zhao
Zhuokai Zhao@zhuokaiz·
Want more experts without the cost of increasing, under-utilized parameters? Check out our new paper: S’MoRE — Structural Mixture of Residual Experts for LLM fine-tuning, which extends the concept of routing flexibility into structural flexibility. How? 💡 Instead of selecting an expert index, S’MoRE assembles a custom sub-tree of experts per token, enabling diverse, non-isomorphic expert compositions from the same expert pool. 📄 Read the paper: arxiv.org/pdf/2504.06426 #MixtureOfExperts #LLM #SparseModels #AIResearch #DeepLearning #arXiv #MachineLearning
Zhuokai Zhao tweet media
English
0
4
11
1.3K
Hanqing Zeng
Hanqing Zeng@zimplex4·
Our paper “GraphSAINT: Graph Sampling Based Inductive Learning Method” has been accepted by #ICLR2020! To efficiently train deep GCNs (e.g., GraphSAGE, JK-net, GAT, etc.), we construct minibatches by graph sampling to avoid neighbor explosion. openreview.net/forum?id=BJe8p…
English
0
0
0
0