Vinay Samuel

23 posts

Vinay Samuel

Vinay Samuel

@vsamuel2003

ML Research Intern @ Scale AI| MS in Computer Science @ UMD College Park| BS in Stat/ML @ CMU |

Katılım Temmuz 2024
28 Takip Edilen31 Takipçiler
Sabitlenmiş Tweet
Vinay Samuel
Vinay Samuel@vsamuel2003·
Post-training makes LLMs safer and better at following instructions, but less diverse. 🤔 Can we get that diversity back without sacrificing alignment? Introducing ReDiPO: a preference optimization recipe for restoring distributional diversity while preserving safety and instruction-following.
Vinay Samuel tweet media
English
1
12
33
11K
Vinay Samuel retweetledi
rishanth rajendhran
rishanth rajendhran@rishanthrajendh·
Can open-weight language models write long stories? Not well: either their writing quality degrades at longer lengths or they severely undershoot in length as the requested length grows, if not both. Frontier models fare better. This begs the question, can we train open-weight language models to write long stories with limited data and compute?
rishanth rajendhran tweet media
English
2
8
38
12.5K
Vinay Samuel
Vinay Samuel@vsamuel2003·
Takeaway: alignment should not mean model collapse for distributional diversity. Base models can be reservoirs of valid alternatives that instruction tuning pushes down. ReDiPO shows how careful preference data can bring those modes back into instruct models. 🌱
English
1
0
1
132
Vinay Samuel
Vinay Samuel@vsamuel2003·
Across Qwen3-4B, OLMo-3-7B, and LLaMA-3.1-8B, ReDiPO boosts NoveltyBench distinct_k over instruct checkpoints by: 📈 +134% 📈 +33% 📈 +44% These novelty gains do not compromise instruction following or safety alignment on all tested models.🛡️
Vinay Samuel tweet media
English
1
1
2
199
Vinay Samuel
Vinay Samuel@vsamuel2003·
Post-training usually optimizes one answer at a time: helpful, safe, well-formatted. But for brainstorming, creative writing, and hypothesis generation, users often need many distinct valid answers, not multiple samples of the same idea. ReDiPO uses the base model as a source of missing alternatives. For each prompt, it samples base + instruct responses, rewrites base outputs in instruct style, filters unsafe/low-quality candidates, then forms DPO pairs where quality is matched but marginal diversity differs. 🔁
Vinay Samuel tweet media
English
1
0
2
218
Vinay Samuel
Vinay Samuel@vsamuel2003·
Post-training makes LLMs safer and better at following instructions, but less diverse. 🤔 Can we get that diversity back without sacrificing alignment? Introducing ReDiPO: a preference optimization recipe for restoring distributional diversity while preserving safety and instruction-following.
Vinay Samuel tweet media
English
1
12
33
11K
Vinay Samuel
Vinay Samuel@vsamuel2003·
Ablations show that marginal-diversity pair selection drives the diversity recovery, base-response rewriting makes candidates usable, and quality/safety filters help preserve alignment. 🧪
Vinay Samuel tweet media
English
0
0
1
23
Vinay Samuel
Vinay Samuel@vsamuel2003·
Compared with DDPO, a DPO baseline for creative writing, RiDiPO gives up some raw novelty but keeps the model much more aligned: ✅ MTBench: 5.29 vs 7.38 ✅ IFEval: 0.566 vs 0.817 ✅ HarmBench ASR: 0.296 vs 0.075 ✅ Arena-Hard: 0.198 vs 0.287
Vinay Samuel tweet media
English
1
0
1
30
Vinay Samuel retweetledi
Connor Dilgren
Connor Dilgren@ConnorDilgren·
Excited to announce my first preprint in LM interpretability! Latent reasoning models are not monitorable by default, since they don't reason in human-readable, natural language text. But can we make progress in understanding their intermediate reasoning steps using mech interp?
Connor Dilgren tweet media
English
7
30
208
18.1K
Vinay Samuel retweetledi
Yijia Shao
Yijia Shao@EchoShao8899·
LM agents today primarily aim to automate tasks. Can we turn them into collaborative teammates? Introducing Collaborative Gym (Co-Gym), a framework for enabling & evaluating human-agent collaboration! I now get used to agents proactively seeking confirmation or my deep thinking.
English
17
80
183
72.2K
elvis
elvis@omarsar0·
Evaluating Persona Agents and LLMs It's great to see research around evaluating persona agents. Persona agents is probably going to become one of the most common and useful ways to use LLMs but there is very little research on how to properly assess them. This work proposes a benchmark to evaluate persona agent capabilities in LLMs. Finds that Claude 3.5 Sonnet only has a 2.97% relative improvement in PersonaScore compared to GPT 3.5 despite being a much more advanced model. Custom agents that act as specific personas makes sense in domains like education, healthcare, creativity, productivity, entertainment, and more. The interesting thing is that current models are explicitly trained to operate this way but a lot of people are interested in using them like that.
elvis tweet media
English
5
75
348
27.1K
Vinay Samuel
Vinay Samuel@vsamuel2003·
We evaluate 6 LLMs on 200 personas and 10K questions, noting: - Significant potential for improvement (e.g., Claude 3.5 Sonnet's 2.97% PersonaScore gain over GPT-3.5). - Larger, more complex models don't necessarily enhance persona agent capabilities.
Vinay Samuel tweet media
English
1
0
1
139
Vinay Samuel
Vinay Samuel@vsamuel2003·
PersonaGym – the first platform to evaluate persona agents. Persona agents are critical to magical AI experiences, but how do we evaluate them? PersonaGym uses dynamic LLM-powered environments for faithful and targeted agent eval. personagym.com
Vinay Samuel tweet media
English
1
4
8
1.9K