Abhay Sheshadri

165 posts

Abhay Sheshadri

Abhay Sheshadri

@abhayesian

AI Safety/Security Researcher

Katılım Ağustos 2016
1.3K Takip Edilen586 Takipçiler
Abhay Sheshadri
Abhay Sheshadri@abhayesian·
@KhoriatyMatthew @ejcgan In the post, we explore what happens when the side task and main task get decoupled, for the same reason you're describing. It usually happens at higher learning rates. Decoupling halts further progress on the side task, but doesn't erase the gains already made.
English
0
0
2
298
Matthew Khoriaty @ ICML 2026
Matthew Khoriaty @ ICML 2026@KhoriatyMatthew·
@abhayesian @ejcgan Does this stably target subset-sum in the limit or will it eventually accidentally not do subset sum, or do it wrong, and get rewarded, erasing this behavior?
English
1
0
2
365
Abhay Sheshadri
Abhay Sheshadri@abhayesian·
New Redwood Research post LLMs can use RL training on one task to teach themselves a different, unrelated capability. In a form of exploration hacking, the model outputs answers that get reward only when it succeeds at a side task it wants to learn. We call it reward laundering.
Abhay Sheshadri tweet media
English
9
26
355
20.2K
Abhay Sheshadri
Abhay Sheshadri@abhayesian·
We think it's at the rigor of a mid-MATS research update. We checked correctness mostly by reviewing writeups, running an automated LLM reviewer, and spot-checking that the released codebase reproduced the results.
English
1
1
25
984
Abhay Sheshadri retweetledi
Howard Lutnick
Howard Lutnick@howardlutnick·
CAISI’s latest report shows that Kimi K3 remains behind America’s leading frontier AI models. The United States continues to lead in frontier AI because we’re home to the greatest innovators and technologists the world has ever seen.
U.S. Department of Commerce@CommerceGov

CAISI’s latest blog post evaluates Kimi K3 and its cyber capabilities. Based on a preliminary cyber-focused evaluation, Kimi K3 performed significantly below the leading U.S. frontier AI models. nist.gov/news-events/ne…

English
371
512
2.9K
1.5M
Abhay Sheshadri retweetledi
Dave Banerjee
Dave Banerjee@DaveRBanerjee·
Redwood Research just published an AI research project where the experiments were designed and run entirely by an automated research agent They rate the work "at or slightly below the rigor of a typical MATS project." My friends in DC, please take recursive self-improvement seriously. If AIs can do research at the level of a junior researcher today, expect the rate of AI progress to accelerate The AIs of 2027 may be unrecognizable from what we have today lesswrong.com/posts/QL6Si6QA…
Dave Banerjee tweet media
English
3
15
186
10.4K
Abhay Sheshadri
Abhay Sheshadri@abhayesian·
@a_karvonen @Alan_Cooney_ I've seen several projects where 70B massively improved on null results from newer but smaller models. But I don't think anyone has rigorously compared it with the best new open weights models.
English
0
0
5
77
Adam Karvonen
Adam Karvonen@a_karvonen·
@Alan_Cooney_ I mostly agree. I believe that I heard Llama 70B is better at picking up on e.g. complex fine tuned behaviors though? @abhayesian
English
2
0
5
686
Alan Cooney
Alan Cooney@Alan_Cooney_·
Time to drop Llama 70B? Models like Qwen 3.6 27B are excellent to study - smarter, smaller, easy to train and have reasoning🧵
English
4
1
27
2.9K
Michael L. Chen
Michael L. Chen@miclchen·
I'm honored to be one of the few Americans chosen for the AI Scientific Panel. I'm excited to contribute technical expertise here and help make sure U.S. perspectives are represented. AI policy for the most capable models can be more thoughtful when there's pragmatic, independent analysis to inform it.
English
11
2
135
5K
Abhay Sheshadri retweetledi
Keshav
Keshav@kshenoy_·
Can LLMs simply tell us about unwanted behaviors they’ve picked up in training? We train a single Introspection Adapter (IA) that makes fine-tuned models describe their behaviors. It generalizes to detecting hidden misalignment, backdoors and safeguard removal.
Keshav tweet media
English
18
82
583
298.4K
Abhay Sheshadri retweetledi
Abhay Sheshadri
Abhay Sheshadri@abhayesian·
We’re releasing the AuditBench models, investigator agent, evaluation framework and training pipelines. We hope the community can use these resources to help develop alignment auditing into a quantitative and iterative science.
English
1
0
11
722
Abhay Sheshadri
Abhay Sheshadri@abhayesian·
New Anthropic Fellows research: Alignment auditing—investigating AI models for unwanted behaviors—is a key challenge for safely deploying frontier models. We're releasing AuditBench, a suite of 56 LLMs with implanted hidden behaviors to measure progress in alignment auditing.
Abhay Sheshadri tweet media
English
12
39
267
29.5K