
Alexandra Souly
28 posts

Alexandra Souly
@AlexandraSouly
Alignment Red Team at @AISecurityInst
Katılım Ağustos 2022
215 Takip Edilen302 Takipçiler

Alexandra Souly retweetledi

My first paper from @AnthropicAI! We show that the number of samples needed to backdoor an LLM stays constant as models scale.
Anthropic@AnthropicAI
New research with the UK @AISecurityInst and the @turinginst: We found that just a few malicious documents can produce vulnerabilities in an LLM—regardless of the size of the model or its training data. Data-poisoning attacks might be more practical than previously believed.
English

If you are excited to work on similar projects with a small, cracked team, apply for our open research scientist position on the Safeguards team! job-boards.eu.greenhouse.io/aisi/jobs/4390… 11/11
English

Full paper: arxiv.org/abs/2510.07192
Work done with @_robertkirk @alxndrdavies @yaringal at AISI, @javirandor + others @AnthropicAI, and Ed Chapman + others @turinginst. 10/11
English

New @AISecurityInst research with @AnthropicAI + @turinginst:
The number of samples needed to backdoor poison LLMs stays nearly CONSTANT as models scale. With 500 samples, we insert backdoors in LLMs from 600m to 13b params, even as data scaled 20x.🧵/11

English
Alexandra Souly retweetledi

We at @AISecurityInst recently did our first pre-deployment 𝗮𝗹𝗶𝗴𝗻𝗺𝗲𝗻𝘁 evaluation of @AnthropicAI's Claude Sonnet 4.5!
This was a first attempt – and we plan to work on this more! – but we still found some interesting results, and some learnings for next time 🧵

English
Alexandra Souly retweetledi
Alexandra Souly retweetledi

We at @AISecurityInst worked with @OpenAI to test & improve Agent’s safeguards prior to release. A few notes on our experience🧵 1/4

English
Alexandra Souly retweetledi
Alexandra Souly retweetledi

Defending against adversarial prompts is hard; defending against fine-tuning API attacks is much harder. In our new @AISecurityInst pre-print, we break alignment and extract harmful info using entirely benign and natural interactions during fine-tuning & inference. 😮 🧵 1/10

English
Alexandra Souly retweetledi
Alexandra Souly retweetledi

Great to see our AgentHarm benchmark mentioned here as an example evaluation of frontier AI systems!
openai.com/index/early-ac…
English




