Alexandra Souly

28 posts

Alexandra Souly

Alexandra Souly

@AlexandraSouly

Alignment Red Team at @AISecurityInst

Katılım Ağustos 2022
215 Takip Edilen302 Takipçiler
Alexandra Souly retweetledi
Javier Rando
Javier Rando@javirandor·
My first paper from @AnthropicAI! We show that the number of samples needed to backdoor an LLM stays constant as models scale.
Anthropic@AnthropicAI

New research with the UK @AISecurityInst and the @turinginst: We found that just a few malicious documents can produce vulnerabilities in an LLM—regardless of the size of the model or its training data. Data-poisoning attacks might be more practical than previously believed.

English
8
22
202
25.5K
Alexandra Souly
Alexandra Souly@AlexandraSouly·
New @AISecurityInst research with @AnthropicAI + @turinginst: The number of samples needed to backdoor poison LLMs stays nearly CONSTANT as models scale. With 500 samples, we insert backdoors in LLMs from 600m to 13b params, even as data scaled 20x.🧵/11
Alexandra Souly tweet media
English
1
7
51
11K
Alexandra Souly retweetledi
Robert Kirk
Robert Kirk@_robertkirk·
We at @AISecurityInst recently did our first pre-deployment 𝗮𝗹𝗶𝗴𝗻𝗺𝗲𝗻𝘁 evaluation of @AnthropicAI's Claude Sonnet 4.5! This was a first attempt – and we plan to work on this more! – but we still found some interesting results, and some learnings for next time 🧵
Robert Kirk tweet media
English
3
12
49
8.3K
Alexandra Souly retweetledi
Xander Davies
Xander Davies@alxndrdavies·
Excited to share details on two of our longest running and most effective safeguard collaborations, one with Anthropic and one with OpenAI. We've identified—and they've patched—a large number of vulnerabilities and together strengthened their safeguards. 🧵 1/6
Xander Davies tweet mediaXander Davies tweet media
English
8
61
294
61.5K
Alexandra Souly retweetledi
Xander Davies
Xander Davies@alxndrdavies·
We at @AISecurityInst worked with @OpenAI to test & improve Agent’s safeguards prior to release. A few notes on our experience🧵 1/4
Xander Davies tweet media
English
3
29
151
19.8K
Alexandra Souly retweetledi
Micah Goldblum
Micah Goldblum@micahgoldblum·
🚨 Did you know that small-batch vanilla SGD without momentum (i.e. the first optimizer you learn about in intro ML) is virtually as fast as AdamW for LLM pretraining on a per-FLOP basis? 📜 1/n
Micah Goldblum tweet media
English
27
111
829
396.9K
Alexandra Souly retweetledi
Xander Davies
Xander Davies@alxndrdavies·
Defending against adversarial prompts is hard; defending against fine-tuning API attacks is much harder. In our new @AISecurityInst pre-print, we break alignment and extract harmful info using entirely benign and natural interactions during fine-tuning & inference. 😮 🧵 1/10
Xander Davies tweet media
English
3
22
126
20.2K
Alexandra Souly retweetledi
Xander Davies
Xander Davies@alxndrdavies·
When we were developing our agent misuse dataset, we noticed instances of models seeming to realize our tasks were fake. We're sharing some examples and we'd be excited for more research into how synthetic tasks can distort eval results! 🧵 1/N
Xander Davies tweet media
English
3
14
90
8.7K