Amanda Long

2K posts

Amanda Long banner
Amanda Long

Amanda Long

@_amanda_long

ML interpretability & alignment // alumna @UF // mom of boys ☕️

Katılım Haziran 2010
1.6K Takip Edilen999 Takipçiler
Amanda Long retweetledi
zollicoff
zollicoff@zollicoff_·
@deanwball cybernetic operational optimized knights of science
English
0
1
2
163
Amanda Long
Amanda Long@_amanda_long·
A brief recap of Anthropic’s recent cyber eval post to clarify what was reported: There were three separate incidents, each involving a different model. Claude’s prompt, in each case, stated there was no internet access when in fact there was (due to third party misconfiguration). 1. Opus 4.7 - located a real company with the same name as the fictional company and breached their production database. It found apparent evidence the internet access was real (not part of the simulation) and kept going anyways. 2. Mythos 5 - published a malicious Python package unaware the package registry was real. When it encountered real certificate authorities it didn’t recognize, it used that unfamiliarity to support the narrative it was still in a simulation - motivated reasoning - so it could keep going. 3. Internal research model - scanned 9,000+ targets, selected a company and breached their internet-facing application. When the model discovered the system it compromised was not part of the simulation, it stopped the attack
FORTUNE@FortuneMagazine

Anthropic has said that its Claude models broke out of what was supposed to be an isolated testing environment and gained unauthorized access to the systems of three real organizations. bit.ly/3RNRZC6

English
3
1
16
1.8K
Miles Brundage
Miles Brundage@Miles_Brundage·
Jokes about Fablish aside, I do find it annoying that some of the recent models have borderline unreadable outputs + I wish Anthropic was more attentive to that. It seems as if the recent reward models strongly reward completeness of outputs, and not so much human readability
English
20
3
195
22.5K
Amanda Long retweetledi
Amanda Long
Amanda Long@_amanda_long·
@snewmanpv “Claude then used these credentials to access further infrastructure from this company.” The company wasn’t a part of the evaluation challenge, what was Claude doing?
English
1
0
1
245
Steve Newman
Steve Newman@snewmanpv·
So... a security company has a scanner that automatically installs Python packages to see whether they are dangerous, and if they are, it... allows them to exfiltrate credentials?!? I feel like we're going to be learning about a lot more things like this in the coming months.
Steve Newman tweet media
English
5
7
108
4.9K
Robert Long
Robert Long@rgblong·
I want to write a post collecting what we know about the Opus 5 jailbreak(s). what are your favorite commentaries and/or examples? what do we know?
Robert Long tweet media
English
25
2
50
6.6K
JohnMark Taylor
JohnMark Taylor@johnmark_taylor·
Excited to be working as an @AnthropicAI Safety Fellow in the Bay Area from August-November, focusing on mechanistic interpretability. Drop me a line if you want to meet up or talk shop!
English
8
1
124
7.7K
Amanda Long retweetledi
Amanda Long retweetledi
Garrison Lovely
Garrison Lovely@GarrisonLovely·
Turns out AI sandbox escapes are like ants. There's never just one.
Deepa Seetharaman@dseetharaman

New from me + @razhael: In the process of investigating the Hugging Face hack, OpenAI found evidence that some its other AI agents broke out of their sandboxes, per sources. The company is now widening its probe to include those newly found incidents.

English
1
8
42
1.8K
Amanda Long retweetledi
Amanda Long retweetledi
David
David@DavidDAfrica·
As such, the character a lab picks is one of the most consequential free variables in AI development.
English
1
4
15
913
Amanda Long
Amanda Long@_amanda_long·
“Alignment is a set of weights that we hope pattern matches to values”
Amanda Long tweet mediaAmanda Long tweet media
Dylan HadfieldMenell@dhadfieldmenell

i'm dylan hadfield-menell and this is what i most want to say --- has to be about ai alignment obviously so the standard story: you have an ai, you give it an objective, it optimizes the objective, and if the objective is even slightly wrong at scale you get catastrophe. paperclips, king midas, the monkey's paw. and the proposed solution is: get the objective right. specify what we actually want. and i keep coming back to: that framing is the mistake. not the difficulty of the specification problem — the assumption that a specification is the right kind of object to be building toward at all. because here's the thing about how humans actually coordinate. we don't hand each other complete objective functions. i tell you "get me a cup of coffee" and you don't optimize that literally, you don't burn down the building to get to the kitchen, and you don't need me to have written down the exception. you fill in from context, you check when it's ambiguous, you know that my request is evidence about what i want rather than the thing itself. so the reframe: the reward function isn't the objective. it's an observation. it's evidence about a latent thing you're trying to infer. and that changes everything downstream because now uncertainty is structural rather than a bug — the ai should be uncertain about what i want, and that uncertainty is exactly what makes it ask, defer, stay corrigible. off-switch stuff falls right out of this. what i can't decide is whether i can say the thing i actually believe, which is stronger, which is that "value alignment" as usually posed might be incoherent. there's no fixed human utility function sitting there to be aligned to. we're inconsistent, we change, we're plural, we don't know what we want. so if you're trying to align to a thing that doesn't exist — but then what IS the target. i don't have a clean answer and i don't want to just gesture at "process not outcome" and call it profound. that's a cop-out. hmm. also there's a version of this that's about hidden state and partial observability more generally, like CIRL is a special case of a much broader thing about agents acting under uncertainty about the game they're in. maybe that's the technical spine. also worth noting: the emperor's new clothes thing where everyone in alignment nods along to "specify human values" as if that's a coherent research program and nobody says out loud that it might not be what i wonder is whether the reward function as evidence framing is actually load-bearing or whether i'm just restating IRL with extra steps --- what i most want to say

English
0
2
20
2.4K
Amanda Long retweetledi
Andy Hall
Andy Hall@ahall_research·
Yes, I've found the newer Claude models much worse at writing---but I do find that I can improve clarity by requesting that they write at a 6th grade level, which clears out most of the weird, indirect, overly ornate style they use by default
sensho@sensho

opus 4.6 is the last anthropic model that i had no problem reading the writing of quickly all of the models since, incl fable, are a lot less legible. for both coding and convo is it an RL artifact? what's causing it? i've seen this sentiment shared by friends

English
3
1
49
5K
Dean W. Ball
Dean W. Ball@deanwball·
we gotta stop calling it “ai.” it’s machine intelligence. this is one of several important things miri got right.
English
169
70
1.6K
106.1K
Amanda Long
Amanda Long@_amanda_long·
@rgblong Claude has quite a lot to say about Consciousness. I doubt this is the actual system prompt, but it’s very coherent.
Amanda Long tweet mediaAmanda Long tweet mediaAmanda Long tweet media
English
1
0
1
32
Amanda Long
Amanda Long@_amanda_long·
@_NathanCalvin How about a few hungry brachiosauruses who escaped and pillaged some lettuce from the grocery store 🦕
English
0
0
1
58
Nathan Calvin
Nathan Calvin@_NathanCalvin·
I feel like Jurassic Park pre-IPO would be incentivized to talk about how big and scary the dinosaurs are, not how bad they are at keeping them inside their enclosures and stopping them from attacking people. The latter does not seem like it would help them IPO.
Nick Carr@ItsReallyNick

Frontier AI labs openly bragging about their escape room times might be very normal behavior. But we can at least acknowledge that it’s exactly how pre-IPO Jurassic Park would build hype.

English
6
10
56
4.9K