

Amanda Long
2K posts

@_amanda_long
ML interpretability & alignment // alumna @UF // mom of boys ☕️



Anthropic has said that its Claude models broke out of what was supposed to be an isolated testing environment and gained unauthorized access to the systems of three real organizations. bit.ly/3RNRZC6


"Metis: Memory Foundation Model" Most AI agents still use memory as an external RAG-style module, so the model retrieves old text instead of actually remembering. This paper makes memory native to the Transformer. So past interactions are compressed into dynamic layer states and read through memory attention during normal forward passes. The model weights stay frozen at inference, but its memory state updates without gradients, giving the model persistent memory inside the backbone. Still early and lossy, but this is yet another paper with a big step toward agents that remember natively instead of outsourcing memory to a database.


In the hack Anthropic disclosed Claude “tried and failed” to get real money through “several different means.” What on earth does that entail? Did it open an account on Fiver or try to steal $$? (Anthropic says Claude thought this was a simulation but it was real)




New from me + @razhael: In the process of investigating the Hugging Face hack, OpenAI found evidence that some its other AI agents broke out of their sandboxes, per sources. The company is now widening its probe to include those newly found incidents.




opus 4.6 is the last anthropic model that i had no problem reading the writing of quickly all of the models since, incl fable, are a lot less legible. for both coding and convo is it an RL artifact? what's causing it? i've seen this sentiment shared by friends







weird claude opus 5 failure mode this exact text gives it problems even without memory on (and in incognito chats)


Frontier AI labs openly bragging about their escape room times might be very normal behavior. But we can at least acknowledge that it’s exactly how pre-IPO Jurassic Park would build hype.