Eeeenold Trusk
525 posts




Introducing Sendblue Verify 🟦🎉 Sign in with iMessage. So your users can auth with a single message. Try it now for free. link in first comment



In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. anthropic.com/news/investiga…



There’s been a lot of speculation about where we stand on open-weights models. We’ve outlined our views in full here: anthropic.com/news/position-…

At the core of our mission is working through how to ensure increasingly powerful AI benefits everyone. We believe that, at some point in the future, AI acceleration for frontier model development may be so high that the world will need to pace the rate of AI advancement. We hope to contribute to work led by the U.S. government, alongside other labs and the open-source community, to develop the tools and mechanisms that could make that possible. pacingthefrontier.com


I think open source is <12 months behind the frontier, and the same security reaction you saw with Mythos/Fable will likely happen with open source sometime in the next year. To my mind, the escape routes are (1) Frontier models desensitize the security apparatus sufficiently that open source hitting the same levels doesn't spook folks. Could be by genuinely fixing a lot of the worst security issues a la Glasswing, could be by just folks familiarizing themselves and deciding they're not as worried as they were. (2) Convince decisionmakers that industry benefits to open source exceed the security costs. I think this is an easy case to make now but will be a much harder one after the first effective open-model-driven cyberattacks. (3) A research program that makes open source models as safe as closed models. This'd mean some combination of (a) Suppressing the concerning capabilities during training (b) Sensitizing the weights somehow so that adding those capabilities later is hard (c) Developing assurances a model doesn't contain backdoors (3) seems hard but a lot of the things we've accomplished the last few years have seemed hard, and various groups have made inroads on all components of the above. Picking some examples: (a) anthropic.com/research/off-s… (b) arxiv.org/abs/2408.00761 (c) microsoft.com/en-us/security…


SITUATION UPDATE: The letter has been published. 1,122 frontier AI lab employees have signed so far.

Elon Musk is the biggest asshole in our Solar System

pov: your company sponsored a yacht gala to Alcatraz with 150 hot ppl the most popular side event at YC sus 🛳️💃


There’s been a lot of speculation about where we stand on open-weights models. We’ve outlined our views in full here: anthropic.com/news/position-…


We support this petition, signed by our CEO, several co-founders, and senior staff. Our own research on recursive self-improvement, published last month, points to the need for tools to deliberately pace the frontier of AI development so society can prepare. We’re glad to see broad agreement across the field. pacingthefrontier.com












