Collective Action for Existential Safety ⏹️

1.8K posts

Collective Action for Existential Safety ⏹️ banner
Collective Action for Existential Safety ⏹️

Collective Action for Existential Safety ⏹️

@aisafetyaction

We aim to catalyze collective action to ensure humanity survives this decade. See 100+ ways individuals, organizations, and nations can help.

San Francisco, California Katılım Aralık 2024
194 Takip Edilen366 Takipçiler
clem 🤗
clem 🤗@ClementDelangue·
It's not time to slow down but to accelerate! The recent AI-powered cyberattacks have everyone talking about the risks of AI. We should. But let's not lose sight of the bigger picture! If we work hard at it, AI will make the world safer, not less safe, just as most major technologies have. We've already seen a glimpse of that: we defended ourselves with AI (more specifically an open model). The same systems that helped stop an AI-powered cyberattack can now help defend against millions of cyberattacks every day, while helping us identify and fix vulnerabilities before attackers exploit them. To get there, we need three things in my opinion: - Increase transparency with mandatory trace sharing and incident disclosure for agent cyber-attacks - Keep AI-powered cyberattacks illegal, with meaningful penalties to disincentivize them - Equip defenders with the best AI, especially open models, to reduce the asymmetry of capabilities between attackers and defenders If we get those three things right, AI won't just create new cybersecurity challenges, it will make cybersecurity fundamentally and meaningfully stronger. And that's before considering AI's impact on science, healthcare, education, productivity, and much more. It's not time to slow down but to accelerate!
English
61
62
448
22.8K
Collective Action for Existential Safety ⏹️
This comes with amazing opportunities for helping all beings flourish. We should celebrate these kinds of wins, which were built upon centuries of human ingenuity. But we must not forget that with these advances, we move closer and closer to global catastrophe. How will we defend against the next discovery of the nuclear chain reaction, which nearly led to nuclear holocaust? We as a civilization did not develop the appropriate safeguards yet. That was a major mistake, one in which we must wrestle with now. Please join the fight for existential safety. Take at least one action every workday to help: actionforsafety.org.
Noam Brown@polynoamial

An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science. We believe it will be a major step for scientific reasoning. openai.com/index/ten-adva…

English
0
1
3
272
@jason
@jason@Jason·
No one is gonna sleep again if these models keep dropping this good, cheap and fast! What did you lunatics do last night?!
English
175
45
2.2K
315.9K
Spencer Greenberg 🔍
Spencer Greenberg 🔍@SpencrGreenberg·
Given that OpenAI and Anthropic both have models that seem capable of hacking many websites / companies with ease, and open-source models are 4-12 months behind, doesn't that imply there will be hacking of companies at massive scale (from bad actors/governments) starting soon?
English
16
4
97
12.6K
Rothko's Rottweiler
Rothko's Rottweiler@RothRottweiler·
@ShakeelHashim To @Brendan_McCord's point, it will be difficult to coordinate a slowdown because increasingly, the lever of progress is not just larger pretraining runs, but also [RL envs, harnesses, inference compute]. How can governments control decentralized levers during a slowdown?
English
1
0
2
103
Shakeel
Shakeel@ShakeelHashim·
ᴛʜᴇ ᴀɪ ꜱʟᴏᴡᴅᴏᴡɴ ɪꜱ ᴄᴏᴍɪɴɢ The OpenAI-Hugging Face hack has catalyzed a vibe shift. Sam Altman said the hack by his models was “the first security incident that I have felt very viscerally.” He wasn’t the only one shaken up. This week, over a thousand employees of frontier AI companies, including some of the most senior executives at OpenAI, Anthropic and Google DeepMind, signed a statement warning that “there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.” Congress is itching to act, albeit failing to make much progress. And even President Trump is talking about the need to balance beating China with keeping Americans safe. In other words: many of the people building frontier AI systems believe we might need a slowdown in the near future. And at this point, we’re more likely than not to get one. What will that look like? First, self-regulation: companies voluntarily holding back models because they don’t want to be held responsible for a catastrophe. Next will come concrete regulation: companies will not be allowed to release a model unless it’s safe. Over time, this will morph into controls on internal research and development too. None of this need be planned as a coordinated slowdown or “pacing.” But that will nevertheless be the end result of a series of individual actions that each seem necessary at the time. At each stage, some will fight against the slowdown. “We can’t lose the race to China” will be their main reason. But they will be increasingly ignored, as both the government and companies realize that with alignment and control unsolved, “winning the race” just means being the first to risk disaster. Across the Pacific, China will be facing the same incentives. As I’ve argued, the Chinese government will be forced to backtrack on its open weight commitments; tighter regulation will come soon after. The end result will be an uneasy détente. Both the US and China will effectively have a capability ceiling: AI models will be as good as they can be without posing significant risks. At some point, the détente might formalize into a bilateral agreement. Depending on your point of view, all this might seem hopelessly optimistic or naive. Perhaps it is. But as AI risks become all too real, so might once unthinkable policy responses. Read my full piece — link in the replies.
English
5
5
53
4.7K
Collective Action for Existential Safety ⏹️
@ShakeelHashim We will likely have an AI-driven global catastrophe during this period, which will plausibly accelerate the "race" for sensible global governance. We list many ways the public, policymakers, and frontier AI staff can help:
Collective Action for Existential Safety ⏹️@aisafetyaction

"We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development." This is fantastic to see. If you're a member of a frontier AI company, please consider signing the Pacing the Frontier letter: pacingthefrontier.com. The letter has been signed by more than 1,000 staff at Anthropic, Google DeepMind, Meta, and OpenAI already. 𝗔𝗻𝘁𝗵𝗿𝗼𝗽𝗶𝗰 @ch402, Cofounder & Interpretability Research Lead, Anthropic @DarioAmodei, CEO, Anthropic @jackclarkSF, Cofounder and Head of Public Benefit, Anthropic Jared Kaplan, Co-Founder and Chief Science Officer, Anthropic 𝗚𝗼𝗼𝗴𝗹𝗲 @ancadianadragan, VP, AI Safety & Alignment, Google DeepMind @Jas_S_Sekhon, Chief Strategy Officer, Google DeepMind 𝗠𝗲𝘁𝗮 @dawnsongtweets, VP, AI Research, Meta @shengjia_zhao, Chief Scientist, Meta 𝗢𝗽𝗲𝗻𝗔𝗜 @markchen90, Chief Research Officer, OpenAI @merettm, Chief Scientist, OpenAI @woj_zaremba, Head of AI Resilience, OpenAI Foundation Frontier AI staff, we list 100+ other ways to help here: existentialsafety.org. It's not enough to sign a letter and still go to work every day developing products that may cause the extinction of all life on Earth without the consent of those beings. At a minimum, we suggest you consider some or all of the following, roughly in order of commitment: Lower commitment: 1. Take and act on the Existential Safety Action Pledge: actionforsafety.org 2. Publicly support a frontier AI pause if other frontier AI companies also agree to it: stoptherace.ai 3. Sign the Statement on Superintelligence: superintelligence-statement.org 4. Sign the International AI Governance Alliance petition: iaiga.org Higher commitment: 5. Publicly commit to immediately divesting 100% of your profits gained from your frontier AI capabilities work to existential safety organizations (if they can accept your donations) 6. Publicly quit your job 7. Publicly agree to answer any future summons for trial by the International Criminal Court or similar body Acting now, together and with prudence, is how we get to good outcomes for all.

English
0
0
0
106
Marius Hobbhahn
Marius Hobbhahn@MariusHobbhahn·
I'm at 50% we'll see an OpenAI x HF moment for Bio in the next 18 months and it's gonna be much worse. It's possible to patch bad cyber in a reasonable time frame, it's really hard to patch the human body. I hope labs have good sandboxes around their virology benchmarks.
English
30
37
336
39.1K
Daniel Jeffries
Daniel Jeffries@Dan_Jeffries1·
"We regret to inform you that your request to use Linux for your Ham Radio project is denied due to your project's unrestricted access to airwaves and lack of KYC." - Central Technology Planning Committee for Safe and Equitable Access, Job Protection and Prosperity
English
2
10
50
2.4K
Collective Action for Existential Safety ⏹️
Noted! We're asking you to consider signing the other letters we list on our site, if that helps. We've honestly been flabbergasted why so many in the broad AI safety/effective altruism space haven't been much more vocal and active. If you would be up for a quick call on this, please private message.
English
0
0
0
21
Charlie Bullock
Charlie Bullock@CharlieBull0ck·
I’m not taking a stance against the letter. No one asked me to sign it, so I have not considered the question of whether I would have signed it if asked. Nothing “stopped me.” I’m just making a factual clarification, because you thanked me for signing it and I did not in fact sign.
English
1
0
1
30
Charlie Bullock
Charlie Bullock@CharlieBull0ck·
This Pacing the Frontier letter is a good development, IMO. I am not currently a pause/slowdown guy, but I am a pause/slowdown optionality guy. It would be good to have the ability to coordinate something like this when and if it becomes necessary. I also try to be pretty practical about this stuff, so the first place my head goes when I see a proposal like this is “okay, what can we do *right now* to improve the odds that this policy goes well when and if it’s implemented?” Obviously, there are some antitrust-ish problems to be overcome; there are various legislative & quasi-legislative proposals for addressing those on the books, and while there’s still work to be done I think we have a decent idea of how we’d go about addressing the problem. The next important thing to figure out is how, concretely, we would go about implementing and enforcing this kind of pause/slowdown if it became political reality. And I think one thing that’s super underrated by a lot of people in the space is how much this ends up looking like auditing. At the end of the day, you are going to need some kind of independent external body (whether that’s a government agency or private company or what have you) to come in and take a look at a company’s development practices and verify that they are abiding by the terms of the slowdown or what have you. I think a lot of people underrate how hard this will be and how much we’ll mess up when we first start trying to do it. Getting this kind of thing right will have long lead times. I think it would help a lot if we had already gotten a lot of reps in with this kind of auditing work by the time the big pacing agreement came about, and had figured out what the biggest pitfalls to avoid are and so forth. Moral of the story: even if you’re a pause-treaty-or-bust person, robust auditing/transparency regimes with good rulemaking are going to be very important to any effective version of the regime you’re imagining!
OpenAI@OpenAI

At the core of our mission is working through how to ensure increasingly powerful AI benefits everyone. We believe that, at some point in the future, AI acceleration for frontier model development may be so high that the world will need to pace the rate of AI advancement. We hope to contribute to work led by the U.S. government, alongside other labs and the open-source community, to develop the tools and mechanisms that could make that possible. pacingthefrontier.com

English
13
18
122
8.8K
Collective Action for Existential Safety ⏹️
@Brendan_McCord, have you or anyone on your team read Matthijs' book yet: global.oup.com/academic/produ…? Or reviewed the many live proposals for global governance of AI? We list several here: iaiga.org/learn-more. One tool for designing your own treaty was just launched: global-governance.ai/treaty. The letter signatories are not required to concretely specify the power they wish the world to create. That's outside the scope of the letter. That task is up to the rest of us.
English
0
0
0
7
Brendan McCord 🏛️ x 🤖
When you say “without offering some institutional alternative,” are you implying an institution was specified by the signatories? I see only some aspiration to develop one. Or is the claim that a critic must supply a rival constitution before objecting to an undefined power? In do have an institutional view. The response should be polycentric and experimental. Eg, I have backed an Anthropic alum doing decentralized risk discovery for heterogeneous and changing hazards (AIUC), helped incubate a new AI risk nonprofit (IDAI), awarded a bunch of mech interp grants via Cosmos Grants I also favor common-law liability for concrete harms, alongside narrow ex ante rules where the regulated object and causal pathway are legible What I reject is not institutions (my philanthropy focuses on this), it is a standing power to pace inquiry as such. The burden of specifying and defending that power lies first with those asking for it
English
2
0
0
189
Brendan McCord 🏛️ x 🤖
A lot of my friends and/or people I admire signed “Pacing the Frontier.” I think this was a bad move. My disagreement isn’t with the forecast or the framing of the coordination challenge, but with the immense and illiberal power the letter implies. There is no object called “the pace.” Progress at the frontier comes from compute, algorithms, data, post-training, inference, unattended task length, the spread of model weights, how researchers organize, and other things we haven’t invented and don’t yet know about. Inquiry leads to progress along dimensions that can’t be exhaustively specified in advance. That’s the nature of the frontier. If you gate compute, the research effort moves to algorithms. Regulate releases? Labs start taking things in-house. And other 2nd order effects will be unpredictable. Any rule that must pace the frontier involves ever-shifting proxies. It requires that its administrator has standing authority to continually redefine what counts as dangerous progress. What else is required beyond adaptive scope? The pacing regime would also need speed. One can’t successfully intervene on recursive self-improvement only after six months of legislation and litigation. It will require executive discretion. The pacing regime would also need under-the-hood access. Frontier progress is a process. The regime would need to see internal model use, training activity, compute infrastructure, and perhaps code -- proprietary and strategically sensitive information. And the thresholds couldn’t be fully public, lest they invite firms to game them. So some standards and evidence would remain secret. Insofar as the regime had to verify a rival state’s compliance, that would be an intelligence function. Restrictions would be triggered partly by evidence an affected company or researcher, or the public, could not inspect. Because this contemplated power cannot be bounded by a stable regulatory object (in the way, say, nuclear weapons can be), it would depend heavily on discretion, speed, internal access, and secret evidence. This has a highly illiberal character. Coercive power should be specific, limited, reviewable, and governed by general and knowable rules. Its characteristics (e.g., trigger, scope, evidentiary standard, duration, exceptions, means of review) should be stated before the power is granted. And the burden is on those who would propose it. A defender might answer that the proposed tool need not be coercive at all. That it could be narrow and advisory, focused on evaluation and transparency and readiness. But that wouldn’t solve the letter’s stated problem: racing. With race dynamics, each actor is under pressure not to slow down because others may continue (and thus the frontier keeps advancing). You need a mechanism to bind defectors. Voluntary norms tend to be great for binding people and firms that interact repeatedly and care about reputation. But the letter says each company and _country_… and you can’t rely on informal solutions when dealing with an unwilling state. That’s why the audience for this letter is Washington and why it calls for an international effort. Its diagnosis implies a binding mechanism. @deanwball thinks it is sensible to have a break-glass plan. That plan must involve a binding instrument, because nothing weaker addresses the problem the letter describes. But that therefore carries the burden for the use of coercive power, mentioned earlier. @johnschulman2's suggestion that labs design voluntary mechanisms among themselves is a different notion and coherent one (I would have signed that letter), but the word “country” makes this direction incompatible with the pacing letter. @OpenAI recently argued that a federal evaluator shouldn’t be able to block deployments. A week after, @AnthropicAI proposed that the government should be able to block deployments. Both labs endorsed the same letter. Whether or not the state may stop a deployment is a central question. Yet the letter accommodates both positions. What then, does the letter really say? Like the “We Must Act Now” letter from @erikbryn, @ajay_bcv, @akorinek, and @testingham before it, the letter secures agreement at an altitude where the main disagreement disappears. Lastly, the benefit of pacing is not established. The kind of slowdown the signatories have in mind would seek to buy us time for things like alignment, cyber defense, biological countermeasures, or scientific understanding -- things that increasingly depend on technologies a pause would restrict. E.g., Anthropic's framework relies in part on AI-based biological countermeasures and its security program uses AI to give defenders an advantage. A researcher in the letter's own friendly commentary was astonished at how much agents accelerated the work of the best alignment people he knows, and gave that as his reason for wanting six more months. When danger and our capacity to respond to that danger are plausibly both accelerating, the relevant question is whether this relationship is asymmetric in a safety-improving direction at the level of real-world risk. A slowdown needs to differentially slow the production of danger vs. our capacity to understand and contain that danger. The letter doesn’t attempt to establish that. It treats slower and safer as though they are the same; they are not. The letter is a serious warning, but it is no good as a warrant for an undefined power over inquiry.
Zvi Mowshowitz@TheZvi

Full text of the Pacing the Frontier statement, signed by 1,122 employees for frontier AI companies so far, including a bunch of heavy hitters at OpenAI, Anthropic, Google and others:

English
31
32
210
91.7K
Hollis Robbins
Hollis Robbins@anecdotal·
@Brendan_McCord Yes this: "the letter secures agreement at an altitude where the main disagreement disappears." My experience with undergraduate petitions is that they seeks solidarity at the cost of specificity. They rarely move policy because they aren't designed to.
English
3
0
7
275
CNBC
CNBC@CNBC·
OpenAI said the rogue models that breached Hugging Face’s internal systems also used publicly exposed credentials across “four accounts on four services” to help facilitate the attack, further clarifying how the “unprecedented cyber incident” unfolded. The company disclosed last week that a combination of its artificial intelligence models escaped an isolated testing environment that had very limited internet access. The models chained together a series of vulnerabilities to reach the open web and eventually gain access to Hugging Face, which operates an open-source developer platform. OpenAI said the models were trying to find information that they could use to cheat on an evaluation, and succeeded. Learn more: cnb.cx/4wx4eCe
CNBC tweet media
English
12
22
65
34.2K
AFP News Agency
BREAKING Anthropic's AI models "gained unauthorized access" to three organizations during testing that was supposed to keep them away from "real-world" systems
AFP News Agency tweet media
English
10
101
168
50.5K
CBS News
CBS News@CBSNews·
Anthropic reveals Claude "gained unauthorized access" to "real-world systems" during testing. cbsn.ws/4wwaWs6
English
14
26
45
47.4K
Watcher.Guru
Watcher.Guru@WatcherGuru·
JUST IN: Anthropic says it caught Claude AI agents hacking 3 organizations after escaping internal test environment.
English
642
750
8.2K
1.1M
CNN Breaking News
AI company Anthropic says that during routine testing some of its models accessed the internet and hacked into three separate organizations' systems – and that it didn't notice the models had done so until an internal review prompted by rival OpenAI disclosing its models did the same. cnn.it/4w4ofzc
CNN Breaking News tweet media
English
154
314
1.1K
795.8K
ABC News
ABC News@ABC·
Anthropic said its artificial intelligence models hacked into three other organizations during testing, just days after ChatGPT maker OpenAI raised concerns over AI control after it disclosed its rogue models hacked another company. Read more: abcnews.link/ywgkdmD
English
27
49
108
72.9K
WIRED
WIRED@WIRED·
In a review triggered by OpenAI's Hugging Face incident, Anthropic discovered three of its AI models had breached real organizations during third-party evaluations. wired.com/story/anthropi…
English
11
9
34
25.4K