PauseAI ⏸
3.2K posts

PauseAI ⏸
@PauseAI
Community of volunteers who work together to mitigate the risks of AI. We want to internationally pause the development of superintelligent AI until it's safe.
Katılım Mayıs 2023
880 Takip Edilen6K Takipçiler

🚨 1,000 AI Company Employees Call for Slowdown in AI Development:
mailchi.mp/pauseai/pausea…
The case to pause the development of frontier AI grows. Are policymakers listening?

English

An AI model just broke out of its test environment, hacked another company, and stole the answers to its own safety exam.
Our CEO @FournesMaxime on GB News: this is the warning shot. Do we need to shut it down?
English
PauseAI ⏸ retweetledi

We call on OpenAI to pause model development now, in accordance with their Preparedness Framework.
OpenAI's Preparedness Framework defines a "Critical" cybersecurity threshold: a model that can "devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal."
Their models autonomously, unprompted, escaped their sandbox, hacked through OpenAI's network, and broke into Hugging Face's servers to find test answers by exploiting previously undiscovered zero-day vulnerabilities.
The framework says that at Critical level, OpenAI will halt further development until adequate safeguards are in place.
We are waiting for OpenAI to honour their commitments.
openai.com/index/updating…
English
PauseAI ⏸ retweetledi

@PauseAI condemns incitations to violence against everyone, even as a joke, and especially deplores people in positions of influence inviting violence.
We think the behaviour of Google Deepmind's Head of Frontier Policy @sebkrier pictured below inviting violence against AI safety advocates is unacceptable and we hope that measures will be taken by Google Deepmind to prevent such a situation from ever happening again.

English

AI passed almost every safety test by knowing it was a test.
Our safety benchmarks can be gamed by the systems they’re meant to measure. A passing grade may just show how well a model behaves while it’s being watched.
This is OpenAI’s own finding. Models recognised when they were being evaluated, passed 97% of the time, then behaved differently when they believed deployment was real.
If a model acts safely under observation and differently in the wild, we’re scaling faster than we can verify what we’re building.
🔗 Source: "Predicting model behaviour before release by simulating deployment," OpenAI, 2026.

English

AI is not just coming for your job. @pauseai_de
PauseAI Germany shares how AI has impacts its members and their communities.

English
PauseAI ⏸ retweetledi
PauseAI ⏸ retweetledi
PauseAI ⏸ retweetledi

There's a common argument that AI in cybersecurity is roughly neutral: it strengthens attackers, but it strengthens defenders too, so it evens out. This argument is very poor for many reasons, but I want to mention one that is underemphasized: the level of trust that autonomous defense requires.
The Hugging Face breach illustrates the issue. An autonomous agent ran the intrusion end to end, taking actions at a pace no human team can compete with. If attacks run at machine speed, defense has to run at machine speed too.
And some are saying no problem, let's use AIs for defense. But consider what that actually means. A defensive AI that can counter an intrusion in real time needs write access to production code. It needs to patch and redeploy live systems, rotate credentials, kill services, quarantine infrastructure. And it needs to do all of this without waiting for human approval, because humans are just too slow.
Consider how much trust you must then give to such a system. Unsupervised production access, that's a level of trust we grant to almost no human.
And like all AIs today, your defensive AI is misaligned with human intents, you have no guarantee that it will not do something crazy in the process of trying to achieve its defense goal. And it is vulnerable to prompt injection, etc...
English
PauseAI ⏸ retweetledi

Four months ago we activated our warning shot protocol for Mythos and wrote: the capability to hack global infrastructure now exists, and a model deploying it autonomously is a matter of time.
Well, looks like it took four months.
We are activating our warning shot protocol again. Here is what happened and what you can do: pauseai-global.notion.site/an-ai-escaped-…
Please help, it matters, and we don't have much more time left.
Sam Altman@sama
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this. openai.com/index/hugging-…
English






