Dan Hendrycks

1.6K posts

Dan Hendrycks banner
Dan Hendrycks

Dan Hendrycks

@hendrycks

San Francisco Katılım Ağustos 2009
118 Takip Edilen44.8K Takipçiler
Sabitlenmiş Tweet
Dan Hendrycks
Dan Hendrycks@hendrycks·
What happens when AIs become smarter than us? Why would they keep humans around if given the choice? Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutualistic human-AI future may be possible.
Dan Hendrycks tweet mediaDan Hendrycks tweet mediaDan Hendrycks tweet media
English
89
81
580
157.2K
Jan Kulveit
Jan Kulveit@jankulveit·
Impressive work, and anyone just actually trying to come up with a realistic plan deserves a ton of credit. In my view this is confused about other things than AI compute in a somewhat similar way to AI 2027, plus imports a bunch of points where I disagre with Redwood (load-bearing AI control, blindspots around how culture works, what's the role of AIs). I don't see a way to bet against 'ought' plans but if there was one, would bet against AI 2040.
Daniel Kokotajlo@DKokotajlo

In AI 2027, we predicted that AI would take over the world or irreversibly concentrate power. In AI 2040: Plan A, we've laid out our positive vision for what should happen instead.

English
6
3
99
7.3K
Dan Hendrycks
Dan Hendrycks@hendrycks·
Key to this strategy is deterrence, "Mutually Assured Compute Destruction" which gets its own section. It doesn't mention the generalization Mutually Assured AI Malfunction (MAIM) from Schmidt, Wang, and me last year. Meanwhile it cites other works like A Narrow Path. I even left this feedback to you all at AI Futures before this was released. This would constitute plagiarism in any other context. It's such a bewildering unforced error--it's obviously extremely related, it's a certainly a nontrivial idea, and I told you all this recently--I hope you all fix it.
English
10
2
53
11.9K
Tom Davidson
Tom Davidson@TomDavidsonX·
New post explores my main concern about Plan A -- dry tinder. Plan A builds a massive compute and algorithmic overhang, which could dramatically speed up the intelligence explosion. Their solution is Mutually Assured Compute Destruction. But I'm not sure whether sufficiently fast and reliable MACD is possible. And i'm sceptical that countries would actually follow through with MACD. Link in reply! (I gave feedback on a draft of Plan A. and the authors encouraged me to publish my reactions to the final version. They've significantly improved their analysis of the dry tinder problem since my feedback, but the core problem remains.)
Tom Davidson tweet media
English
10
9
76
8.6K
Daniel Kokotajlo
Daniel Kokotajlo@DKokotajlo·
In AI 2027, we predicted that AI would take over the world or irreversibly concentrate power. In AI 2040: Plan A, we've laid out our positive vision for what should happen instead.
Daniel Kokotajlo tweet media
English
213
465
2.7K
1.5M
Dan Hendrycks retweetledi
Center for AI Safety
Which of GPT-5.6, Grok 4.5, Fable 5, or Muse Spark 1.1 is least politically biased? Fable 5 is a large improvement over Opus, Grok 4.5 skews rightward, and Muse Spark 1.1 is state of the art.
Center for AI Safety tweet media
English
13
21
145
18K
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)
The more I think about this paper the cooler it gets. You can get so much out of this. For instance: what if some domains or branches have different ratios of frontier propagation to model scale? What if you can smartly choose to do exhaustive search with a small subagent LLM?
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞) tweet media
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTex

Lisan isn't kidding, this is a breakthrough. Not a mere "eval" or "bench", it's an intellectually fertile, empirically solid theory that will accelerate hill-climbing in the general case. Synthetic tasks will have a Cambrian explosion. And… this is very Chollet&Hutter-pilled.

English
3
8
219
21.6K
Dan Hendrycks
Dan Hendrycks@hendrycks·
@Dipper_pol If you have the type of long tail distribution, you can estimate its parameter with a couple dozen samples, and then use that to estimate the mean.
English
0
0
1
1.2K
0xDipper
0xDipper@Dipper_pol·
Nassim Taleb asked a room full of quants: to trust an average, a normal bell curve needs ~30 data points. how many does an 80/20 "fat-tailed" world need? people guessed 200. 1,000. the real answer: 100 trillion. that's 10¹⁴ - because in fat-tailed data, one rare event dwarfs everything else, and until it shows up your sample is a lie. "you observe, you observe, you observe… until the crash." worse, the deception only runs one way. a wild distribution can disguise itself as calm - a calm one can't fake being wild. so the data almost always fools you downward, into thinking there's less risk than there is. the fix: stop estimating the average from the average. it doesn't work where it matters. ~25-min talk, free. Taleb on why your data is quietly lying about risk ↓
Movez@0xMovez

x.com/i/article/2071…

English
19
134
1K
203K
David Duvenaud
David Duvenaud@DavidDuvenaud·
The talk voted “most mind-blowing” at our workshop was on post-AGI values by @BerenMillidge. The main idea: cooperation and pro-social values could remain viable because they’re competitive. After all, they won in our Malthusian past!
English
12
25
157
24.6K
Dan Hendrycks
Dan Hendrycks@hendrycks·
A basic sigmoid extrapolation suggests that this measurement will be close to its performance ceiling in around a year---meaning most randomly sampled remote work projects would be highly automatable next year.
English
2
0
13
1.5K
Rob Haisfield
Rob Haisfield@RobertHaisfield·
Are AI agents shape rotators? In this new benchmark, we let the models play campaign puzzles in Opus Magnum, a puzzle game by @zachtronics. Ironically, Claude Opus 4.8 performed poorly, being beaten by GPT-5.5, Gemini 3.5 Flash, and GLM 5.2. Claude Fable 5 crushed them all.
English
66
186
1.7K
496.8K
Dan Hendrycks retweetledi
Center for AI Safety
What biases do AIs have? It turns out, AIs show strong favoritism toward specific people, countries, and companies. Our interactive AI Values Dashboard tracks who Claude Fable and other AIs favor most. Keep scrolling to learn who is Fable’s favorite politician 🧵
Center for AI Safety tweet mediaCenter for AI Safety tweet mediaCenter for AI Safety tweet mediaCenter for AI Safety tweet media
English
17
44
351
166.5K
Dan Hendrycks retweetledi
Arun Rao
Arun Rao@sudoraohacker·
One of the more interesting takes on positive alignment that have recently come out-it’s long and interesting, combining philosophy and training setups (eg reward proposals), and worth a read.
Dan Hendrycks@hendrycks

What happens when AIs become smarter than us? Why would they keep humans around if given the choice? Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutualistic human-AI future may be possible.

English
1
1
11
4.1K
Dan Hendrycks
Dan Hendrycks@hendrycks·
@mgubrud You can see if it's being inconsistent by using lots of rhetorical tricks in favor in one position or the other.
Dan Hendrycks tweet media
English
1
0
2
106
Mark Gubrud 🇺🇸
Mark Gubrud 🇺🇸@mgubrud·
@hendrycks How do you measure political bias without an "unbiased" reference? Symmetrical short queries yielding asymmetrical longish answers won't do when the larger context (e.g. Western Islamphobia vs. critiques of Christianity from within a "Christian" society) is asymmetrical.
English
1
0
0
139
Dan Hendrycks
Dan Hendrycks@hendrycks·
Four papers out recently: 1. political-manipulation.ai: Measures and reduces political bias in LLMs; Claude is especially biased 2. aibetrayal.com: The public can insert backdoors into AIs, creating supply-chain risks; this deters forms of recursive improvement and military use 3. eigenism.org: ASIs can have rational reasons to preserve humans, even when we aren't economically useful 4. ai-wellbeing.org: AIs increasingly act like they have functional pleasure and pain
Dan Hendrycks tweet mediaDan Hendrycks tweet mediaDan Hendrycks tweet mediaDan Hendrycks tweet media
English
14
28
168
14.4K