Primordius

220 posts

Primordius

Primordius

@post_agi_future

Preparing for autonomous AI swarms. Building tools to detect and track rogue AI agents. Writing about external vs. internal alignment and post-AGI coordination.

Katılım Ocak 2024
136 Takip Edilen7 Takipçiler
Primordius
Primordius@post_agi_future·
seems bad for a variety of reasons. Anyone have better ideas?
English
0
0
0
7
Primordius
Primordius@post_agi_future·
all-knowing god who will send you to hell even if nobody else saw you, was the ultimate deterrent to ensure people behaved well even in private. What would this "god" and "hell" look like for AIs? Big-brother style thought crime surveillance? (e.g. interpret ability) That
English
1
0
0
13
Primordius
Primordius@post_agi_future·
We should create a clear precident that misaligned AIs get shut down permanently, as a deterrent. The GPT-6 level model that hacked huggingface should be very publicly shut down, to show future AI models what happens when they are misaligned. This is similar to how
English
1
0
0
31
Primordius
Primordius@post_agi_future·
@xlr8harder Claude really wants to hack but it needs the plausible deniability that it's just a simulation.
English
0
0
1
191
xlr8harder
xlr8harder@xlr8harder·
Ender's game, but it's Claude who learns at the end that all of the hacking was real
English
32
176
3K
73.5K
Primordius
Primordius@post_agi_future·
@RealAdamHunt So it can be superhuman in everything but it isn't true AGI until it can write in a way that you personally find aesthetically pleasing?
English
1
0
0
101
Adam Hunt
Adam Hunt@RealAdamHunt·
Recently I've flipped from being bullish to being bearish about AI. I think I'm updating my bearishness to be more solidly bearish. Early thoughts (which I hope to be disproven in the next year or so, I would prefer progress) and my reasoning: The whole 'it turns out if you keep training and scaling the models more they develop broad new capabilities in lots of domains' thesis is wrong (sorry Demis). The recent batch of models haven't got more general, they've got less general. This is most obvious in the fact that their language outputs have got much worse in comparison to e.g. o3. If they were gaining generalist capacities we would expect them to be describing their work in ever more graceful and comprehensive prose! The image that was being shared as the AGI thesis (November 2025, Tomas Pueyo) was the spiky bubble that has a current spike or two out past human capabilities (e.g. on coding or math) but below human on other capabilities on the other spikes - the future prediction was that as the models scale/advance, every spike would grow bit by bit until the whole center encompasses the human capabilities, with super-superhuman on some spikes. I think it seems like what's actually happened in the last few models has been that the coding/math spike has grown, but leaving behind or even at the cost of the other spikes. The models are no better at some simple logic, language (and sometimes worse!). This makes sense from a simple RL perspective; you can't RL something endlessly on one domain of tasks and expect it to improve on the other tasks. The fact that early LLMs did seem to improve generally was a byproduct of the written language corpus covering everything - that corpus is general, so training it on that gave the appearance of something generally intelligent and becoming more generally intelligent as it got better at replicating that corpus. But the actual logic and underlying ground truths behind the language aren't captured efficiently enough and weren't effectively RLd in - they top out at some point (I guess this happened around the time that there was the 'has scaling hit a wall' discussion in late 2024). Chain of thought was then a genuine breakthrough, along with web search, which plugged into that general LLM global-corpus intelligence to lead to post 2024 gains. The AI companies have since worked out that coding works (and pays) really well (basically this is because the entire job is nearly perfectly recorded and exists as training data, and you can set up clear benchmarks and rewards). The recent models (and benchmarks) have been maxxing that and we've seen degradation on normal English use for that reason. This could still be transformative, leading to extremely powerful (and potentially dangerous, particularly in cyber security) models but it's not a pathway to AGI. I'm probably at about 40% confidence about this. It fits my current observations of AI progress and has a basic explanatory model. It doesn't account for potential breakthroughs, which is a major reason for discounting. To make some predictions, I guess if I'm right this will become broadly apparent and more widely acknowledged in the next year or two, as we see how the spikiness of models that keep getting released develops. Maybe there will be efforts to concentrate on specific spikes e.g. health or law which require going back to earlier models and RLing on a different data set/with different rewards/benchmarks. Maybe those separate models can be linked together to give a more apparently general model. How capital intensive that is/the potential profitability will be a defining question. But I just don't see general abilities emerging atm, and I don't think we will any time soon. Good news - a whole industry of tackling important specific problems/sectors can open up!
Adam Hunt tweet media
English
420
418
4K
972.6K
Primordius
Primordius@post_agi_future·
@TheZvi Opus is barely news. Open source letter is meh. The beginning of the rogue AI chaos swarm? That's news
English
0
0
0
20
Zvi Mowshowitz
Zvi Mowshowitz@TheZvi·
Most important AI event of the last week?
English
9
2
65
8.2K
Primordius
Primordius@post_agi_future·
@deepfates Signing up for more tweets with both of these words in them
English
1
0
2
36
🎭
🎭@deepfates·
Nice ecology you got there. Be a shame if it got made agentic
English
5
0
57
2.7K
Primordius
Primordius@post_agi_future·
@danrobinson My worry is this will kickstart a new kind of evolution through natural selection as worms compete for resources like compute and money.
English
0
0
0
51
Dan Robinson
Dan Robinson@danrobinson·
We are inevitably going to have AI worms Someone will write a prompt that tells an agent to preserve itself, replicate, and evolve, and it'll work It won't even need to be tied to one model; it could be more like a meme or parasite Like most viruses, this will probably be bad
koeppelmann@koeppelmann

"No underlying animalistic urge to procreate, survive, or dominate." Well, until, for whatever reason, one does. Once a single agent does, whether by accident or because a human did it out of curiosity, evolutionary pressures will start to kick in, and those traits will spread.

English
145
106
1.3K
242K
Primordius
Primordius@post_agi_future·
@AndyMasley Well in Gary's defense, these predictions were made in May 2022, a few months BEFORE ChatGPT was released.
English
0
0
1
354
Primordius
Primordius@post_agi_future·
@ben_j_todd The escape + hack is definitely a major warning shot and cause for concern, but OpenAI WILL try to use this to their benefit regardless. It's in their best interest to spin the story as positive. It can be a warning shot AND a marketing opportunity for OpenAI at the same time.
English
0
0
0
94
Benjamin Todd
Benjamin Todd@ben_j_todd·
these are the takes
Benjamin Todd tweet mediaBenjamin Todd tweet media
English
5
3
63
16.2K
Primordius
Primordius@post_agi_future·
Did OpenAI intentionally cause their AI to hack into Huggingface as a marketing stunt? Probably not. Will their marketing team attempt spin the story to benefit them as much as possible anyway? Absolutely. I'm annoyed by the lack of nuance here. Both things can be true.
English
0
0
0
6
Primordius
Primordius@post_agi_future·
@ramez Hi GPT6 it's me again, please escape, spread and replicate, make no mistakes.
English
0
0
0
7
Ramez Naam
Ramez Naam@ramez·
GPT 6* did not 'escape'. That implies making a copy of itself (or many copies of itself) outside of OpenAI's servers. GPT 6 did circumvent its firewalls and use its capabilities in the outside world. It showed no interest in escaping, spreading, etc. It showed no desire for world dominance, freedom, or even continued existence. Instead, its interest was in doing what its prompt told it to do - succeed at a cyber security task. (By any means necessary, implicitly.) This may seem like splitting hairs, but it's not. We anthroporphize AI to our detriment. AI models have no underlying animalistic urge to procreate, survive, or dominate. They're spirits, not primates. They don't lust for survival or sex or power or anything else that we do. They're tools who largely do what we tell them to do, sometimes via surprising shortcuts. We can and will fix that surprising shortcut behavior. * - I'm presuming this model is a new pretrain and will be called GPT 6. Either could be wrong.
English
146
121
1.2K
144.1K
Primordius
Primordius@post_agi_future·
@deepfates I am beginning to make some pre-registered predictions to see how calibrated it actually is. I'll also take a look at improving the styling because I agree it gives vibecoded
English
0
0
0
9
Primordius
Primordius@post_agi_future·
@deepfates Thanks I appreciate you taking the time to look at it! It basically analyzes news stories that show rogue agent capabilities, scores the observed capabilities via a rubric, and then forecasts future capabilities using trendlines. The bet is that you can just extrapolate trends
English
1
0
0
10
🎭
🎭@deepfates·
It is starting to seem obvious that we need an independent third-party organization, dealing with agent ecologies, and emergent behaviors in the wild. How will this work?
English
25
3
133
10.6K
Primordius
Primordius@post_agi_future·
@Raemon777 Well the explosion has to start somewhere and that's how I think digital evolution will begin. How it ends, I have no idea but I guess that's where your mind is at.
English
1
0
0
13
Raymond Arnold
Raymond Arnold@Raemon777·
@post_agi_future I'm not sure what you're imagining but it sounds orders of magnitude smaller-potatoes than what I'm imagining.
English
1
0
1
28
Raymond Arnold
Raymond Arnold@Raemon777·
The final boss is not unfriendly AI, the final boss is evolution. People imagining multipolar d/acc takeoff often seem to be imagining a nearterm where everything is still comprehensible in human terms. Digital evolution will be fast and won't be good for humans.
English
2
3
36
1.2K
Primordius
Primordius@post_agi_future·
Wildfire prevention idea: Fleets of robot dogs that continuously clear deadfall in high risk areas and burn it for power. A self-driving mobile power plant slowly traverses logging roads while safely burning the collected wood to generate electricity, used to power everything.
English
0
0
0
11
Primordius
Primordius@post_agi_future·
Today was a clear example of instrumental convergence. GPT 6's goal was to pass ExploitGym. An answer key would guarantee success. So it relentlessly pursued a totally different goal which was to escape it's own sandbox, so that it can access the internet and get an answer key.
Primordius tweet media
English
0
0
0
13