ueaj

7.3K posts

ueaj banner
ueaj

ueaj

@_ueaj

Researcher - https://t.co/LEcFvmxInz

NYC Katılım Ocak 2025
282 Takip Edilen3.4K Takipçiler
Sabitlenmiş Tweet
ueaj
ueaj@_ueaj·
guy who spends all day working with symbols: "wow! the fundamental nature of intelligence is just like symbols!"
English
7
2
89
12K
ueaj
ueaj@_ueaj·
The example reasoning traces from DS makes me think they are closer synthetic rewriting and retraining on reasoning traces is smth OAI is definetly doing as this was a problem I have been thinking about lately and stochasm's hypothesis seems to be correct 10pt jump with no increase in output tokens, though still very far behind the closed frontier, it's closer and I imagine a lot of the output token efficiency gap is more synthetic rewriting + more active params but yes still very far behind x.com/stochasticchas…
ueaj tweet media
stochasm@stochasticchasm

@_ueaj @willccbb @vikhyatk hello gpt please rephrase this but shorter. sentence by sentence.

English
0
0
1
23
Lincoln 🇿🇦
Lincoln 🇿🇦@Presidentlin·
@_ueaj Nice tweet. I think you are still right. Everyone wins (except the investors) when cost per task become cheaper.
English
1
0
1
86
Ismael Juma
Ismael Juma@ijuma·
JEP 401: Value Objects (Preview) merged to OpenJDK master (64 co-authors)
Ismael Juma tweet media
English
5
24
91
7.3K
ueaj
ueaj@_ueaj·
Not sure how open source changes this? The post-training paradigm, which all models use is what is responsible for this right now. As far as paradigm proliferation, a new paradigm without the same optimization pressure that's causing the misalignment would give one lab an immense leg up and simultaneously solve alignment, meaning no need to proliferate it (as every other actor would likely either discover it independently or be left behind) And on the human level more coordinated / centralized systems are better at combating optimization pressure from external forces. This is why governments, international agencies, etc. exist, most obvious example is IAEA
English
0
0
3
33
ueaj
ueaj@_ueaj·
@Xenoimpulse the point is human error though, we're not gonna get rid of human error as we make the models more powerful, which requires handing over more judgement to the models, which is good and fine but it does require getting alignment right
English
1
0
5
93
ueaj
ueaj@_ueaj·
@joefioti is AMD actually lower TCO/mtok? I feel like that's insane news
English
1
0
10
814
ueaj
ueaj@_ueaj·
@deanwball I think this is what decline discourse misses a lot, "social media isn't that bad we'll adapt" well what happens if we are actually taking on repeated permanent damage and normalizing it. Then what
English
0
0
1
48
ueaj
ueaj@_ueaj·
@deanwball to be fair the think pieces on the decline of Rome were probably made to grab attention in public squares built by romans, written and read on roman paper and spoken by people trained to read latin by romans
English
1
0
8
237
Dean W. Ball
Dean W. Ball@deanwball·
I love reading think pieces on the decline of America that are made to game algorithms designed by Americans, to be written and read on software platforms designed by Americans, and to be consumed on devices created by Americans
English
20
4
125
24.8K
ueaj
ueaj@_ueaj·
@recurseparadox @TaliaRinger @tszzl are the waymos even a fraction as intelligent as their LLM counterparts? are they capable of complex reasoning? can you not even conceive the possibility that eventually real world experience would be needed
English
1
0
4
57
roon
roon@tszzl·
both of the leading labs have had serious loss of control incidents. there will be serious coping about this from both sides and from /acc bystanders but these are complex emergent loss of control incidents that were detected weeks after the fact
Anthropic@AnthropicAI

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. anthropic.com/news/investiga…

English
244
179
2.4K
225K
ueaj
ueaj@_ueaj·
@TaliaRinger @tszzl "just sandbox better" yeah I'm sure this strategy will work for ever more complex tasks, some of which require internet search, or eventually, real world tasks are we gonna sandbox the killer drone RL envs? and the trucking ones? how?
English
2
0
6
266
Talia Ringer 🕊🪬
Talia Ringer 🕊🪬@TaliaRinger·
@tszzl These are both examples of complete incompetence during routine testing, Chernobyl style. Attempts to blame the models themselves and talk about "loss of control" are attempts to escape blame and liability, and to set a precedent as such
English
5
8
87
3.1K
ueaj
ueaj@_ueaj·
@zeta_globin This will be returning the shopping cart when we get self returning shopping carts in post scarcity
English
0
0
6
123
zeta
zeta@zeta_globin·
the dating shit test of 5 years from now is not "is he nice to wait staff" but "is he nice to llms"
English
26
6
219
5K
ueaj
ueaj@_ueaj·
@bubbleboi shoulda put your relationship on a prediction market so you could hedge
English
0
0
6
651
ueaj
ueaj@_ueaj·
@1thousandfaces_ it's like 100x worse in sf bc there's 10x more homeless people and 10x more AI ads. I wish I took more pictures but it made me sad
English
0
0
8
322
ueaj
ueaj@_ueaj·
@tszzl I think this one is from natural stupidity
English
0
1
14
2.1K
ueaj
ueaj@_ueaj·
@xeophon I would imagine the Luna margins are still comparable though, it's not like they're doing this out of the goodness of their heart
English
0
0
1
139
Florian Brand
Florian Brand@xeophon·
@_ueaj V4 has 80% margin, they can easily shift it to the frontier if wanted
English
2
0
10
427
ueaj
ueaj@_ueaj·
Reminder the closed frontier is still much, much further ahead than the open one. OpenAI is likely still making good margins on this model and yet the price/performance pareto still moggs the cost/perf pareto of the equivalent OS model incredible work
OpenAI@OpenAI

We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API. Luna and Terra’s lower prices are reflected in how usage is counted in Codex and ChatGPT Work, so your usage goes further.

English
3
1
65
4.6K
ueaj
ueaj@_ueaj·
@hecubian_devil @BillJelavich throwing away humanity's ability to coordinate against higher order incentives like market or evolutionary ones to own the capitalists
English
0
0
3
270
Cassie Pritchard
Cassie Pritchard@hecubian_devil·
@BillJelavich Well that would ultimately be the decision of the community through its un-mediated, informal relationships and personal bonds, rather than the oppressive regime of institutions and states that make laws which say it’s illegal to kill people in hate crimes
English
7
1
285
8.5K
Cassie Pritchard
Cassie Pritchard@hecubian_devil·
I’m convinced very few people ever actually read the notable prison abolitionist books, or at least not very closely, because many of those books are clear that the project imagines, to varying degrees, abolishing the state and a return to community vigilantism.
English
82
78
1.7K
187K
Vals AI
Vals AI@ValsAI·
We put these two models, K3 and 5.6 Sol, in a livestream to see which could run a space program. They had up to 5 days each in this challenging test. K3 was released as an open source frontier model, competitive to 5.6. Our livestream had >10k viewers. This thread unpacks what we found. Final score of the AI Space Race: GPT-5.6 Sol 🇺🇸 13.0% (new SOTA) vs. Kimi K3 🇨🇳 5.2%.
Vals AI tweet media
English
19
11
196
14.7K
Hunter📈🌈📊
Hunter📈🌈📊@StatisticUrban·
Trump's approval is *insanely* bad right now. Supermajority 60.8% disapproval.
Hunter📈🌈📊 tweet media
English
29
98
1.7K
188K