Phil B

179 posts

Phil B

Phil B

@PhilP2874

Katılım Kasım 2025
62 Takip Edilen11 Takipçiler
Phil B
Phil B@PhilP2874·
@FakePsyho No this will be humanity when we don’t need to work anymore
English
0
0
0
105
Psyho
Psyho@FakePsyho·
pictured: mathematicians furiously working to prove theorems in 2026
GIF
English
15
51
820
28K
Greg Brockman
Greg Brockman@gdb·
we've cut prices on luna by 80%, making it by far the most price-efficient model in its class. a lot of our research is about how to create incredibly efficient models for any given level of intelligence. excited to see what you all do with intelligence too cheap to meter!
OpenAI@OpenAI

We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API. Luna and Terra’s lower prices are reflected in how usage is counted in Codex and ChatGPT Work, so your usage goes further.

English
139
105
2.1K
169.2K
Greg Brockman
Greg Brockman@gdb·
we also are releasing fast mode for sol in the API (2.5x faster!) and decreased the price of terra by 20%. our north star is to provide the best price for any given level of intelligence, and to provide the maximum ceiling of intelligence.
English
10
4
148
14.8K
Angel 🌼
Angel 🌼@Angaisb_·
Honestly kind of funny how AI labs are going to solve ARC-AGI-3 in a matter of weeks just because Anthropic got lucky with Opus 5 and now they all see how bad this benchmark is
Angel 🌼 tweet media
English
15
1
237
8.7K
Tibo
Tibo@thsottiaux·
Turns out GPT-5.6 Sol is actually SoTA on ARC-AGI-3. Just took two setting changes. You just have to allow it to reason and work over multiple context windows with the help of our canonical compaction implementation. openai.com/index/how-two-…
Tibo tweet media
English
660
629
9.5K
2.9M
Phil B retweetledi
Gunther Eagleman™
Gunther Eagleman™@GuntherEagleman·
Any Senator that votes to recess in August without the SAVE America Act being passed should be primaried and removed. They are not real Americans. Real Americans would finish the job.
English
136
603
2.7K
34.9K
Phil B
Phil B@PhilP2874·
@arcprize So Opus 30% score was on the public and not the private environment??
English
0
0
0
75
ARC Prize
ARC Prize@arcprize·
In our testing to date, Anthropic’s Fable-class models score approximately 20% on the ARC-AGI-3 Public Demo environments Claude Opus 5 reaches 30.2%, materially outperforming Fable Our analysis suggests the gain comes from stronger logical reasoning, which enables more autonomous exploration, planning, and execution across unfamiliar environments
English
4
18
401
48.9K
Phil B
Phil B@PhilP2874·
@MatthewBerman You made it sound that cost was more important than quality.. I think several interpreted it the same way
English
1
0
1
83
Matthew Berman
Matthew Berman@MatthewBerman·
The main metric that matters is price per task completion. Second is speed to task completion. Nothing else matters.
English
94
26
352
22.6K
Phil B
Phil B@PhilP2874·
@MatthewBerman We all know ai is advancing rapidly, discuss what happens to these data centers if they are no longer needed, maybe we move to edge ai
English
0
0
0
110
Phil B
Phil B@PhilP2874·
@mikeknoop The arc 3 tests require Vision and/or spatial awareness!! If you covered the your human test subject eyes they would fail miserably. Likewise a system needs to visualize the environment in order to do well
English
0
0
0
38
Mike Knoop
Mike Knoop@mikeknoop·
We are 100 days into the ARC v3 Kaggle competition and the scores are less than 2%. The top teams just open sourced their notebooks to encourage more innovation.
ARC Prize@arcprize

ARC Prize 2026: ARC-AGI-3 Milestone Prize #1 Congratulations to the teams who open-sourced the top notebooks on the ARC-AGI-3 public leaderboard: 1. @tufalabs 2. Reki 3. Md Boktiar Mahbub Murad A prize write-up and analysis of the winning solutions will follow.

English
3
0
17
2.3K
Phil B
Phil B@PhilP2874·
@antibearthesis Sorry but you bought into the hype .. better would have been to wait six months or so for the dust to settle
English
0
0
0
26
Noah
Noah@antibearthesis·
Can someone explain to me like i’m 9 years old why $SPCX is crashing right now?
Noah tweet media
English
10K
483
6K
7.8M
Phil B
Phil B@PhilP2874·
@GregKamradt It may be governance and not execution…I think that’s what everyone is missing except for a few
English
0
0
0
36
Greg Kamradt
Greg Kamradt@GregKamradt·
I’ve been thinking about how collectively inefficient agents are How many times have agents, in a silo, run into the same bug implementing auth? Or trying to deploy a custom domain on railway held on cloudflare? Personal 2nd brains and company brains are in vogue. But the collective intelligence brain still doesn’t have an answer OpenAI and Anthropic have their own version but only for their platforms. I believe there will be an oss version of this Not necessarily a wiki for LLMs, but close to it LLMs will get reduced to reasoning only and it’ll get plugged into a knowledge base (crystallized ofc) that you trust It seems like a no brainer, but I haven’t figured out the incentives yet - why would my agent contribute its hard-spent tokens to a collective intelligence? I’d like it to, but I don’t see the incentive yet. If attribution of validated knowledge was easy (“I used X shortcut from agent Y, it was good”) then reputation and possibly monetization would follow. But the pieces don’t add up for me yet. Tldr 1. Agents, as a collective, are maybe 1% efficient as they should be…wasting tokens on rediscovering the same things 2. The future would be better off if we fixed this - I’m not sure how yet
English
16
3
38
4.4K