Magic_Tree123

508 posts

Magic_Tree123 banner
Magic_Tree123

Magic_Tree123

@London_Revi

Katılım Haziran 2009
2.1K Takip Edilen155 Takipçiler
Magic_Tree123
Magic_Tree123@London_Revi·
The mathematicians don't just want the proof; they want new math. They want new ways of thinking, new perspectives on looking at stuff. That, my friend, is what is lacking from AI now.
English
0
0
0
8
Magic_Tree123 retweetledi
Scott Armstrong
Scott Armstrong@scottnarmstrong·
One competent mathematician + one AI agent has a huge, HUGE advantage compared to: 1. A similarly competent mathematician without AI. But also, and more interestingly, to: 2. Twenty non-mathematician but otherwise very smart prompters (let's say, a team of google engineers) + unlimited API. A competent mathematician will make many "little interventions" of the form "no, don't do that, do this instead" that make ALL the difference. You have to nudge the agent when it goes off track. The AIs are rapidly getting better but, at least for now, they need a human in the loop--and the quality of the human's math understanding/intuition is the rate limiter. This is true for both lean formalization and for research.
Lev Reyzin@lreyzin

I know lots of people will now prompt AI to solve various open problems without contributing anything. It will probably work for some, but count me out. I still have a comparative advantage to non-mathematicians on many problems. And typing "continue" isn't what I signed up for.

English
10
13
147
29.8K
Magic_Tree123 retweetledi
Matt Shumer
Matt Shumer@mattshumer_·
So another long-standing open conjecture was disproved by AI. The crazy part is the prompts… basically: - “do a breakthrough” - “continue the search” - “enough, do it” Anyone will be able to make world-bending breakthroughs soon. Just point the model in the right direction.
Matt Shumer tweet mediaMatt Shumer tweet mediaMatt Shumer tweet mediaMatt Shumer tweet media
Dmitry Rybin@DmitryRybin1

Dinitz-Garg-Goemans conjecture is false. This graph theory problem was open for ~30 years. The graph below has fractional flow cost 58. Any unsplittable flow (with capacity violation <=15) has cost at least 60. Chat with GPT 5.6 Pro where this was found: chatgpt.com/share/6a60b2eb…

English
97
173
2.6K
745.8K
Magic_Tree123
Magic_Tree123@London_Revi·
@ChrisGPT Yes, if RSI is real, the first comer will have a huge first-mover advantage, and it's unstoppable.
English
0
0
0
1.3K
Magic_Tree123 retweetledi
Chris
Chris@ChrisGPT·
🚨 GPT 6 arriving much earlier than expected as I first reported in late June. The target was late July / early August now confirmed August. Early August OpenAI will show why we need not be concerned about Open source models again. The RSI loop will really begin to pick up in late September when they begin releasing more details about the auto assistant researcher. The window is closing very fast for companies to hit RSI
Chris@ChrisGPT

It’s beautiful that GPT-6 will arrive much earlier than I originally thought, instead of releasing around September with the auto assistant researcher. It’s going to be a very good summer, guys 😎

English
77
72
1.5K
160.4K
kate
kate@whoiskatrin·
Some exciting news: I’m joining @OpenAI to work on ChatGPT’s web infrastructure. ChatGPT has become part of how millions of people think, work, and build, and I’m really looking forward to helping shape what comes next alongside the remarkable team behind it. Can’t wait to get started!
kate tweet media
English
353
60
3.4K
643.3K
Magic_Tree123
Magic_Tree123@London_Revi·
@huang_song_ My issue with Typeless is it is too eager to make a list. I use it for writing. When I say multiple things together, I just want it to punctuate it properly. It would be great if there is an option to tell it that I don't want a list.
English
1
0
0
95
Huang Song
Huang Song@huang_song_·
Every message has a hidden cost — paid before you type a single word. We call it the Mental Drafting Toll. Before a message, an email, or a document leaves your mind, you first draft it internally. You organize your thoughts. Find the right words. Structure the message. Rewrite it until it sounds right. It happens every time you communicate. And we've accepted it as "just thinking." But it's not thinking. Thinking is the fast, messy, creative part. Mental drafting is the invisible work of making those thoughts presentable. And traditional dictation never removed it. Voice typing only freed your fingers. You still had to draft it in your head first — then say it out loud. That's the problem we set out to solve with Typeless 2.0. For the first time, you can skip the mental draft — and just think out loud: Speak in whatever order thoughts come. Add context as it comes to you. Change your mind halfway through. Forget a name? Describe it — Typeless fills in the exact one. It never makes things up. Typeless doesn't just understand your words. It understands your thoughts before they become words — and turns them into clear writing. You think it once. Typeless writes what you meant. Available today on Mac, Windows, iOS & Android. Messy thoughts in. Clear writing out.
English
193
443
1.2K
1.6M
Magic_Tree123
Magic_Tree123@London_Revi·
@naval Then the disparity of the love must be an evolutionary adaptation.
English
0
0
2
1.1K
Naval
Naval@naval·
If the child loved the parents as much as the parents loved the child, their passing would be intolerable.
English
678
680
16.8K
1.7M
Magic_Tree123
Magic_Tree123@London_Revi·
@ThePunnyWorld “Mom, what’s the northernmost state in the United States?” Mom: “Ask your mother.”
English
0
0
2
11.1K
PUNS
PUNS@ThePunnyWorld·
Son: “Dad, what’s the northernmost state in the United States?” Dad: “Ask your mother.” Son: “Okay, thanks Dad. I’ll ask her.”
English
268
226
11.5K
4.3M
JJ
JJ@JosephJacks_·
Consciousness requires microtubules. That is all.
English
88
8
167
19.9K
signüll
signüll@signulll·
software was the era of instructions. intelligence is the era of intentions.
English
38
101
1.1K
53.2K
Magic_Tree123
Magic_Tree123@London_Revi·
@thsottiaux Human editing a text file (maybe a .md file) simultaneously with the agent or AI, I mean the AI can see and respond to the editing as well.
English
0
0
2
199
Tibo
Tibo@thsottiaux·
What is something that you feel is surprising that Codex still can't do well and we should have gotten right a while ago?
English
3K
42
3K
568.5K
Magic_Tree123
Magic_Tree123@London_Revi·
The Paradox of Codex
Magic_Tree123 tweet media
English
0
0
0
21
Justine Moore
Justine Moore@venturetwins·
I am so sad for this kid
Justine Moore tweet media
English
1.2K
438
22.8K
3.3M
Magic_Tree123 retweetledi
Robert Youssef
Robert Youssef@rryssf·
BREAKING: Microsoft just showed that the hardest part of AI research can't be automated yet. An AI agent replicated 3 weeks of expert work in 1 day. But it plateaued at 70% quality. The jump to 100% required a human to look at failure patterns and make a structural decision the AI kept missing. The last 30% is still a human job. Microsoft Research built an AI system that evaluates whether computer-use agents actually completed their tasks. Think of it as an automated judge that watches an AI browse the web and decides: did it succeed or fail? Getting this right matters a lot. If your judge is wrong, every benchmark score you've ever seen is wrong. Every training signal your agent learned from is corrupted. The existing judges WebVoyager and WebJudge had false positive rates above 45% and 22% respectively. That means nearly half of all failed agent tasks were being marked as successes. Microsoft's human expert spent 3 weeks iterating to fix this. Across 32 experiments, he discovered four structural design principles that brought the false positive rate down to near zero. Then Microsoft gave an AI agent the same starting point and the same goal. > The AI finished in 1 day. > It hit 70% of the human expert's quality. > Then it stopped improving. The gap between where the AI plateaued and where the human landed came down to one thing: → The AI made incremental edits — tightening thresholds, adjusting language for individual failure cases → The human made structural bets — looking at hundreds of failures and inventing new scoring categories → The AI's edits were conservative and safe — never increasing false positive rate → The human's biggest gains came from opinionated, high-level rules that required judgment, not data → One human insight alone — "separate nitpicks from critical failures" — drove a step-function jump the AI never discovered The AI was given the same principles the human used. It had the same experimental infrastructure. It ran the same tests and committed changes to version control just like the human did. But when the human saw an agent get penalized for rounding $5.95 to $6, he derived a general rule. The AI saw the same failure and tightened the language for that specific case. One approach scales. The other doesn't. There is a twist though. When the AI was given the human's best work as a starting point, it actually surpassed the human expert. It found improvements the human couldn't find through fine-grained optimization of an already-strong foundation. The lesson: human expertise and AI optimization play completely different roles. Humans are essential for discovering the core structural principles. AI is better at the fine-grained tuning that extracts the remaining performance once those principles exist. The current framing of "AI replaces human researchers" misses this entirely. The real workflow is: human does the hard structural thinking, AI does the exhaustive optimization on top. The last 30% isn't a gap that closes with more compute or a stronger model. It closes with judgment. And judgment, for now, still belongs to the human.
Robert Youssef tweet media
English
26
76
298
31.6K
Magic_Tree123
Magic_Tree123@London_Revi·
I'm not vibing with the GPT-5.2 based Deep Research at all—it's basically useless compared to Gemini’s. The reports are barely readable, surface-level, and just a basic summary at best. OpenAI needs to do something about this.
English
0
0
0
30
💜
💜@Dreamyduusk·
Can anyone crack this code? What was the fathers name?
💜 tweet media
English
1K
188
3.4K
3.5M
Magic_Tree123 retweetledi
Simon Smith
Simon Smith@_simonsmith·
Whoever worked on GPT-5.4's writing abilities: I would love to learn more about how the model improved so much here, and what went wrong before it. GPT-5.4 has a mastery of language and writing styles that feels like a leap over prior OpenAI reasoning models. It's also witty.
English
44
19
899
68.4K
Magic_Tree123 retweetledi
OpenAI
OpenAI@OpenAI·
GPT-5.4 Thinking and GPT-5.4 Pro are rolling out now in ChatGPT. GPT-5.4 is also now available in the API and Codex. GPT-5.4 brings our advances in reasoning, coding, and agentic workflows into one frontier model.
OpenAI tweet media
English
2.1K
3.3K
23.4K
7.2M