Lorenzo

1.3K posts

Lorenzo

Lorenzo

@Lorentzano

Katılım Kasım 2025
458 Takip Edilen39 Takipçiler
Lorenzo retweetledi
Gary Marcus
Gary Marcus@GaryMarcus·
Brilliant, must read lecture on AI and math by leading mathematician Terence Tao. Among the key points: (a) we are risk for “proof indigestion”; not every proof is useful and (b) math is not uniform (see figure below), and AI might be good at some part of what needs to be done but not all. In keeping with this (and with what I argued yesterday), Astra (his essay was written just before Astra was announced) is good at solving certain types of problems but there is no evidence thus far it can build new theories. teorth.github.io/tao-web/slides…
Gary Marcus tweet mediaGary Marcus tweet media
English
11
31
145
19.3K
Lorenzo retweetledi
Specter
Specter@DemonFramed·
Maturing is realizing that 99% of people are programmed bots, and there’s nothing you can do about it. They wake up at the same time, follow the same routine for 60 years, never question anything, never think for themselves, let their bodies fall apart, and obsess over things that don’t improve their own lives. They don’t want freedom. They want instructions. They don’t want to build a different life. They want someone else to tell them what success looks like. There is no revolution coming, it’s over.
English
4
22
170
4.3K
Lorenzo retweetledi
Epic Maps 🗺️
Epic Maps 🗺️@theepicmap·
The 4 Europes.
Epic Maps 🗺️ tweet media
English
606
1.3K
21.8K
1.3M
Lorenzo retweetledi
Jonathan Blow
Jonathan Blow@Jonathan_Blow·
Hello, we have posted another TECHNICAL UPDATE for the Demo of the Sinking Star on Steam. This version contains Shader Model 6.5 support (compatibility with more GPUs!) as well as some further optimizations and bug fixes. Let us know how it works for you. To try it, you go to the Properties for the game, go to Game Versions & Betas, and pick public-testing (no password).
Jonathan Blow tweet media
English
26
12
449
30.9K
Lorenzo retweetledi
Paul Graham
Paul Graham@paulg·
I was curious why LLMs have gotten so good at math and still aren't that good at writing, so I asked an expert. It's not because math is easier. It's because it has clear right and wrong answers and so is easier to train on. But they are coming for me next.
English
458
141
4.7K
279.7K
Lorenzo retweetledi
Polymarket
Polymarket@Polymarket·
JUST IN: Italy moves toward reviving nuclear power after four decades, with next-generation reactors potentially operating by 2033.
English
86
283
4.7K
133.6K
Lorenzo
Lorenzo@Lorentzano·
LLMs without human made tools cannot achieve even 2% of ARC-1 performance. Despite massive amount of data
Gary Marcus@GaryMarcus

@benglickenhaus are you literally so stupid that you don’t understand the difference between an LLM and LLM that has been supplemented with other things such as tools and harnesses? or are you just pretending? blocking for intellectually dishonesty or stupidity but i can’t tell which.

English
0
0
0
0
Lorenzo retweetledi
Gary Marcus
Gary Marcus@GaryMarcus·
@benglickenhaus are you literally so stupid that you don’t understand the difference between an LLM and LLM that has been supplemented with other things such as tools and harnesses? or are you just pretending? blocking for intellectually dishonesty or stupidity but i can’t tell which.
English
10
4
116
5K
Lorenzo retweetledi
Gary Marcus
Gary Marcus@GaryMarcus·
Yet another paper argues that LLMs aren’t close to doing real discovery.
How To Prompt@HowToPrompt__

MIT and Harvard argue LLMs are nowhere near doing real scientific discovery. They published a paper called “Evaluating Large Language Models in Scientific Discovery.” Every week, tech labs claim an LLM has made a breakthrough in biology, physics, or chemistry. But this proves they are faking it. For years, AI benchmarks have tested models using static, multiple-choice science trivia. Models ace these tests, leading everyone to believe AI is right on the verge of autonomous scientific discovery. Researchers built a new evaluation framework called SDE to test what happens when you take LLMs out of the multiple-choice quiz and put them into real, open-ended research projects. They tested frontier models across biology, chemistry, materials science, and physics. The results are sobering. When forced to handle the actual loop of discovery—proposing a testable hypothesis, designing simulations, running experiments, and interpreting ambiguous results iteratively, current LLMs fall apart. There is a massive, glaring performance gap between passing standard science benchmarks and doing real science. Why do they fail? Because real science requires iterative reasoning, handling imperfect evidence, and adapting to unexpected observations. LLMs are built to predict the next token based on existing internet data. They can regurgitate a textbook explanation of photosynthesis or quantum mechanics instantly. But when placed inside an uncharted loop where the textbook doesn't have the answer yet, they hit a wall. Worse still, the researchers discovered diminishing returns. Simply scaling up model sizes and adding raw compute isn't fixing the gap. Top-tier models from different providers share the exact same blind spots. We are miles away from general scientific superintelligence. The tech industry is selling a narrative that AI is about to automate labs, run clinical trials, and invent materials on autopilot. But right now, AI isn't doing science. It's just remembering it.

English
61
82
451
70K
Lorenzo retweetledi
Lorenzo retweetledi
Specter
Specter@DemonFramed·
Every girl you meet has a toxic ex boyfriend who was supposedly an unemployed, alcoholic bum that had her paying all the bills. “I can’t believe I dated that guy.” Yes, you can. And you’re still letting him sleep with you for free while using the guys you meet on Hinge for free meals. Women will have no problem dating an unemployed bum as long as he’s a Chad who knows how to fuck.
English
24
16
283
19.3K
Lorenzo
Lorenzo@Lorentzano·
I get your point, agreed
antirez@antirez

@fchollet That's undisputed, but the argument of the skepticals was a bit different: that even in the local space represented by the training set, LLMs could not interpolate on the realm of a deeply inferred representation, just on a shallow almost-already-seen one (parrot model).

English
0
0
0
0
Lorenzo retweetledi
François Chollet
François Chollet@fchollet·
LLMs have failed every single benchmark and experiment focused on generalization, since their inception. It's not just ARC -- this is documented in literally hundreds, possibly thousands of papers. The ability of LLMs to solve a task is entirely dependent of their familiarity with the task (local generalization).
English
12
88
945
133.7K
Lorenzo retweetledi
François Chollet
François Chollet@fchollet·
Worth noting that to this day, base LLMs (no test time compute) *still* perform poorly on the ARC 1 benchmark from 2019 (on unseen tasks) despite continued massive scaling (~100,000x since 2019). We can say that without the switch to the test-time compute paradigm, AI would still not be capable of the kind of advanced reasoning that sota systems display today. Scaling the single-pass, static, next-token prediction paradigm of the GPT-2 through GPT-4 era was running into a capability asymptote. To bypass this plateau, test-time adaptation was a necessary evolutionary step.
English
7
12
131
12.8K
Lorenzo retweetledi
Pavel
Pavel@baltabaev·
I’ve spent well over 10,000 hours studying math in my life, yet I can’t understand these proofs, at least not without weeks of digging deep into each topic. What’s more, none of my math PhD friends know much about these problems either, and they can’t verify most of them without working directly in the field (yes, math is VERY diverse). LLMs are getting smarter than the experts themselves, and I’m not sure we have enough bright human minds to verify everything that will come out of them in the coming years. Remember when we compared AI intelligence to PhD students? I think we’re past that.
Noam Brown@polynoamial

An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science. We believe it will be a major step for scientific reasoning. openai.com/index/ten-adva…

English
315
829
7.8K
901.5K
Lorenzo retweetledi
François Chollet
François Chollet@fchollet·
From 2013 to 2016 I was actually quite convinced that RNNs could be trained to learn any program. After all, they're Turing-complete (or at least some of them are) and they learn a highly compressed model of the input:output mapping they're trained on (rather than mere pointwise associations). Surely they could perform symbolic program synthesis in some continuous latent program space? Nope. They do in fact learn mere pointwise associations and completely useless for program synthesis. The problem isn't with what the function space can represent -- the problem is the learning process. It's SGD.
English
21
24
410
187.4K