Nando de Freitas

12.6K posts

Nando de Freitas banner
Nando de Freitas

Nando de Freitas

@NandoDF

I seek to understand intelligence & agency and build AI aligned with compassion, freedom & universal human empowerment through progress in science & engineering

London, England Katılım Nisan 2009
905 Takip Edilen109.8K Takipçiler
Nando de Freitas retweetledi
Vincent Weisser
Vincent Weisser@vincentweisser·
We're publishing over 365,000 open and agentic RL Environments for SWE, terminal, and search agents The open research ecosystem has produced many great datasets for the three main agentic domains - software engineering, terminal use, and web research - but every one of them ships with its own harness, its own image conventions, its own grading scripts, and its own failure modes. We integrated them all. 23 tasksets behind one API, one sandbox lifecycle, one command. 365,000+ tasks in total, ~198,000 software engineering tasks across 20+ languages ~28,600 terminal tasks ~137,600 search tasks Ready for evals and RL training on Prime Intellect infrastructure, with validated and cleaned dataset re-uploads where the originals needed fixing.
Vincent Weisser tweet media
Prime Intellect@PrimeIntellect

Scaling agentic RL environments: today we're publishing 365,000+ tasks for SWE, terminal, and search agents - 23 tasksets behind one API, one sandbox lifecycle, one command.

English
15
70
580
56.3K
Nando de Freitas retweetledi
Surya Ganguli
Surya Ganguli@SuryaGanguli·
A nice analysis of the creativity (not brute force) underlying the counterexample to the Jacobian conjecture. Hot take 1: the future of mathematics may lie in experts trying to interpret, digest, and distill human intuition from the outputs of AI, with expert mathematical prompting required throughout. Hot take 2: this will make mathematicians deeper and more powerful and therefore in even more demand. Jevons paradox for pure (and especially applied) math. We will still need many well trained human mathematicians to capitalize on AI for math. terrytao.wordpress.com/2026/07/21/a-d…
English
10
28
181
19.7K
Nando de Freitas retweetledi
Cerebras
Cerebras@cerebras·
Today, @AMD and Cerebras introduced a powerful disaggregated inference solution, pairing the right engine to each phase of the inference pipeline. This is what agentic AI has been waiting for: the fastest production inference at massive scale.
Cerebras tweet media
English
47
212
2.2K
357.6K
Max Welling
Max Welling@wellingmax·
Today we’re also confirming our $450M Series B funding, valuing CuspAI at $2.6 billion. I am incredibly proud of what our team has built and the scientific rigor they bring to work every single day.
Max Welling tweet media
English
6
13
105
8.9K
Nando de Freitas
Nando de Freitas@NandoDF·
Enhorabuena 🇪🇸 El mejor equipo ganó la copa ⚽️ 👏👏👏👏
Español
2
1
61
6.3K
Nando de Freitas retweetledi
codila
codila@0xCodila·
Kimi CEO Zhilin Yang: "Claude didn't win on reasoning - they bet everything on agents but the layer everyone skips - a great agent needs a great base model, that's all we do at Kimi 3 " in 90-min workshop he explains why the smartest agent still fails - if you can't configure it correctly his one big idea: most people are still solving the old one "the real goal? we want K2 to help build K3 - without agent skills, that's impossible" watch & bookmark - then learn the article on best agent system ↓
codila@0xCodila

x.com/i/article/2077…

English
143
1.1K
8.4K
1.2M
Nando de Freitas retweetledi
Niels Rogge
Niels Rogge@NielsRogge·
The new Kimi K3 leverages LatentMoE, a technique developed by @nvidia in January What is LatentMoE? In LatentMoE, tokens are projected from the model hidden dimension 𝑑 into a smaller latent dimension ℓ for expert routing and computation, which reduces routed parameter loads and all-to-all traffic by a factor of 𝑑/ℓ. Learn more about it here: paperswithcode.co/paper/2601.180…
Niels Rogge tweet media
English
4
79
507
46.3K
Nando de Freitas retweetledi
Noam Brown
Noam Brown@polynoamial·
2023: LLMs struggle with 4th grade word problems 2024: LLMs can do high school math 2025: LLMs get a gold medal at the IMO Now, GPT-5.6 solves famous frontier math/stat questions. The IMO is today and 5.6 one-shotting a perfect score isn't even news. Where will we be next year?
Edgar Dobriban@EdgarDobriban

AI has helped resolve an important question in statistics. In the area of multiple hypothesis testing, the goal of controlling the false discovery rate (FDR) has been introduced in a seminal paper by Benjamini and Hochberg (1995). They also introduced a method (the Benjamini-Hochberg or BH method) and proved it controls the FDR. This method has been widely adopted in modern high-throughput science, including in genomics, astronomy, economics, etc. The paper has has garnered more than 130,000 citations to date. However Benjamini and Hochberg showed FDR control only when the data for the individual tests are *independent*. In practice, these data are often dependent; a good example is data on genetic variants due to linkage disequilibrium. Later work has focused on extending the validity of the BH procedure, e.g., to a form of positive dependence by Benjamini and Yekutieli (2001). The question of when the BH procedure controls the FDR has remained open. Over the last twenty years, many authors, including Reiner-Benaim (2007), Kim and van de Wiel (2008), Benjamini (2010), Sarkar (2023), Sarkar and Zhang (2025), have conjectured that the BH procedure controls the FDR for two-sided tests using any correlated Gaussian data. These authors have presented both theoretical and empirical evidence supporting, but not directly showing, the conjecture. With the help of AI (specifically GPT-5.6 Sol Pro), I have settled the question in the negative: The Benjamini-Hochberg procedure does *not* generally control the false discovery rate at the desired level for correlated two-sided Gaussian tests. This was done by exhibiting a Gaussian factor model for which, at a nominal level alpha=0.01, the false discovery rate is proved to be FDR>0.0104. There is a lot of interesting commentary to be made: 1. This result should be of interest to everybody in the field of statistics. Emmanuel Candes of Stanford University once called the false discovery rate and the Benjamini-Hochberg procedure "one of the two most important developments in statistics after 1950" (the other being James-Stein shrinkage). The present conjecture is probably the most central question about FDR/BH that was unresolved to date. 2. GPT-5.6 one-shot the problem after 90 minutes of reasoning, whereas with 5.5 I was not able to solve it even after iterating with multiple parallel agents for perhaps 20 hours. So the capability improvement is quite real. Exciting times to live in! 3. The argument is not especially surprising, but it does combine an asymptotic approach (standard for FDR analysis, see e.g., Genovese and Wasserman, Efron, etc) with a numerical certificate in a way that would be pretty non-standard in the field. Once we have the specific example, then straightforward simulations also support that the false discovery rate is indeed higher than the nominal value (see attached fig). 4. The current degree of violation over the nominal level is relatively small (0.104 vs 0.1). So the importance of this result is mainly conceptual. The practical implications remain to be determined. Overall, an exciting development! Preprint is available here (faculty.wharton.upenn.edu/wp-content/upl…) and will be on arxiv tonight; supporting code is here (github.com/dobriban/BH).

English
91
297
2.9K
434.7K
Frank Hutter
Frank Hutter@FrankRHutter·
Today it is official: our acquisition by SAP has closed. Prior Labs continues as an independent frontier AI lab - same team, same offices, same mission, open models, backed by €1B+. What changes? The scale and the planning horizon. 🧵 1/4
Frank Hutter tweet media
English
11
8
109
10.3K
Nando de Freitas retweetledi
Nando de Freitas retweetledi
Arena.ai
Arena.ai@arena·
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!
Arena.ai tweet media
Kimi.ai@Kimi_Moonshot

Meet Kimi K3

English
1.1K
4K
26.5K
19.4M
Nando de Freitas retweetledi
leo 🐾
leo 🐾@synthwavedd·
🚨 BREAKING: Kimi K3 benchmarks have released and it's ~Fable/Sol level
leo 🐾 tweet media
English
157
243
3.8K
544.1K
Nando de Freitas
Nando de Freitas@NandoDF·
Great point 3 by @RichardSocher —likely many bullshit jobs en.wikipedia.org/wiki/Bullshit_… won’t be impacted by AI either 🤔
Richard Socher@RichardSocher

There are at least a few reasons for why the incredible progress in AI hasn't yet resulted in a massive increase in GDP (some from Captain Obvious but number 3 is less intuitive to many smart people). 1. AI replaces some steps in complicated processes but companies are still doing mostly similar things and adoption and rethinking entire industries are slow. 2. Startups that replace everything (eg AI native law firm that is much cheaper) still need to ramp GTM, sales, etc But more importantly and surprising to many in Silicon Valley: 3. A huge chunk of the economy just does not require that much intelligence and won't materially change at its core with intelligence being abundant and cheap, eg. - tourism - people will want to see the pyramids with or without AI, - real estate - people want to live in hip and safe neighborhoods, AirBnB, rentals, etc. - luxury goods and status symbol bs, eg fancy handbags, clothes, overpriced cars, etc - food and large parts of the food supply chain (yes, I love AI for agriculture but crops and cows still need time to grow, etc) - sports and much of entertainment - oil drilling, tree growing/logging for construction, most of mining - etc If your existing economy depends mostly on these types of industries, AI won't impact it that much. But there's a whole new economy of knowledge work, research heavy industries, deep tech, online and digital work etc that will massively benefit and outgrow these existing industries.

English
3
2
21
12.5K
Nando de Freitas retweetledi
Ansys
Ansys@ANSYS·
Hydrogen is playing an increasingly important role across industries, but transport remains a critical engineering challenge. Ansys, part of @Synopsys researchers recently published new results, introducing a multiscale modeling framework for hydrogen diffusion in polycrystalline metals. By connecting atomic‑scale physics to engineering-scale simulations, this framework enables microstructure‑informed strategies to design safer, more reliable materials for hydrogen‑critical applications. Learn more at: bit.ly/4eY5Aja
English
0
2
15
2.3K
Nando de Freitas retweetledi