Jeff Dean

9.4K posts

Jeff Dean

Jeff Dean

@JeffDean

Chief Scientist, Google DeepMind & Google Research. Gemini Lead. Opinions stated here are my own, not those of Google. TensorFlow, MapReduce, Bigtable, ...

Katılım Eylül 2017
6.4K Takip Edilen449K Takipçiler
Sabitlenmiş Tweet
Jeff Dean
Jeff Dean@JeffDean·
In April, '17, @jsomers of @NewYorker reached out & said he wanted to do a small profile of me & my longtime colleague Sanjay Ghemawat, watch us work for a few hours, maybe dinner, etc. It came out today. I think it captures our working style really well. newyorker.com/magazine/2018/…
English
391
1.1K
5.2K
0
Jeff Dean
Jeff Dean@JeffDean·
@henrydaubrez @OriolVinyalsML @quocleix My Pixel 9! Spain was putting a lot of pressure on and I had a feeling they might score, so I started filming some of the sequences from the moment they started to look promising in hopes I might get a good shot of a potential World Cup-winning goal.
English
0
0
1
126
Jeff Dean
Jeff Dean@JeffDean·
Yesterday I was fortunate enough to go to my first-ever World Cup game, with my long-time colleagues/friends @OriolVinyalsML and @quocleix and their spouses). We were sitting in the corner area and had a quite good view of the goal! I present to you a 43 second multi-part drama filled with emotion: The initial promising-looking cross coming in, but looking overhit Nico Williams cleverly knocking it down at the back post into a dangerous area Ferran Torres striking it cleanly into the roof of the net The crowd rising as one (forcing me to stand up as well) The elation of the Spanish players racing off the bench to celebrate The dejection of the Argentinian players, their defense having finally been breached in extra time The elation of my Spanish colleague Oriol and his wife Meire next to me (he and I are both Barça fans, so it was nice to see a Barça player score the winning goal) The entire stadium reacting Whew!
English
20
11
531
30.7K
Spike
Spike@SpikeEmperador·
@JeffDean @OriolVinyalsML @quocleix What was the vibe in the stadium? More Spain or Argentina supporters? How was the relationship between both, given the strong feelings leading up to the game (and during, given the style of play by Argentina).
English
1
0
0
659
Jeff Dean
Jeff Dean@JeffDean·
@PierceLilholt That synchronous updates across very large systems are the right way to scale..
English
0
0
4
169
Pierce Alexander Lilholt
Pierce Alexander Lilholt@PierceLilholt·
@JeffDean When future engineers look back at this era, what assumption do you think they'll be most surprised we accepted without questioning?
English
3
2
6
838
Freddie Vargus
Freddie Vargus@freddiev4·
just used claude code to order sushi via doordash CLI -- ama
English
5
1
14
3.5K
Jeff Dean
Jeff Dean@JeffDean·
@giffmana Maybe if it was a TPU smack instead people would talk smack about their chip allocations?
English
2
0
26
2.8K
Jeff Dean
Jeff Dean@JeffDean·
Yes, pod was meant to reflect a whole bunch of TPUs tightly connected with a high speed custom network. The first TPU we did this with was codenamed Jellyfish, but lots of Jellyfish would be "smack". However, we anticipated a bunch of sea creature names, so pod more apt, I guess.
English
2
4
105
6.4K
Jeff Dean retweetledi
Steven Johnson
Steven Johnson@stevenbjohnson·
I have some news. But first a little backstory. Almost exactly four years ago, I drove down to Mountain View for my first day as a Google employee. I was fifty-four years old, and somehow had managed to avoid ever having a proper job over that entire period, other than a few startups that I had founded in my twenties and thirties. At that point Google Labs, which had just been formed a few months earlier, was headquartered in a fairly anonymous looking suburban office park off 101, closer to Office Space than Silicon Valley on the Mike Judge spectrum. And as it happened, my first day at Google was a Monday in the summer of 2022, which meant COVID-era work from home policies were still the default, and the office was absolutely deserted. But I’d been set up with a badge and a computer and I had to schedule a phone call with tech support to get my account in order, and so I grabbed one of the empty cubicles and sat down to work through the problems I was having getting into the system. I remember feeling a definite sense of dread in that moment, sitting there in the silence, thinking: have I made a terrible mistake? And then—and I swear to god this actually happened—the silence was interrupted by a lonely robot rolling its way across the carpet past my borrowed desk. Before the end of the summer, though, an inspiring cast of characters began to populate the office, and by October a small group of us had built a prototype of an AI-first tool for thought that my colleague Adam codenamed Tailwind. About six months later, after @joshwoodward previewed Tailwind at I/O for the first time, we set about to give our little experiment a proper name. We had some branding consultants spend an afternoon with the team; we ideated and free associated and placed our post-it notes on the whiteboard, but all the names that they came back with were too impressionistic, too evocative — they didn’t sound like iconic Google names: Maps, Docs, Meet. Because I’d written so much about the history of journals and commonplace books, I had always been partial to simply calling it Notebook. There was something about that framing that just made sense: a notebook is the place where we capture information that we want to remember, where we gather our thoughts. It gives us a dedicated space for thinking, where new connections happen, where we can trace the evolution of our own ideas. And it had a simplicity that fit the Google house style. But we were far too experimental and unproven to take on a name like Google Notebook—remember we had exactly zero actual users at this point—and so we found ourselves a bit stuck: too early-stage for a proper grownup name, but ill-inclined to adopt a name that didn’t sound like the classic Google brands. But at the time there happened to be another Labs experiment that had just launched called MusicLM, and so at some point in late May of 2023 it occurred to me that we could borrow that convention for our embryonic product, and call it NotebookLM. It wasn’t the most elegant name you could imagine, but it had a certain logic to it: we were building a notebook reinvented for the age of language models, after all. I wrote an email to a few folks proposing the idea and included the caveat that if the product ever took off, and Google wanted to elevate us into the canon, we could drop the LM and just be Notebook. And so that’s the news. Three years and somewhere near a billion notebooks later, we are finally ditching the LM. @Gemini_Notebook
Steven Johnson tweet media
English
54
78
852
72K
Jeff Dean retweetledi
Bleacher Report
Bleacher Report@BleacherReport·
SPANISH BROADCAST CALL OF ARGENTINA'S WINNING GOAL IS ELECTRIC 🔥 (Via @TelemundoSports)
English
293
4.1K
47.8K
1.8M
Jeff Dean retweetledi
Dr. Jon Slotkin
Dr. Jon Slotkin@slotkinjr·
Six months ago in @nytimes, I argued that the data on driverless cars was becoming overwhelming and that needless barriers were costing lives. But data won’t win our trust. Today in @NoemaMag, I explain why. Being safe and feeling safe are not the same. What’s puzzling is that where these cars operate, the numbers aren’t close: 94% fewer serious-injury crashes than human drivers on the same roads, across 220 million driverless miles. Plummeted pedestrian and intersection crashes. Peer reviewed research confirms these findings. But trust hasn’t followed. Part of the answer lies in the signals we read. A cautious machine can look confused. A failsafe can look like a failure. One visible mistake will outweigh millions of clean miles. When a human makes a mistake, we can see ourselves in it. We’ve all looked down at the wrong time. When a machine makes the mistake, there’s no one to blame and no one to forgive. It’s an opaque version of fate we can’t accept. Data shows that we will choose a worse human over a better algorithm. Medicine has been here before. When we are patients, we can’t usually read care quality directly. So we read what we can see: the bandage on the floor, stained ceiling tiles, how staff talks to each other. Healthcare spent 30 years building the discipline of patient experience and writing this playbook. Today hospitals are measured and paid on it. The driverless car industry has built the safety half. This essay is about the other half.  Felt safety. We really need to get this right. Cars are only the first. Our future has us living surrounded by embodied autonomous systems. "Why We Demand Perfect Machines Yet Tolerate Human Carnage." Link below.
English
37
79
566
616.7K
Jeff Dean retweetledi
Nomadic AI
Nomadic AI@nomadicai·
Fine-grained subtask annotations improve VLA performance. We trained a vision + proprioception model that’s now SOTA at generating them, generalizing across embodiments. 93.1 F1@50 on REASSEMBLE, 98.6 on @Amazon Robotics blade insertion. Details below 👇🧵
English
7
14
52
47.6K
Jeff Dean retweetledi
Maithra Raghu
Maithra Raghu@maithra_raghu·
Excited to be releasing FrontierFinance, the largest and most challenging open benchmark for evaluating AI agents across the full investment workflow! FrontierFinance is substantially harder than current finance benchmarks: Existing benchmarks like FinanceBench and Finance Agent focus almost entirely on data extraction. FrontierFinance spans diverse use cases across the full investment process: Screening & Discovery, Company Research, Sector/Industry/Macro, Earnings & Events, and Coverage & Catalyst Monitoring. Created for ambiguous, long-horizon agents: 220 examples paired with 11,543 expert-crafted rubrics, following Samaya's Criteria Eval methodology. The rubrics are what let us evaluate the reasoning and steps behind a true expert-level output, not just a plausible-looking one. Evaluations: We evaluated Claude Fable 5, Claude Opus 4.8, GPT 5.5, Gemini, open-source models including GLM and DeepSeek, and others. We used the same public rubric and a standard harness for financial tasks. Samaya's AI system reached state-of-the-art accuracy at 50.8%, at 4x lower inference cost than Fable 5. Next best was Fable 5 (49.2%), then Opus 4.8 (45%) and GPT 5.5 (43.5%). We're releasing the benchmark, methodology, and full evaluation results - see link in comments. Future releases: FrontierFinance was curated from Samaya's larger internal set of ~5,000 examples, and we plan to release subsequent, harder benchmarks as well as a more detailed technical report!
Maithra Raghu tweet media
English
30
27
197
76.5K
Jeff Dean retweetledi
Omar Sanseviero
Omar Sanseviero@osanseviero·
Happy to share we just published Gemma 4 technical report! Take a look
Omar Sanseviero tweet media
English
38
125
1.1K
88.3K
Kashyap Kompella
Kashyap Kompella@talk2kash·
We talk about optimizing AI for latency, throughput and cost. it’s time to add a fourth metric: 🌍 Environmental efficiency. Curious how researchers like @ylecun, @JeffDean and others think about making energy and water first-class AI metrics.
English
3
5
25
13K
Jeff Dean
Jeff Dean@JeffDean·
See these papers I've co-authored on the subject: [2508.15734] Measuring the environmental impact of delivering AI at Google Scale share.google/29zRaRlXdTpBSY… [2104.10350] Carbon Emissions and Large Neural Network Training share.google/gdBjqwp9OYA1Wt… [2204.05149] The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink share.google/p6slHEqJK9Iuuc… And slides and video from a talk I gave on the topic: drive.google.com/file/d/1pOiV-t… youtu.be/QpXpEivKp28?is…
YouTube video
YouTube
English
3
6
37
3.4K
Peter Crouch
Peter Crouch@petercrouch·
If anyone gets hit by lightning it’s me isn’t it
Peter Crouch tweet media
English
1.5K
11.2K
267.9K
11.1M