Alex Ratner

1.9K posts

Alex Ratner banner
Alex Ratner

Alex Ratner

@ajratner

@SnorkelAI @uwcse / prev @StanfordAILab – Interested in data management systems for machine learning, weak supervision, and impactful applications.

Menlo Park, CA Katılım Kasım 2013
707 Takip Edilen6.8K Takipçiler
Sabitlenmiş Tweet
Alex Ratner
Alex Ratner@ajratner·
This week we launched the Open Benchmarks Grant with a $3M initial commitment from @SnorkelAI + partner support from @huggingface @togethercompute @PrimeIntellect @PyTorch @harborframework & others, in order to close the evaluation gap in AI. Our ability to measure AI has been outpaced by our ability to develop it - and open benchmarks are one of several critical, complementary tools to fix this. We're particularly interested in novel benchmarks that push and probe the frontier along three key vectors: (1) Environment complexity --> E.g. complex, domain-specific context and tool/action spaces, human interaction, world modeling) (2) Autonomy horizon --> E.g. long horizon, non-stationary goals (3) Output complexity --> E.g. complex outputs with nuanced, rubric-based evaluation / reward signals Check out more detail + link to apply here! benchmarks.snorkel.ai
English
1
7
46
8.2K
Alex Ratner retweetledi
Snorkel AI
Snorkel AI@SnorkelAI·
@Ryfeus talk at @aiDotEngineer World’s Fair explores the next era of agent evaluation: turning production traces into repeatable simulations for testing agents in realistic environments. He covers how teams can compare models and harnesses, evaluate long-horizon tasks, uncover failure modes, and bring agent testing into the CI pipeline. Watch the talk: youtu.be/Ib5t2RLtxvM
YouTube video
YouTube
English
2
3
10
647
Alex Ratner
Alex Ratner@ajratner·
One of the main goals in AI has always been to get subject matter expertise into computers as efficiently as possible. Data is the most natural, universal transfer medium for subject matter expertise. But it's never really been about data :)
English
0
4
12
1.3K
Alex Ratner
Alex Ratner@ajratner·
All enterprises need some kind of specialized AI to be non-commodity in the AI era - and for that, they need a data flywheel. However: data flywheels are *not* built by just passively collecting user/agent interaction traces and then tuning on these. That is like telling a student to study only using their ungraded practice exams. Most of the alpha and effort in AI today is getting *ground truth* - i.e. the grading keys/grades/answers to the practice exams - that enable tuning/RL to work. This is all about *high quality data* in the format of rubrics (grading keys), human evals (grades), and/or "gold" traces (correct answers). Without this feedback signal, a data flywheel is fundamentally incomplete - and tuning on it will only reinforce inaccurate agent behavior. Specialized AI that works is built with data flywheels that are continuously *developed* with the right annotation and data!
English
5
14
49
3.1K
Alex Ratner retweetledi
Henry Kiss Ehrenberg
Henry Kiss Ehrenberg@henryehrenberg·
Opus 5 debuts at #2 on Senior SWE-bench, just behind the leader Fable 5. Opus 5 improves on performance-per-compute, achieving its score at 32% of the output tokens of Fable 5 on average. It also excels on senior-level bug investigations, taking the #1 spot across all models.
GIF
English
3
13
74
19K
Alex Ratner retweetledi
Snorkel AI
Snorkel AI@SnorkelAI·
Frontier-Bench is live. We're proud to have been a task author and data partner, with support through the Open Benchmarks Grants program, contributing to benchmark-wide testing and corrections and building the category taxonomy. Congrats to @frontierbench, @harborframework, @laudeinstitute, and the whole team on the launch!
Ryan Marten@ryan_marten

We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%

English
1
14
51
3.2K
Alex Ratner
Alex Ratner@ajratner·
Incredibly excited for a new frontier benchmark from the incredible TBench/Harbor team!! Proud that @SnorkelAI contributed as a task author and data partner, with support through our Open Benchmarks Grants program.
Ryan Marten@ryan_marten

We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%

English
5
10
40
2.3K
Alex Ratner retweetledi
Mark McQuade
Mark McQuade@MarkMcQuade·
A big day for Arcee and for American open-weight AI. Extremely proud to be partnering with the U.S. Department of Energy to help build the next generation of open models for scientific research.
Arcee.ai@arcee_ai

Today we are announcing a partnership with the Department of Energy to build Genesis-Science-1, an open model for scientific research. GS1 is an American open-weight AI model and governed research harness designed to complete scientific computing workflows while preserving a reproducible record of its work. This model will be shaped by the people who know scientific work inside and out. @ENERGY is opening a contributor program for researchers, laboratories, universities, companies, and nonprofits, and we're speaking with infrastructure partners who can add training or evaluation capacity. There’s still lots of work ahead, and we hope you’ll help us build in the open, starting with GS1.

English
4
6
25
2.5K
Alex Ratner retweetledi
Lucas Atkins
Lucas Atkins@latkins·
We've been quiet since Trinity-Large-Thinking came out in early April - but for good reason! I'm excited to finally share that not only have we @arcee_ai joined the DOE's Genesis Mission, but to also announce the development of Genesis-Science-1.
English
39
50
444
102.6K
Alex Ratner
Alex Ratner@ajratner·
Some notes on the idea that fine-tuning your own specialized agent is as easy as just capturing usage data. Setting up your own Data Flywheel is critical- but takes real data development work to close the loop, and correct for sample incompleteness and bias!
Alex Ratner@ajratner

x.com/i/article/2078…

English
1
2
16
1.4K
Mehrdad Farahmand
Mehrdad Farahmand@mifarion·
Given the current direction of AI, this post from over two years ago has aged remarkably well. Despite the constant hype cycles and changing narratives over the past two years, it turned out to be one of the best predictions of where LLMs and AI agents were headed. Well predicted, Alex Ratner @ajratner
Alex Ratner@ajratner

1/ Prediction: Everyone will soon be using foundation models (FMs) like GPT-4. However, they'll be using FMs trained on their own data & workloads: "GPT-You", not GPT-X Tl/dr: - Closed APIs aren't defensible - The durable moat is data - The last mile generates the real value

English
1
0
3
828
Alex Ratner
Alex Ratner@ajratner·
@HollyRoseFaith I remember!! Little did we both know how much I'd be pinging you for help down the line... :). Honored to get to continue working together so closely!
English
1
0
1
16
Holly Rose Faith
Holly Rose Faith@HollyRoseFaith·
@ajratner 💯 Our desks were next to each other when I first joined Greylock and you all were incubating the company out of the office. The best.
English
1
0
1
38
Alex Ratner retweetledi
Alex Ratner
Alex Ratner@ajratner·
Excited to see the GPT-5.6 launch using multiple recent @SnorkelAI Open Benchmarks Grants-backed benchmarks: - Agent's Last Exam - Terminal-Bench 2.1 - OSWorld 2.0 Open benchmarks are critical guideposts for advancing the science of both AI model *and* data/env development!
English
3
9
36
2.7K
Alex Ratner retweetledi
Logan Kilpatrick
Logan Kilpatrick@OfficialLoganK·
it’s surprising to me how many people seem to not understand that great models are built with super high quality curated data finding novel ways to create / get this data is a huge edge
English
395
210
4.8K
705K