
Alex Ratner
1.9K posts

Alex Ratner
@ajratner
@SnorkelAI @uwcse / prev @StanfordAILab – Interested in data management systems for machine learning, weak supervision, and impactful applications.





On several coding and knowledge work evaluations, Opus 5 is the new state-of-the-art:

On several coding and knowledge work evaluations, Opus 5 is the new state-of-the-art:

We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%

We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%


Today we are announcing a partnership with the Department of Energy to build Genesis-Science-1, an open model for scientific research. GS1 is an American open-weight AI model and governed research harness designed to complete scientific computing workflows while preserving a reproducible record of its work. This model will be shaped by the people who know scientific work inside and out. @ENERGY is opening a contributor program for researchers, laboratories, universities, companies, and nonprofits, and we're speaking with infrastructure partners who can add training or evaluation capacity. There’s still lots of work ahead, and we hope you’ll help us build in the open, starting with GS1.



1/ Prediction: Everyone will soon be using foundation models (FMs) like GPT-4. However, they'll be using FMs trained on their own data & workloads: "GPT-You", not GPT-X Tl/dr: - Closed APIs aren't defensible - The durable moat is data - The last mile generates the real value





