J 🦀

1K posts

J 🦀 banner
J 🦀

J 🦀

@olorosia

https://t.co/0N0S4yrLZn

Katılım Eylül 2022
1.2K Takip Edilen348 Takipçiler
Andrew Ho
Andrew Ho@andrewho03·
Today is my last day at @OpenAI. I'm glad to have spent the last eight months of my life working here! I'm starting a new company focused on the production of high-quality reinforcement learning datasets: 1. The generalization ability of LLMs is clearly very poor, with "spiky" capabilities even in areas that have received tremendous amounts of investment and attention. For example, despite multiple years with tens (if not hundreds) of billions invested, even coding capabilities don't demonstrate "generality" -- even if every model can solve Codeforces questions or port C++ to Rust better than I can, I still have to manually "deslop" pull requests. 2. The vast majority of economically productive capabilities are not well represented in existing data offerings. First, there's a certain art to the design of an RL dataset which most vendors, not having upstreamed data into large training runs themselves, don't really understand. Second, and more importantly, most work is highly contextual and not easily encoded into a gradable environment; even if we can observe a "golden path" taken by a human which we believe to be good, it's challenging to understand whether alternate, counterfactual paths produce good or bad outcomes. The basic premise here is that I have a clear understanding of what labs need/want, having explicitly been on the other side and having been involved at every level from procurement all the way through training, and I'm able to provide it. I also believe that data needs will grow tremendously in the coming years, especially as frontier labs face increasing pressure toward profitability, and that they won't get the relevant capabilities "for free" through scaling alone; instead, they'll need to spend >$100B on precise, well-targeted data acquisition. Our first products will be focused on biology and statistical reasoning: 1. First, datasets that address long-horizon scientific reasoning, drawing on my work on GeneBench-Pro with @jeremyli__. Frontier models are still unable to reliably execute "messy" data analyses that require judgment, exploration, and adaptive revision (GB-Pro passrate on GPT-5.6 Sol scarcely exceeds 30%); to address this, we have the ability to generate thousands of high-quality problems with known ground truths which can be reliably graded. (In contrast, most existing RL data for bioinformatics is either massively over- or under-specified, and will probably break your model when you train on it.) Moving the "reliability gap" from 30% to >90% is obviously required for scientific acceleration, and -- despite my skepticism about generalization of RL -- is one of the *most promising datasets* conceivable when it comes to yielding generalization benefits for models' overall reasoning capabilities. 2. Second, datasets that address capabilities relevant to day-to-day workflows. Imagine a scientist snapping a picture of some experimental process or result -- say, a cell culture plate or a Western blot -- and asking Claude a question. Frontier models remain quite bad at these questions, especially those with multimodal components. But they're obviously required for acceleration of scientific discovery; before we can dream about automating science, we have to begin with shoring up these basic, generalist capabilities. Beyond these two, we hope to expand to adjacent fields (chemistry, materials science, etc.), and then even further into fields with more direct economic applicability like healthcare and white-collar office work. I strongly encourage labs with data needs to reach out. We offer industry-standard pricing and terms, and like I said -- I know how this process works, what good data looks like, and how to demonstrate to you, convincingly, that you'll be able to upstream our data into your training processes without issue. My DMs are open!
English
245
291
5.6K
1.2M
J 🦀
J 🦀@olorosia·
@exoaursen Is there a smart contract programming model that isnt retarded?
English
1
0
3
1.2K
Aursen
Aursen@exoaursen·
Another Solana program rekt. Another six-figure audit. Still no one willing to say the quiet part: this programming model is fucking retarded
English
12
1
82
18.4K
J 🦀
J 🦀@olorosia·
I give up on the job search. To the trenches we go
English
1
0
2
126
J 🦀
J 🦀@olorosia·
@gakonst Literally true of any software
English
0
0
0
173
Georgios Konstantopoulos
Georgios Konstantopoulos@gakonst·
codex is awesome, but it's almost 1M Rust LoC - i suspect it could be better if there was a rewrite from scratch that didn't care about backwards compat and was focused only on latest generation of models
English
19
0
197
43.2K
J 🦀
J 🦀@olorosia·
What French attack? Where is mbappe???
English
0
0
2
127
Taelin
Taelin@VictorTaelin·
... honestly, I'm too naive to be a founder. People are raising serious money (money I never raised) to create startups based on algorithms *I invented*, and open-sourced. That's going to be an expensive lesson. Taking a morning to appreciate people are just evil and it is all a game of interests, after all
English
171
32
2.4K
139K
Bryan Johnson
Bryan Johnson@bryan_johnson·
If you get cancer, your risk of Alzheimer's drops by about a third. If you get Alzheimer's, your risk of cancer drops by about half. The body playing tug of war with itself. Two diseases pulling against each other, using the same cellular mechanism in opposite directions.
Bryan Johnson tweet media
English
341
300
6.3K
616.6K
J 🦀
J 🦀@olorosia·
"Spawning a codex worker to apply the audit — mechanical edit, so it gets a mid-tier (codex:gpt-5.5:medium, the cheap-sweep case #34 was built for):" 🤌
English
0
0
1
64
J 🦀
J 🦀@olorosia·
Why the ref have shoulder pads on tho
English
0
0
1
122
J 🦀
J 🦀@olorosia·
@frabigol Literally no player has a higher peak than r9. Stats wise many surpass him though, injuries
English
0
0
0
15
Fran
Fran@frabigol·
Ya podemos decir esto o algún gordo obsesionado con la nostalgia se enoja?
Fran tweet mediaFran tweet mediaFran tweet media
Español
1.9K
683
22.5K
1.5M
J 🦀 retweetledi
Andrew Côté
Andrew Côté@Andercot·
Technology is downstream of science fiction
English
9
15
103
7.6K
J 🦀
J 🦀@olorosia·
@redacted_noah Its also up there in terms of stress if youre on a timeline
English
0
0
2
51
Noah
Noah@redacted_noah·
There is nothing that tops the high of learning a new field/skillset. The best part of joining any new company is the feeling of going from "I'm overwhelmed, how does any of this work?" to a deep, nuanced understanding. Coding is fun, but lifelong learning is the prize.
English
10
1
77
2.9K
J 🦀
J 🦀@olorosia·
Embrace the degen within
English
0
0
0
78
J 🦀
J 🦀@olorosia·
@mfranz_on What's "serious hardware stuff"?
English
1
0
1
21
Marco Franzon
Marco Franzon@mfranz_on·
I need a garage-lab to make serious hardware stuff.
English
16
2
58
4.4K
FreeSolDev
FreeSolDev@FreeSolGamesDev·
@joshyote That's the question of the centurary,but it's the one thing i could think of that sol prob cant do without externals right now
English
2
0
3
80
josh
josh@joshyote·
What’s a moonshot that could never be built on Solana?
English
26
1
56
4.5K
Paras Chopra
Paras Chopra@paraschopra·
Today I learned that in English: words starting with "gl-" often relate to light: gleam, glint, glance, glare, glaze, glimpse, glitter, gloss, glower, glisten, glaze, glitz, glazed. words starting with "sn-" often relate to nose: sniff, sneeze, snuff, snore, snort, snot.
English
354
2.4K
27.7K
2.5M