Zain

5K posts

Zain banner
Zain

Zain

@zainhas

I build and teach AI • AI/ML @togethercompute • EngSci ℕΨ/PhD @UofT • Previously: vector DBs, data scientist, lecturer & health tech founder • 🇺🇸🇨🇦🇵🇰

SF⇄Toronto Katılım Ağustos 2012
2K Takip Edilen6.3K Takipçiler
Sabitlenmiş Tweet
Zain
Zain@zainhas·
If you cannot explain something in simple terms, you don't understand it.
Zain tweet media
English
7
8
79
21.4K
Zain
Zain@zainhas·
#llm-variance" target="_blank" rel="nofollow noopener">danluu.com/ai-coding/#llm
ZXX
1
0
1
150
Zain
Zain@zainhas·
awesome writeup on ai coding makes a great point: all benchmark results should be presented as distributions instead of a nice and clean single number > distribution of performance across many runs > distribution of cost per task > distribution of wall clock times
Zain tweet media
English
2
4
11
593
Zain
Zain@zainhas·
pretty mindboggling finding in this new science paper even with 60 years of modern medical advancements 90% drugs still fail human trials, same number as in 1963
Zain tweet media
English
1
0
2
224
Zain
Zain@zainhas·
U Chicago law school is adding a in-person oral exam to combat LLM slop submissions I think LLMs are a tool that can 10x learning power for students if incentives are aligned and they're using them for the right reasons
Zain tweet media
English
1
1
5
344
Zain
Zain@zainhas·
@bleysg AGI has already been achieved internally!
English
0
0
1
12
Zain
Zain@zainhas·
LLM's be like:
Zain tweet media
English
5
0
10
927
Zain
Zain@zainhas·
@CompleteSkeptic Very insightful read, thanks for sharing! Can't wait to see the models from typesafe!
English
0
0
0
24
Zain
Zain@zainhas·
men lie, women lie, numbers never lie. the numbers:
GIF
English
0
0
7
372
Zain
Zain@zainhas·
Interesting paper on the impact of overthinking in reasoning models
Zain tweet media
English
1
0
2
241
Zain
Zain@zainhas·
this graph is interesting because effectively it says that Kimi K3 will be the best cyber model that people can use to perform/protect against cyber attacks b/c the blue line is gated and not generally available
Zain tweet media
English
0
4
9
498
Zain
Zain@zainhas·
oh wow was not expecting Claude Opus 5 to 3x performance of the next best frontier model on ARC AGI 3
Zain tweet media
English
2
0
2
259
Zain
Zain@zainhas·
if you take all the popular coding benchmarks and classify every task in them you can get a pretty good idea of the hills agentic models are currently climbing👇 tldr: > lots of implementing features + bug fixes + implement algo + perf optim > not enough security, devops, ML
Zain tweet media
English
2
2
19
1.1K
Zain
Zain@zainhas·
This is pretty foretelling, businesses will need: > Ways to amplify thier content to be picked up by LLMs + thier search api's > LLMs recommending options genuinely vs. ads that show up on the side of ChaGPT will be what you need to optimize for > Embed their content in high signal sources that get scraped for pre/post training so you can bias the model like reddit ex. Fascinating how quickly human behavior can change when a technology is 10x better than incumbents.
Hamel Husain@HamelHusain

35% of my business revenue in the last 60 days is from Claude / ChatGPT recommending me 🤯

English
0
0
1
518
Zain
Zain@zainhas·
does opus 5 have an overthinking problem? on FrontierCode main + extended set - Opus 5 @ med reasoning effort > high, xhigh and max pay to lose?
Zain tweet media
English
3
1
12
562