ali

626 posts

ali banner
ali

ali

@waterloo_intern

ml research, kernels, and the occasional peer-reviewed shitpost inference @baseten || eng @uwaterloo

San Francisco Katılım Ekim 2024
89 Takip Edilen30.3K Takipçiler
Thomas Børstad
Thomas Børstad@tboerstad·
@waterloo_intern Haha, now I understand how you were always available in my EU timezone. Keep cooking Ali!! 🐐
English
1
0
0
72
ali
ali@waterloo_intern·
holy shit i'm going to print this and stick it on all my gym fits and frame it and hang it on my wall and look at it every time i take creatine and show it to every one i meet holy fuck bryan johnson my goat i love you
Bryan Johnson@bryan_johnson

@waterloo_intern @ghostrodeo0 incredible project ali and your intense focus lowered pro inflammatory genes expression and extended you life

English
3
0
112
16.8K
ali
ali@waterloo_intern·
@sohailmo alhamdulillah, i see you grinding too, waiting for the next banger from you 🫶
English
1
0
2
116
ali
ali@waterloo_intern·
@dhruvbhatia0 so you're saying we draw the line at 3.6 billion times
English
0
0
5
159
dhruv bhatia
dhruv bhatia@dhruvbhatia0·
@waterloo_intern Since he didn't do it I will. The inner lining of plastic in those cans is harmful, and if you do it 3.7 billion times more in your lifetime your balls will be 1% plastic by weight.
English
1
0
3
283
ali
ali@waterloo_intern·
@easonjiangcc super cool, do you mind putting up a version translated to english? i'd love to read the extra notes you have here!
English
1
0
3
374
ali
ali@waterloo_intern·
@Hi_Mrinal That's so funny. You can't blame me, honestly. You both have very similar profile pictures haha
English
1
0
6
167
ali
ali@waterloo_intern·
@IKorovinsky you're training that rn bro wdym
English
0
0
2
869
ali
ali@waterloo_intern·
@hakmgpt thank you man haha
English
0
0
0
571
ali
ali@waterloo_intern·
I spent 48 hours with the Kimi K3 modeling code. It took: - 650 mg of caffeine (mandatory) - 40 cans of LaCroix (optional... world record (?)) - 8 papers - 6 months off my lifespan Finally grokked the entire lineage of Kimi K3 and how we got here... every single step, since 2019 GPT-2
ali@waterloo_intern

x.com/i/article/2077…

English
248
728
9.7K
2M
ali
ali@waterloo_intern·
@geetkhosla means a lot, thanks!
English
0
0
0
274
ali
ali@waterloo_intern·
@saranormous hahahaha "intern" technically still true for two more weeks
English
0
0
4
374
ali
ali@waterloo_intern·
@baseten capacity-mogged the competition
English
0
0
0
288
Baseten
Baseten@baseten·
"The central change is not scale alone. Each architectural step changes what the model stores, how it updates that state, or how it retrieves information that a fixed-size state cannot preserve. Kimi K3 combines constant-state recurrent memory, periodic softmax retrieval, sparse expert capacity, and selective depth-wise residual access. The result is a system that spends additional capacity where it has a specific functional role."
ali@waterloo_intern

x.com/i/article/2077…

English
1
4
65
6.9K
Trivikrama M
Trivikrama M@trivikrama_m·
@waterloo_intern this is brilliant! thank you for putting this together. Glad I spent 30 mins reading this right away, otherwise this gem would have been left rotting in my bookmarks
English
1
0
4
4.4K
ali
ali@waterloo_intern·
@thepatelkush Attention is all you need Linear Attention Gated Delta Attention Kimi Linear Kimi Residual Attention preferably in that order
English
0
1
7
3.4K
Kush Patel
Kush Patel@thepatelkush·
@waterloo_intern Excellent writeup, where can i read foundation of this in details? What papers you recommend to start with.
English
1
0
0
4K
ali
ali@waterloo_intern·
@zycxfyhxzh you flatter me too much
English
0
0
1
287
zycxfyh
zycxfyh@zycxfyhxzh·
This is an absolutely outstanding piece of work. A crystal-clear, beautifully structured walkthrough of seven years of architectural evolution—from GPT-2 all the way to KimiK3. The way you break down linear attention, DeltaNet, gated variants, MLA, MoE, and AttnRes is both rigorous and incredibly readable. Truly a masterclass in technical writing. Massive respect. This is the kind of worklog that deserves to be bookmarked, shared, and studied.
ali@waterloo_intern

x.com/i/article/2077…

English
1
0
1
514
ali
ali@waterloo_intern·
@nhuber short answer is no, its not just scale, many other optimizations in place, primarily smarter ways to do attention...but scale is obviously needed and helps
English
0
0
0
928
Nick Huber
Nick Huber@nhuber·
@waterloo_intern Is there a short answer somewhere to the “Is it just scale?” question I missed. Would be helpful for scanning to have some kind of summary/answer to that tantalizing question.
English
1
0
0
1K
ali
ali@waterloo_intern·
@hanul93 Of course, sounds good!
English
1
0
4
2.5K
Kei Choi
Kei Choi@hanul93·
Hi! I really enjoyed reading your article—it's one of the clearest explanations of the progression from GPT-2 to Kimi K3 that I've seen. Many Korean AI engineers and researchers would benefit from it, so I'd like to translate it into Korean and share it with the Korean community. Of course, I will clearly credit you, include the original article link, and state that it is a Korean translation of your work. Would that be okay with you? Thank you!
English
1
0
6
3.1K