Alex Sherstinsky

3.1K posts

Alex Sherstinsky

Alex Sherstinsky

@AlexSherstinsky

Scientist, Engineer, Manager, Musician, Athlete, Thespian, Father, Husband, Son, Friend

Silicon Valley Katılım Ağustos 2008
6.9K Takip Edilen815 Takipçiler
Alex Sherstinsky
Alex Sherstinsky@AlexSherstinsky·
@ZainHasan6 There are many to respect, the giants, and the disciples, who have toiled using the resources they had in many decades to get the field to its present moment of shine.
English
0
0
1
193
Alex Sherstinsky
Alex Sherstinsky@AlexSherstinsky·
@NathanZ721 I see -- nice! Raw implementation (i.e., not even using something like Unsloth); "Monophonic excerpts only" -- great -- will work for clarinet and/or violin. :) Huge thanks!
English
0
0
0
49
Nathan Zhao
Nathan Zhao@NathanZ721·
I trained an SVLM to do one thing really well: read music Sounds simple, but frontier models consume millions of tokens over hours to write scripts and analyze glyphs, only to hit <50% accuracy. My model takes 10 seconds on one GPU to hit 98% accuracy
Nathan Zhao tweet media
English
15
12
440
159.3K
Alex Sherstinsky
Alex Sherstinsky@AlexSherstinsky·
@rasbt Huge congratulations, @rasbt ! Unfortunately, I will miss the next PyTorch Conference (have to be at a different conference), but hopefully will get your new book signed by you in person at a future opportunity!
Alex Sherstinsky tweet media
English
1
0
2
253
Sebastian Raschka
Sebastian Raschka@rasbt·
After 18 months of writing, coding, and experimenting, Build a Reasoning Model (From Scratch) is finally out! My first copies just arrived! 📚 440 full-color pages. Inference scaling, reinforcement learning, and distillation from scratch.
Sebastian Raschka tweet media
English
279
560
6K
527.9K
Alex Sherstinsky
Alex Sherstinsky@AlexSherstinsky·
@maxmbeck Humangous congratulations! So excited to learn from your thesis report, papers, talks, and other materials!
English
0
0
1
93
Maximilian Beck
Maximilian Beck@maxmbeck·
👨‍🎓Last week, I successfully defended my PhD thesis - an incredibly exciting and rewarding milestone after 3.5 years of work on xLSTM: Recurrent Neural Network Architectures for Scalable and Efficient Large Language Models
Maximilian Beck tweet media
English
16
3
138
8.7K
Alex Sherstinsky
Alex Sherstinsky@AlexSherstinsky·
@askalphaxiv I have been wondering about trying this for years! The inspiration for trying this has been the Hopfield Network, which excludes self-connections. It always seemed like bias to me (of course the token will be most strongly correlated with itself!). So glad to see this result!
English
0
0
4
783
alphaXiv
alphaXiv@askalphaxiv·
"Exclusive Self Attention" This paper proposed Exclusive Self-Attention (XSA), which is a tiny two-line change that stops attention from looking at itself. This forces it to focus on the rest of the sequence, and can make transformers more effective! This improves the performance at long context at almost no extra cost.
alphaXiv tweet media
English
15
135
825
43.7K
Alex Sherstinsky
Alex Sherstinsky@AlexSherstinsky·
@zhaisf Thank you for sharing; is the full content available?
English
1
0
1
134
Shuangfei Zhai
Shuangfei Zhai@zhaisf·
Found this half page note I wrote ~6 years ago. Describes basically linear attention but half a year before the “Transformers are RNNs” paper came out. Sadly I didn’t take it too seriously at the time because I didn’t have any use cases for it and was also too busy with GANs.
Shuangfei Zhai tweet media
English
4
31
395
26.2K
Alex Sherstinsky
Alex Sherstinsky@AlexSherstinsky·
Claude Code has been 'ruminating,' 'deliberating,' 'pondering,' 'reflecting,' and 'strategizing.' Never once 'hallucinating.' Not even when it absolutely was. #ClaudeCode #AI
English
1
1
3
139
Yuchen Jin
Yuchen Jin@Yuchenj_UW·
Another xAI co-founder just left. This is the 3rd co-founder to leave this month. 6 out of 12 xAI co-founders are now gone. Building a startup is intense. Building an AGI lab is on another level. Sleeping 8 hours a day is impossible. Best of luck to Grok and the xAI team.
Toby Pohlen@TobyPhln

Three years, thousands of PRs, and a million jokes. Today was my last day @xai. To the team: you rock, no one burns the midnight oil better. To @elonmusk, thanks for taking me on board. I've learnt more about execution, speed, and product perfectionism than I could ever have imagined. Thanks for everything. My next priorities: sleep for more than 8h, write down all the things I've learnt (I have a list), and then think about what I want to do next. @gork wdyt?

English
39
39
1.1K
147.3K
Alex Sherstinsky
Alex Sherstinsky@AlexSherstinsky·
@CalcCon @Yuchenj_UW Charles, I am sorry! I have been consistently guilty! I know you are 100% right, there is just so much to do! :)
English
1
0
0
6
Alex Sherstinsky
Alex Sherstinsky@AlexSherstinsky·
The encouraging part is that regular up-skilling, which must now be part of a standard development cycle, is also more readily accessible -- as self-guided, AI-assisted, enjoyable part of an engineer's work.
English
0
0
0
35
Alex Sherstinsky
Alex Sherstinsky@AlexSherstinsky·
This happens much more frequently than before, because the pace of development is so much greater.
English
1
0
0
38
Alex Sherstinsky
Alex Sherstinsky@AlexSherstinsky·
My reflections on using AI-assisted software development tools thus far is that without an experienced operator, any such tool quickly gets out of control and generates mountains of error-prone, unmaintainable "code" (in quotes, because there is no quality in that "code").
English
1
0
2
57
Yung-Sung Chuang
Yung-Sung Chuang@YungSungChuang·
🎓 Life updates: I defended my PhD last week at MIT. Now I am joining OpenAI to continue my research on building trustworthy LLMs! Thanks my advisor James Glass and my committee @yoonrkim, @jacobandreas and all my friends for continual supports throughout this journey!
Yung-Sung Chuang tweet mediaYung-Sung Chuang tweet mediaYung-Sung Chuang tweet media
English
49
46
1.1K
62.3K