driss guessous

553 posts

driss guessous

driss guessous

@drisspg

bytes and nuggets @pytorch https://t.co/gWVJmW741f

Katılım Aralık 2023
267 Takip Edilen1.6K Takipçiler
driss guessous
driss guessous@drisspg·
@kalomaze Hmmm open an issue with what you got would be fun to figure out what’s going on
English
0
0
0
168
kalomaze
kalomaze@kalomaze·
bf16 MoE routers b like
kalomaze tweet media
English
3
2
72
6.9K
driss guessous
driss guessous@drisspg·
Honestly the most fun part was that the lil intra kernel profiling was actually helpful here and led me to try different ideas after looking at pipeline dependencies
driss guessous tweet media
English
0
0
2
241
driss guessous
driss guessous@drisspg·
Also forgot to post the receipts: github.com/drisspg/transf… I was claudxing here playing around with ideas. M=1, pure memory bound -> do we really need tcgen? Not using cublaslt autotuning and this is against pt nightly with 13.2 on B200. Im sure folks can do much better :)
driss guessous tweet media
English
2
0
8
303
driss guessous
driss guessous@drisspg·
"ohh I beat cublas by 2% by implementing Hilbert curves" The trick to beating cublas is to go where they aint looking; And pray the next toolkit hasn't caught back up (it will)
driss guessous tweet media
English
10
5
113
12.4K
Rémi Ouazan
Rémi Ouazan@remi_or_·
@drisspg Wow, I missed this! Will definitely try it out. For bidirectional, I admit I haven’t looked into it too much, not a lot of model run on it these days… thanks for the tips!
English
1
0
1
116
Rémi Ouazan
Rémi Ouazan@remi_or_·
There is no point running a model with SDPA attention. Running an LLM without flash is like nerfing your own hardware. 3 years ago, sure, it was hard to setup, but not anymore. Just look at the impact of flash on throughput... and im not even going to talk about CB... 🚀
Rémi Ouazan tweet media
English
3
5
37
4.2K
driss guessous
driss guessous@drisspg·
@remi_or_ Makes sense, the Api forsure is missing features needed for autoregressive decoding. We did add docs.pytorch.org/docs/2.13/nn.a… which is much more aligned, supports FA3/FA4 and cudnn. For dense bidirectional attention SDPA is often quite hard to beat I find. Regardless, happy coding :)
English
1
0
4
127
Rémi Ouazan
Rémi Ouazan@remi_or_·
@drisspg It means that it’s often a better idea to use flash attention (through the package `flash_attn` or `kernels`) than `torch.sdpa` ! Although it can fall back to flash using it explicitly is better IMO
English
1
0
1
201
Kimi.ai
Kimi.ai@Kimi_Moonshot·
ZXX
696
1.1K
15.6K
3.2M
Lucas Beyer (bl16)
Lucas Beyer (bl16)@giffmana·
@drisspg Even the chatbot is not good daily use, i a/b test with chatgpt a lot and all that's missing is the nice auto memory from past chats and a bunch if UI features.
English
1
0
2
1.2K
eli
eli@_seemethere·
@drisspg Determined not a shill.
eli tweet media
English
1
0
2
81
tender
tender@tenderizzation·
@drisspg is this with or without cublas autotuning
English
2
0
2
502
driss guessous
driss guessous@drisspg·
@GoonGarrett I think it just gives it time to ferment slowly, and build a more "complex flavor" not to sound too douchey. The one I let go for like 4 days deff tasted the most unique
English
0
0
1
16
Garrett Goon
Garrett Goon@GoonGarrett·
@drisspg What does the overnight refrigeration add? Surprised by that step
English
1
0
0
19
driss guessous
driss guessous@drisspg·
If PyTorch added cutedsl as a required Runtime dependency for our Cuda wheels, how much would this mess you up? Would >= X.y for latest X.y at time of pt release work?
English
6
2
44
7.1K
driss guessous
driss guessous@drisspg·
@tenderizzation Never attribute to malice that which is adequately explained by stupidity. To my thought
English
0
0
1
168
tender
tender@tenderizzation·
@drisspg it can’t be a “i’m breaking up with you next week” ahh relationship without the next week part
English
1
0
9
385