Cam

22 posts

Cam banner
Cam

Cam

@RealCamBrown

I work in both CPU & SoC design at @AmpereComputing. In addition to computers, I enjoy baseball, music, and going up and down mountains

Portland, OR Katılım Mart 2020
1.1K Takip Edilen118 Takipçiler
Cam
Cam@RealCamBrown·
Don't try to make assumptions about what is vs. is not a microadjustment. Apply something like k-means per pitch and let the data speak. Cross-correlate with the location of the catcher's mitt to add a validation layer to the clustering (i.e. you should not count up&in when the catcher is set up low&out as a member of an up&in cluster). Spitballing here.
English
0
0
1
40
DT K
DT K@tomdoyo·
I spent 100s of hours grinding out manual labels this summer so that we can measure command! How? By measuring the distance between the catcher's target VS where the ball landed. There's more to come, but here's 10 pitches from the All-Star Game: (📽 @MLBONFOX)
English
98
153
2.6K
631K
Cam
Cam@RealCamBrown·
@tomdoyo @8_Maxx @MLBONFOX Use aggregate data to infer intent. Look for clusters for a given pitch (maybe a guy throws his curveball for a strike and also to try and elicit chases, in separate situations) and treat some radius around the center of the cluster as accurate execution.
English
1
0
2
225
DT K
DT K@tomdoyo·
@8_Maxx @MLBONFOX Yeah definitely. It's a huge problem I haven't solved yet, any ideas?
English
21
0
67
10.7K
Cam
Cam@RealCamBrown·
@elliotarledge it's a fun post but thinking you're putting 24 LPDDR PHYs on a 144 mm^2 die is a pipe dream
English
0
0
2
246
Elliot Arledge
Elliot Arledge@elliotarledge·
your homework for today is understanding this
Big Boss@0xBADB01E

First you have to understand that modern LLM inference already disaggregates weights as models outgrew single chips years ago. You shard either by layer (pipeline parallelism) or by slicing every layer (tensor parallelism), and the two do very different things. As an example, let’s look at Llama 3.3. It has 70B of weights and at FP8 that’s 70 GB of memory which is enough to fit on a single H100. Now that H100 has 3.35 TB/s of HBM, so the fastest it can ever decode for one user is 70/3.35 ≈ 21 ms/token or ~48 tok/s while using under 1% of its FLOPs. Now if we pipeline it across 8 chips: each chip holds ~8.75 GB, which means it only needs 1/8th the bandwidth and 1/8th the FLOPs to sustain the same aggregate throughput. Now crucially the token/sec a user gets is limited by the amount of data that crosses the link. In current LLMs all that is a small amount of activations for LLama 3.3 it’s ~8 KB per token…. Yes, you read that right it’s 8 KILOBYTES we are sending over a <900 GB/s link. That’s only 9 ns of serialization time but the overhead of 224G PAM4 SerDes adds ~100 ns per link traversal with RS-FEC which is 11x longer than the payload itself. And then you have the NVSwitch adding ~300 ns per hop and you need to pay twice. That’s ~600 ns of just hardware latency wrapped around 9 ns of data making a 98% tax before software even shows up. Then NCCL’s collective stack turns 600 ns into 10-20+ us… all to move 8 kilobytes lol. For comparison 8 KB serializes over 10 Gigabit Ethernet NRZ, in just 6.6 us. Pipeline parallelism however doesn’t make a single user faster as the token still needs to visit every layer in the sequence, so per-user speed is still weights / per-chip bandwidth. To get more speed per user token you need to use tensor parallelism and have all the chips work on the same layer simultaneously. TP costs you 2 all reduce OPs per layer, 160 per token on llama 3, that’s still kilobytes of traffic but with NVLink overhead it’s a massive tax and why pipeline parallelism on most models still gives more interactivity per user. However, this gives you a huge latency lever to pull that scales tokens per second with interconnect speed instead of memory BW. The clever amongst you might have also realized that sharding doesn’t just cut memory bandwidth per chip it also cuts FLOPs per chip and is why we have such bad MFU on decode. So once you’ve sized the link for the memory, you need to size the compute for it too. This is called “balancing the pipeline”, and currently no shipping chip does it because they were all designed as standalone monsters. Remember Tokens/sec = ~aggregate memory BW / bytes touched per token. At batch 64 in FP4 you need ~250 FLOPs per byte, and Blackwell ships 1,250. Provisioned 5x more than the narrow pipe of HBM. Nobody saturates shit cause they are all building around HBM. So now it all comes full circle. Parallelism reduces memory bw pressure and thus FLOPs but increases interconnect latency pressure. Despite having HBM and GigaSERDES we aren’t actually doing more work lol. But if you really wanted to balance the pipeline you need to match the memory bandwidth, the flops, and most importantly the interconnect. So what does that look like ? Well if you build around LPDDR’s lower bandwidth, lower your interconnect latency, you actually can beat Nvidia on decode with a fraction of the silicon.

English
20
50
1.5K
280.2K
Cam
Cam@RealCamBrown·
@itsclivetime My knee jerk speculation is that because the heavy computational lift in AI/ML Python is all just precompiled C, you’re losing most of the ROI on a pretty considerable design undertaking. Have to make a coprocessor and intf too. Nontrivial integration task
English
1
0
3
677
Clive Chan
Clive Chan@itsclivetime·
why isn't there a startup doing a RISC-V extension for the Python Virtual Machine put a bunch of PyCores on a chip and you win the whole "Agentic CPU" market
English
21
7
210
27.4K
Cam
Cam@RealCamBrown·
@nikolasklimis My college baseball coach had us watch this 5x in a row before a practice
English
1
0
6
262
nick
nick@nikolasklimis·
Midday finds KK to Adames to D’Arnaud’s Relay during the 2019 ALDS. This team was so sneaky good.
English
8
11
309
11.9K
joe stybrenner
joe stybrenner@Joe_Dagostinooo·
@TalkinBaseball_ Thought it was closer than that? U can see how low it is as the catcher tries to frame it
English
1
0
0
1.7K
Talkin' Baseball
Talkin' Baseball@TalkinBaseball_·
This call on Nick Kurtz was so wrong they didn’t even show the measurement
English
116
408
15.1K
1.2M
Cam
Cam@RealCamBrown·
@mreynolds0504 @KutterIsKing that's what happens when it's spinning off the bat and there's high wind above the park
English
0
0
0
7
KutterIsKing
KutterIsKing@KutterIsKing·
HOLY SHIT HOW AM I JUST SEEING THIS 😭😭😭😭😭😭
English
28
11
2.5K
274.2K
Codify
Codify@CodifyBaseball·
a week's worth of successful defensive ABS challenges
Codify tweet media
English
101
230
6.3K
2.3M
Cam
Cam@RealCamBrown·
@YourAnonCentral That’s called 9:00 AM Moscow time buddy
English
0
0
0
87
Anonymous
Anonymous@YourAnonCentral·
Weird how thousands of virulently pro-MAGA accounts wake up at exactly 8:00 AM Israeli time.
English
130
5.8K
74.5K
970.3K
Cam
Cam@RealCamBrown·
@oxcrowx an* ls
Filipino
0
0
1
35
Cam
Cam@RealCamBrown·
@oxcrowx There are a lot of confounding variables in this. Was this run on a Mac? I.e. a large TAGE predictor that can take more than the execution of and ls to warm up. Notice how many more cycles the parser is. Interesting nonetheless.
English
1
0
1
317
oxcrow
oxcrow@oxcrows·
nxn's parser has less % branch misses than ls. Mildly interesting. lol.
oxcrow tweet media
English
5
0
141
5.7K
Cam
Cam@RealCamBrown·
@rohandevs @bubbleboi eh, I feel like 12 year CEO of half the EDA duopoly is a decent credential
English
1
0
1
207
rohan
rohan@rohandevs·
@bubbleboi are u not worried that gelsingher’s period of influence is over? idt the current ceo is anywhere near as competent
English
1
0
2
4.7K
bubble boi
bubble boi@bubbleboi·
The first upgrade is the cheapest. If you don’t own Intel now you might join the permanent under class forever. I fully expect the entire narrative to change now that the first domino has fallen. People aren’t even pricing in 18 AP yet they aren’t pricing in 14A. They aren’t pricing in packaging. They aren’t pricing in GPU. I’m raising my price target to $200 a share by end of the year, making Intel a 857 billion dollar company.
English
41
22
1.1K
150.5K
Cam
Cam@RealCamBrown·
@dioscuri I mean yeah if you want something built on the latest node. But that's just if you need the bleeding edge. Plenty of fabrication done on older nodes for things like consumer appliances. I don't need my smoke detector fabbed on N3P.
English
0
0
2
88
Henry Shevlin
Henry Shevlin@dioscuri·
The more I learn about the semiconductor supply chain, the more implausible it all seems. There’s a small island vulnerable to invasion where all the chips are made? And the machines to make them all come from one firm in the Netherlands? Using lenses made by one firm in Germany?
English
527
3.4K
60.9K
2.5M
Cam
Cam@RealCamBrown·
@schelldave @SteveHusseyMusi @barstoolsports what if CF wants to throw behind a lazy round of first base on a single to right-center where baserunner puts his head down and walks back to first? i'd routinely ask my first basemen to be ready for a throw behind on balls hit to right center and they'd usually just post up
English
1
0
2
119
ChillOnes
ChillOnes@ChillOnes8ohm·
As a baseball umpire for years, I have to remind almost every first baseman when a ball is hit to outfield there is literally zero reason for you to be on the bag slowing the runners route around base ,zero reason unless deliberately trying to stop a double and if the kid tries to take a double I’m not calling him out on a close play if the first baseman was occupying the bag on an outfield hit
English
3
0
5
2K
Barstool Sports
Barstool Sports@barstoolsports·
In a shock to no one, having two first bases is a BAD idea
English
1K
982
37.5K
5.9M
Soho Hank
Soho Hank@elusivecurrent·
~20 hrs from now, I’ll be sitting in a midtown office scrambling to answer an obscure IC question my Principal doesn’t know
Soho Hank tweet media
English
18
8
1.6K
151.1K
Cam
Cam@RealCamBrown·
@itsclivetime @skibidblackhole Our Altra (80 core) and Altra Max (128 cores) are based on ARM Neoverse, similar to Google Axion and NVIDIA Grace
English
1
0
0
94
Clive Chan
Clive Chan@itsclivetime·
@skibidblackhole true true. are their cores close to apple's in single core? qualcomm's certainly aren't
English
4
0
77
53.8K
Clive Chan
Clive Chan@itsclivetime·
wtf is in apple silicon? my single threaded code is 3x faster on macbook than on server while burning like half the power. if apple bothered making a 64-core part they'd take over every datacenter
English
158
283
9.7K
1.1M
Cam
Cam@RealCamBrown·
@UmptyMaTwo @SwiftOnSecurity Physically closer and integrated into the pipeline as part of hardware prefetching as well!
English
0
0
1
12
UmptyMaTwo
UmptyMaTwo@UmptyMaTwo·
@SwiftOnSecurity Ik this is a bit, but L1 cache is only fast because it's smaller than L2.(Assuming modern architecture matches my 20 year old comp org textbook on mips, lmao)
English
1
0
0
52