Cameron Thacker

367 posts

Cameron Thacker banner
Cameron Thacker

Cameron Thacker

@CameronMThacker

PhD | Co-founder @playmythical | 0-to-1 builder | recovering physicist | AI since ML | explore gt exploit | type II fun

Los Angeles, CA Katılım Ağustos 2013
190 Takip Edilen26K Takipçiler
Cameron Thacker
Cameron Thacker@CameronMThacker·
@charliermarsh I have a fond spot for Licanius. People lost their mind over “will of the many” but I just didn’t love it. To really know if we are on the same team, do you recognize Zorian?
English
0
0
1
120
Charlie Marsh
Charlie Marsh@charliermarsh·
Developing a really bad habit of naming internal prototypes after increasingly obscure fantasy references. If you recognize names like Zvaelar or Alaswha, you're not currently on my team.
English
6
0
58
7.1K
Cameron Thacker
Cameron Thacker@CameronMThacker·
@TensorTonic Disagree. If you have taken linear algebra and calculus you have essentially derived these. Now for a lot of us this was a long time ago but still.
English
0
0
5
277
TensorTonic
TensorTonic@TensorTonic·
7 math ideas every ML engineer uses daily and almost nobody has actually derived: 1. Why gradient descent moves in the direction of steepest descent, not just downhill, but provably the steepest direction, straight from the definition of a directional derivative. 2. Why softmax plus cross-entropy collapses into that suspiciously clean gradient of pred minus true, and what breaks the moment you swap the loss function. 3. Why the chain rule is backprop, not an analogy for it, the same operation applied mechanically to a computation graph. 4. Why dividing attention scores by root d_k isn't arbitrary, it's variance control, derivable from how dot products scale with dimension. 5. Why KL divergence isn't symmetric, and what that asymmetry actually costs you when you pick forward vs reverse KL. 6. Why Adam's second moment estimate quietly approximates a diagonal Hessian, making it quasi-Newton in disguise. 7. Why eigenvectors are the directions a matrix doesn't rotate, the one geometric fact that makes SVD, PCA, and spectral clustering all click at once.
English
7
20
247
11.8K
Carl Kolon
Carl Kolon@carlkolon·
This has to be a joke right?
Carl Kolon tweet media
English
2
0
15
2.2K
Cameron Thacker
Cameron Thacker@CameronMThacker·
@willccbb That’s crazy, it must have been tough. I definitely went the common route of physics > machine learning > gaming > back to AI.
English
0
0
6
1.7K
will brown
will brown@willccbb·
i took a fairly unconventional path into ai research. i did my undergrad in computer science, and then did my phd in computer science. i spent some time in big tech, and also in quant finance. but ultimately, i realized that my true calling was working at a neolab
English
44
24
1.6K
151K
Cameron Thacker
Cameron Thacker@CameronMThacker·
@andrewgwils I'm not sure what the point of this statement is. Are you trying to imply the goal should be about finding ever better inductive biases? Otherwise, sure, everything has an inductive bias including our perception of the World.
English
0
0
0
263
Andrew Gordon Wilson
Andrew Gordon Wilson@andrewgwils·
All of machine learning is about inductive biases. It's wild to think otherwise. If all we cared about was building an expressive model, a model that can fit any data, we wouldn't need neural networks, we could do just basic linear interpolation.
English
16
18
315
34.5K
Dillon Mulroy
Dillon Mulroy@dillon_mulroy·
@mitsuhiko hugs would be great but consistency and certainty would be even better
English
7
2
96
6.7K
Dillon Mulroy
Dillon Mulroy@dillon_mulroy·
yall im ngl its way harder to get joy and satisfaction out of building with ai than it was before constantly straddling burn out, being far less immersed in hard problems, constant context switching i’m tired of the uncertainty of where this is going and how to do it well
English
283
230
3.8K
388.7K
Cameron Thacker
Cameron Thacker@CameronMThacker·
@zeeg Will do thanks! I recently tried out oh-my-pi, but my assumption is that pi tends to work better due to the simplicity of the harness.
English
1
0
0
185
David Cramer
David Cramer@zeeg·
@CameronMThacker tbqh id point coding agent at the repo and ask as i dont even remember. i THINK we use pi-coding-agent here but i could misremember we then layer a meta prompt on top, and then the skill on top of that the latter two things are the same tho as with claude sdk
English
1
0
2
1.4K
David Cramer
David Cramer@zeeg·
now that everyones on this "use harnesses vs native providers with agents", here's your important reminder that its not always a good idea... Pi vs Claude SDK: warden.sentry.dev/benchmarking
English
10
12
270
78.2K
Cameron Thacker
Cameron Thacker@CameronMThacker·
You probably won't see this, but I will say it anyway. The goal of exercise - especially for you with such high fasting insulin - shouldn't be to burn calories/lose weight. It should be to slowly attune your muscles to better utilizing glucose so your body doesn't have to produce as much insulin to regulate it. 1. During severe calorie restrictions, exercise is just walking. Do not lift weights or do intense cardio. Just walk. 2. Even when you aren't restricting, start very slow. Just focus on being more active regularly throughout the day and slowly build up for low intensity cardio time per week. No silver bullet obviously, but it will likely provide benefits even off target ones like mental clarity etc.
English
0
0
3
103
Eliezer Yudkowsky
Eliezer Yudkowsky@allTheYud·
@basar94921 5-6 weeks on 700cal/day. The first week is just burning off 10lb water weight and doesn't count as weight loss. Exercise is debilitating, does not produce strength gain, and does not enhance weight loss presumably because base metabolism drops to compensate per multiple studies.
English
5
0
8
1.7K
Eliezer Yudkowsky
Eliezer Yudkowsky@allTheYud·
Claude: What the fuck? Gemini: Seriously, what the fuck? Me: The blood test said I had fasting insulin levels of 23.4 uIU/mL, 5 weeks into my very-low-calorie ketogenic diet, with blood ketone levels of 2.4 earlier that morning. Claude: That shouldn't be physiologically possible. Gemini: It really shouldn't. Are you taking exogenous ketones? Or exogenous insulin? Me: Nope. Previously 12mg/wk retatrutide didn't do anything helpful for me either, if that helps pin anything down about how my body functions? That's why I still have to drastically keto/starve for 1 month per 2 months of eating enough to work, just to maintain my current level of overweight? A few months earlier my non-fasting insulin was 82.5. Claude: Get the insulin test repeated, and this time test C-peptides. That'll determine if you're actually manufacturing that amount of insulin or if something weirder is going on. Gemini: Yeah, I'll say exactly the same thing when you prompt me without reference to Claude's take. Me: *Sigh.* Trudge trudge trudge, test test test, over and over for years and years with nothing ever helping. Is this really worth it? Gemini: THOSE LAB NUMBERS DO NOT GO TOGETHER. Me: Okay. Sure. Whatever. Anything confusing is worth a second look.
English
77
2
341
110.7K
Cameron Thacker
Cameron Thacker@CameronMThacker·
New gpt 5.6 sol cheats so much, METR can't reliably benchmark it. Evals are so critical, both for understanding the capabilities of models, but also improving them. Yet reliable measurements seem to be inversely proportional to model intelligence? It's not concerning ... it's not concerning ...
Cameron Thacker tweet media
METR@METR_Evals

OpenAI gave METR early access to GPT-5.6 Sol for testing including raw chain-of-thought, a railfree version of the model, and internal information about the model. With this access, METR conducted a pre-deployment evaluation of GPT-5.6 Sol, including an attempted measurement of its 50%-Time Horizon. However, the measurement depends heavily on our treatment of cheating attempts, and GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated.

English
0
0
0
381
will brown
will brown@willccbb·
something tells me the next OpenAI model might be a large sparse Mixture-of-Experts Transformer 👀
English
44
9
552
44.3K
Cameron Thacker
Cameron Thacker@CameronMThacker·
@ZyphraAI Very cool. Going to have to test it in my toy model. Personally, I tend to think the residual is cool hack but we just haven’t discovered it is a projection of a better way. Maybe this is leading us closer to that!
English
0
0
0
267
Zyphra
Zyphra@ZyphraAI·
Zyphra Research is releasing Norm-AGnostic residual networks (NAG) - a new architecture that mitigates the diminishing returns of deeper residual models by controlling the residual stream geometry. NAG makes Mixture-of-Depths practical for pretraining.
Zyphra tweet media
English
4
25
184
130.8K
Cameron Thacker
Cameron Thacker@CameronMThacker·
@4confusedemoji Models are sharded across many nodes. Look up expert parallelism and tensor parallelism. For example deepseek had decode served over 144 gpus I think.
English
0
0
12
598
olivia
olivia@4confusedemoji·
how in the fuck does fable/mythos fit in a b300 rack with kv for 1m token context? its not on vera rubin
English
7
0
44
8.4K
Cameron Thacker
Cameron Thacker@CameronMThacker·
@suchenzang @geteviapp For (2) Im sure you are aware, but it just transforms the features of a given token independently. There is literally no mixing so calling it mixing is misleading for people that don’t really understand. The only mixing happens in attention.
English
0
0
0
11
Susan Zhang
Susan Zhang@suchenzang·
1) i noticeably didn't cover activations 2) how else would you describe FFNs/two-fully-connected all-to-all transformations? maybe knowledge absorbers? it's usually viewed as the memorization/knowledge storer, but that's also a distraction from the overall jailbreak point (already assume "everything is memorized")
English
2
1
21
9.8K
Susan Zhang
Susan Zhang@suchenzang·
current LLMs fundamentally consist of four main components: - input layer: where input "words" (prompt) get mapped to "latents" aka some-model-representation-you-don't-understand-unless-you-start-reading-tea-leaves-of-spurious-correlations (some quite compelling à la word2vec style; latents is also unnecessary lingo so i will refer to these as "inputs" with quotes from now on) - mixing layers: where you jumble all your "inputs" together to see if any correlations between "inputs" can become useful (commonly used to compress or expand dims; predicting a single classification target == compress to a single dim, etc) - attention layers: where you learn how "inputs" relate to each other (aka discern what's important to remember vs fluff) - residuals: where you short-circuit a mixing/attention layer because it's probably adding too much confusion (aka avoid overthinking for simple things) ----- a "big" LLM simply scales two things: - width == how many dimensions you give to your "inputs" (the more dims, in theory the more unique/discerning/precise/complex your knowledge can become) - depth == how many mixing/attention/residual layers you can stack/loop between (aka "reason" over, where more of these ~= more "reasoning" abilities) "capabilities" that seem impressive to humans usually arise from taking advantage of both depth & width: where a model seemingly makes connections between disparate ideas, beyond what an average human can hold in working memory. this requires models to "completely light up" when responding to a "hard prompt", where effectively no param/layer goes unused. ----- the anatomy of a "model capability" is precisely the same mechanism that can be co-opted for a jailbreaking exploit: your goal is simply to "light up" as much of the model as possible, dodging any shallow input-classifiers at the beginning by triggering as many disparate "input ideologies" as possible, and subsequently have these "inputs" relate to each other in seemingly unrelated-yet-related ways that ideally have similar "complexity" as your jailbreak goal (to make it past enough layers of the model). think of the attack-vector as bundling your goal in a series of schizo-nerd-snipes: a sufficiently capable model will try to reason through everything all at once, eliminate the dead-ends, and successfully deliver the one jailbreak use-case you bubble-wrapped for. of course, there's an art to the above, and some are already extraordinarily proficient at the trojan-horse-packaging, but at some point there's no difference between "a capability" and "a jailbreak", though i'll be happy to be proven otherwise. ----- tl;dr ant flew too close to the sun, better kiss the ring or get buried.
English
22
91
1.1K
168.7K
Cameron Thacker
Cameron Thacker@CameronMThacker·
@francoisfleuret If it works, it isn’t dumb. Even if it is some approximation of an underlying truth
English
0
0
1
157
François Fleuret
François Fleuret@francoisfleuret·
Me after 1h of intense thinking on how to deal intelligently with a sophisticated identified gradient amplification that crashes my training: clip: 0.1
English
3
0
37
4.5K
Cameron Thacker
Cameron Thacker@CameronMThacker·
High signal, not low. Assuming it's not a permanent 100% of their role, getting your hands in the data is very underrated. The idea that it's beneath an engineer is pure ego. I spent weeks hand-labeling data on some previous projects. High quality data is priceless.
Gergely Orosz@GergelyOrosz

Just learned: Software engineers used to do manual data labeling at Scale AI while Alex Wang was CEO. After he left, new leadership joined, and were HORRIFIED to learn this. Stopped it ASAP Now at Meta, software engineers are assigned manual data labeling... see the pattern?

English
0
0
0
208
Cameron Thacker
Cameron Thacker@CameronMThacker·
If you learn from experience and outcomes, you must still possess some introspection or else you will develop a giant blind spot. You will be able to see what happened, but not understand your own role in the outcome. People like this tend to blame others for problems of their own making.
English
0
0
0
15
Alps
Alps@alpaysh·
i still think about this
Alps tweet media
English
353
1.1K
7.6K
1.3M
Cameron Thacker
Cameron Thacker@CameronMThacker·
@varunneal Well you don't get any gradient signal on experts that aren't chosen so ... you won't optimize the right thing unless you force exploration. But I agree this isn't satisfying. I think the solution to sparsity will end up a lot more dynamic personally.
English
0
0
2
323
varun
varun@varunneal·
its so interesting to me that MoE router weights are the one place we don't really trust gradient descent. We add our own biases, sorting functions, heuristics because regular descent wants to be totally degenerate
English
12
0
84
17.6K