Maxime Bergeron

4 posts

Maxime Bergeron

Maxime Bergeron

@maxobergeron

Topology & Machine Learning

Katılım Mart 2019
328 Takip Edilen15 Takipçiler
Maxime Bergeron
Maxime Bergeron@maxobergeron·
@ch402 There is probably a way to construct a universal disentangled model. The projections from this disentangled model onto observed ones are then analogous to covering space maps from a universal cover. Isomorphism of models should correspond to isomorphisms of covering spaces.
English
0
0
0
48
Chris Olah
Chris Olah@ch402·
But as we push in this direction, we'll want to think carefully about what makes a "good" isomorphism. I suspect there's some important notion of "mechanistic faithfulness" to be pinned down.
English
1
0
27
5K
Chris Olah
Chris Olah@ch402·
Early days, but I'm pretty excited about crosscoders (transformer-circuits.pub/2024/crosscode…) as a way to start to create a more universal language of features, less tied to a specific embedding in a specific layer of a specific model... Model diffing is a striking consequence.
Anthropic@AnthropicAI

One neat thing is that, very experimentally, crosscoders can be used to "diff" models: comparing between, say, a pretrained model and a fine-tuned one, to see how they differ at a more basic level.

English
6
62
475
63K
Maxime Bergeron
Maxime Bergeron@maxobergeron·
@neilchriss There’s a sneaky fourth possibility: the observations are implicitly conditioned on an event C that causes the correlation between A and B but does not actually cause either of them.
English
0
0
1
0