Armin Bazarjani

92 posts

Armin Bazarjani banner
Armin Bazarjani

Armin Bazarjani

@arminbazza

PhD Student @USC

Los Angeles, CA Katılım Ağustos 2021
105 Takip Edilen49 Takipçiler
Armin Bazarjani retweetledi
Just Jack
Just Jack@7Veritas4·
Hard to believe Spencer Pratt lost the election despite having no experience, being highly unpopular and running in a 4:1 democratic stronghold. Makes no sense.
English
798
3.8K
40.5K
483.8K
Armin Bazarjani retweetledi
Neuron
Neuron@NeuroCellPress·
What can brains teach machines & machines teach brains? We launch the first of a three-part special issue series on neuroscience & AI, featuring reviews on embodied intelligence and the schema spectrum, plus Q&As w/ leaders in the field. hubs.li/Q04ksJl-0
Neuron tweet media
English
1
9
49
4.2K
Armin Bazarjani retweetledi
hardmaru
hardmaru@hardmaru·
For over a decade, we’ve accepted that end-to-end backprop is the only way to train deep networks. But holding the entire network in memory all at once is why AI training is hitting a resource wall. We found a new way to break the network into blocks and train them independently. The trick? Treating the network’s forward pass like a diffusion model denoising a signal. This reinterpretation slashes the memory needed to train deep models. In our #ICLR2026 paper (arxiv.org/abs/2506.14202), we matched end-to-end performance across ViTs, DiTs, and LLMs. We did this while training just one isolated block at a time.
Sakana AI@SakanaAILabs

Introducing DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation pub.sakana.ai/diffusionblocks What if we didn’t have to hold an entire neural network in memory to train it? Standard neural net training optimizes all parameters jointly. As a result, the memory required during training grows linearly with the depth of the network. In our #ICLR2026 paper, we propose DiffusionBlocks, a principled framework to train networks one block at a time, drastically reducing memory requirements while matching end-to-end performance. With DiffusionBlocks, we split the network into blocks and train them one at a time, so you only need memory for a single block. How? We explicitly assign each block a role: to move the representation a little closer to the target than the block before it did. That role turns out to be precisely what a diffusion model does, step by step. Each block only needs to optimize its own objective and can be trained independently. We validated this across five different architectures: • ViT • DiT • Masked diffusion • Autoregressive transformers • Recurrent-depth transformers In each case, performance is competitive with end-to-end training while using a fraction of the memory. This perspective also extends naturally to recurrent-depth (Looped) transformers, which apply the same network iteratively and normally require expensive backpropagation through time (BPTT). Viewed through DiffusionBlocks, we can replace those multiple iterations with a single forward pass during training. Read our paper and code, to learn more. Paper: arxiv.org/abs/2506.14202 GitHub: github.com/SakanaAI/Diffu… 🐟

English
154
628
5.7K
752.1K
Armin Bazarjani retweetledi
james hawkins
james hawkins@james406·
AI just saved me ~4 hours of driving on my vacation i was planning a trip to the Grand Canyon and noticed the driving route was incredibly inefficient i spun up a few claude code agents and used agentic ai to find this faster route it will take me just a few minutes to drive, compared to 4 hours if anyone has contacts at the US government lmk - happy to relay this to them
james hawkins tweet media
English
934
789
28.7K
5.7M
Armin Bazarjani retweetledi
Ali Behrouz
Ali Behrouz@behrouz_ali·
We keep scaling model parameters by increasing width and stacking more layers, but what if the truly missing axes for continual learning are compression and stacking the learning process? Excited to share the full version of Nested Learning, a new paradigm for continual learning and machine learning in general. Paper: nestedlearning.net/paper
Ali Behrouz tweet media
English
43
176
1.1K
147.5K
Armin Bazarjani retweetledi
Yashar Ali 🐘
Yashar Ali 🐘@yashar·
BREAKING - IRAN UPDATE HRANA’s latest numbers. Reminder: HRANA death confirmations take time as they go through multiple layers of confirmation. Total confirmed fatalities: 3,919 Deaths under investigation: 8,949 Total: 12,868 Protesters: 3,685 Children under 18: 25 Military and regime forces: 178 Non-protesting civilians: 31 Severely injured: 2,109 Arrested: 24,669 Broadcast forced confessions: 145 cases
English
21
277
907
66.9K
Armin Bazarjani retweetledi
kev joon
kev joon@never_oppressed·
Reza Pahlavi’s daughter is posting messages on IG about how people in Iran are killing themselves because they don’t think Trump is going to attack. Incredible stuff.
kev joon tweet media
English
1
264
3K
228.4K
Armin Bazarjani retweetledi
Yashar Ali 🐘
Yashar Ali 🐘@yashar·
UPDATE on IRAN HRANA’s latest numbers. This is absolutely horrific. Reminder: HRANA death confirmations take time as they go through multiple layers of confirmation. Confirmed fatalities: 3,308 people Deaths under investigation: 4,382 cases Protesters: 3,097 Children under 18: 22 Military and regime forces: 166 Non-protesting civilians: 23 Severely injured: 2,107 people Arrested individuals: 24,266 people Broadcast forced confessions: 132 cases
English
83
592
2.2K
253.6K
Armin Bazarjani retweetledi
Nature Human Behaviour
Nature Human Behaviour@NatureHumBehav·
In this study, Piray shows a problem of low statistical power in many studies that use Bayesian model selection with computational modelling in psychology and neuroscience. nature.com/articles/s4156…
English
0
19
39
6.6K
Armin Bazarjani retweetledi
Danijar Hafner
Danijar Hafner@danijarh·
Excited to introduce Dreamer 4, an agent that learns to solve complex control tasks entirely inside of its scalable world model! 🌎🤖 Dreamer 4 pushes the frontier of world model accuracy, speed, and learning complex tasks from offline datasets. co-led with @wilson1yan
English
85
353
2.6K
458K
Armin Bazarjani retweetledi
Payam Piray
Payam Piray@payampiray·
New paper with @nathanieldaw in Nature Communications: an RL model that builds a successor map compositionally: it plans as well as the best models, and links components of the map used for planning to neural codes in the medial entorhinal cortex. rdcu.be/eAofi
English
1
9
38
3.4K
Armin Bazarjani retweetledi
Marlos C. Machado
Marlos C. Machado@MarlosCMachado·
📢 I'm happy to share the preprint: _Reward-Aware Proto-Representations in Reinforcement Learning_ ‼️ My PhD student, Hon Tik Tse, led this work, and my MSc student, Siddarth Chandrasekar, assisted us. arxiv.org/abs/2505.16217 Basically, it's the SR with rewards. See below 👇
Marlos C. Machado tweet media
English
1
13
104
9.3K
Armin Bazarjani retweetledi
Alex Kendall
Alex Kendall@alexgkendall·
Here's my favourite scenario from autonomous driving in Tokyo last week! 🇯🇵
English
5
43
399
87.4K
Armin Bazarjani retweetledi
François Chollet
François Chollet@fchollet·
Much of the field obsesses over end-to-end learning. But strong generalization requires compositionality: building modular, reusable abstractions, and reassembling them on the fly when faced with novelty. The models of the future won't be just pipes, they will be Lego castles.
English
58
162
1.3K
136.8K