Ionut-Vlad Modoranu

16 posts

Ionut-Vlad Modoranu

Ionut-Vlad Modoranu

@ionutmodo

PhD student @ ISTA 🇦🇹 Former Research Intern @ Together AI 🇳🇱 Studied Computer Engineering & Machine Learning 🇷🇴

Vienna Katılım Nisan 2018
271 Takip Edilen52 Takipçiler
Ionut-Vlad Modoranu
Ionut-Vlad Modoranu@ionutmodo·
If you’re at the conference and want to geek out over LLM optimizers, send me a DM. I'll be around and happy to meet. I’m also nearing the end of my PhD and actively exploring Research Scientist opportunities. If you're building a great team, let's chat! [3/3]
English
0
0
0
21
Ionut-Vlad Modoranu
Ionut-Vlad Modoranu@ionutmodo·
Training LLMs just got faster I'm at @icmlconf in Seoul this week presenting DASH - our latest improvement to the well-known Shampoo optimizer. By using DASH, we're achieving 5x lower running time and higher accuracy than the standard Distributed Shampoo implementation! [1/3]
English
2
1
5
100
Ionut-Vlad Modoranu
Ionut-Vlad Modoranu@ionutmodo·
Poster details: 🗓️ Thursday, July 9th ⏱️ 2:30PM - 4:15PM 📍 Hall A, Session 7, Poster 405 [2/3]
English
0
0
0
25
Ionut-Vlad Modoranu retweetledi
Dan Alistarh
Dan Alistarh@DAlistarh·
We're releasing DASH (Distributed Accelerated Shampoo), an improved implementation of the Shampoo optimizer that achieves up to 4.83× faster optimizer steps, while matching or improving final model quality. [1/6]
Dan Alistarh tweet media
English
1
14
68
3.1K
Oncel Tuzel
Oncel Tuzel@OncelTuzel·
Come work with us! The Machine Learning Research (MLR) team at Apple is seeking a passionate AI researcher to work on Efficient ML algorithms, including models optimized for fast inference and efficient training methods. Apply here: jobs.apple.com/en-us/details/…
English
6
41
366
32.6K
Sam Bhagwat
Sam Bhagwat@calcsam·
icymi we wrote a new agents book: patterns for building ai agents it has everything you need to take your agents from prototype to production, like agent design patterns, the basics of security, etc reply to this tweet with BOOK and we'll dm you so you can get a copy
Sam Bhagwat tweet media
English
5.3K
738
7.9K
1.1M
Ionut-Vlad Modoranu
Ionut-Vlad Modoranu@ionutmodo·
@asmah2107 Hi! I'm currently a PhD student working in optimizers for deep learning. I'm focused on decreasing the memory usage and speeding up the runtime.
English
0
0
0
6
Ashutosh Maheshwari
Ashutosh Maheshwari@asmah2107·
Twitter is cool. But it’s 10x better when you connect with people who like building and scaling GenAI systems. If you’re into LLMs, GenAI, Distributed Systems or backend. say hi.
English
1K
64
2.8K
171.1K
Ionut-Vlad Modoranu
Ionut-Vlad Modoranu@ionutmodo·
@iamgrigorev This is a great tip for productivity! Please check out our GridSearcher project, it is completely written in Python and allows you to run jobs in parallel or sequentially using dictionaries for hyper-parameters by employing a basic scheduling: github.com/IST-DASLab/Gri…
English
0
0
1
64
George Grigorev
George Grigorev@iamgrigorev·
one interesting trick that i constantly use when running experiments: 1. I create bash script with a command to launch an experiment 2. while it's running, I do some data work, implement another feature in training framework etc 3. I then add new line to the existing bash script that is currently running with new experiment launch 4. when current experiment finishes the new one launches automatically 5. you could continue on and on and play catch-up with your runs. You try to setup a new experiment before the current one finishes. (wonder if there's a better way - like an ongoing queue) P.S. There's an alternative -- sometimes I open a new pane in tmux and write sleep X && bash Y but you need to calculate time
English
2
0
18
1.2K
Ionut-Vlad Modoranu retweetledi
Dan Alistarh
Dan Alistarh@DAlistarh·
We're releasing the DASLab GGUF Quantization Toolkit! 🚀 First open-source toolkit bringing GPTQ + EvoPress to @ggerganov's GGUF format, enabling heterogeneous quantization based on importance. Result: Better models at the same file size. [1/5]
Dan Alistarh tweet media
English
4
49
267
66.4K
Ionut-Vlad Modoranu retweetledi
Dan Alistarh
Dan Alistarh@DAlistarh·
🚀 We are releasing state-of-the-art post-training quantization (PTQ) algorithms for Microscaling FP4, together with kernels: - First study focused on MXFP4/NVFP4 PTQ for LLMs - New Micro-Rotated (MR) format and GPTQ algorithm - QuTLASS GPU kernels with up to 3.6x speedups.
Dan Alistarh tweet media
English
2
27
150
9.5K
Ionut-Vlad Modoranu
Ionut-Vlad Modoranu@ionutmodo·
@KwangjunA This seems to be similar to our recently shared work, where we use Discrete Cosine Transform (DCT) to perform a cheap low-rank projection of the momentum buffer, followed by Newton-Schulz orthogonalization. Check out our paper here: arxiv.org/abs/2505.17967
English
1
0
0
63
Kwangjun Ahn
Kwangjun Ahn@KwangjunA·
New improvement in Dion leads to a speedup that makes orthonormal updates (eg. Muon) more scalable for larger matrices. The trick: carefully using Newton-Schulz (on smaller matrices) as Dion's backend. Updates to our microsoft/dion codebase are coming soon---stay tuned!
Kwangjun Ahn tweet media
English
1
4
27
2.5K
Warren Buffett 📈 $SPY 👑
Warren Buffett 📈 $SPY 👑@OptionsBuffett·
I’m officially restarting the $100 to $10k Account Challenge for 2025 Tomorrow 🚨 This challenge is going to retire dozens and will change lives 💰 Like & comment “Trade” to join! ❤️ (Notifications must be on 🔔)
Warren Buffett 📈 $SPY 👑 tweet media
English
548
51
985
100.2K
Deedy
Deedy@deedydas·
Tried this 5 times on ChatGPT 4.5 and it gets it wrong every time.
Deedy tweet media
English
693
117
3.8K
654.4K
Ionut-Vlad Modoranu retweetledi
Dan Alistarh
Dan Alistarh@DAlistarh·
Introducing Panza, a personalized LLM email assistant, running entirely on-device! [1/6] * Panza adapts LLaMA-3-8B to match your unique writing style; * Can be fine-tuned and executed on a single GPU (free Colab version available!). Give it a try: github.com/IST-DASLab/Pan…
Dan Alistarh tweet media
English
4
30
139
39K