Chris Jericho

48 posts

Chris Jericho

Chris Jericho

@chrisjericho603

Katılım Mart 2014
243 Takip Edilen47 Takipçiler
Chris Jericho
Chris Jericho@chrisjericho603·
@giffmana What's the connection between Curry finding a dead bounce and LLM solving ghost font?
English
1
0
12
5.5K
Florian Brand
Florian Brand@xeophon·
I have a simple setup with GPT-5.6: I use Sol Pro [standard, normal speed] to plan, which I then give to Terra Ultra [fast] to explore deeper, instructing it to use Sol Medium [fast] for deeper consideration. Once they are done, they let Luna xhigh [normal] implement it.
English
20
5
193
16.4K
Chris Jericho
Chris Jericho@chrisjericho603·
@gneubig The idea that it was for all and not just the privileged few was marketing BS and ivory tower delusion. Like google "don't be evil".
English
0
0
0
333
Ben Clavié
Ben Clavié@bclavie·
One of the most interesting papers of the last ~2 years in IR only has 8 citations.
English
8
0
36
12.5K
Garry Tan
Garry Tan@garrytan·
@DWFlickinger My 150k markdown file personal knowledge wiki in GBrain
English
5
0
21
3.3K
Garry Tan
Garry Tan@garrytan·
Thinking Machines is impressive. In a couple hours I just fine tuned my own Qwen3.5-397B model this afternoon. Fast usable multimodal is also going to enable very mind-blowing personal AI.
Thinking Machines@thinkymachines

People talk, listen, watch, think, and collaborate at the same time, in real time. We've designed an AI that works with people the same way. We share our approach, early results, and a quick look at our model in action. thinkingmachines.ai/blog/interacti…

English
114
200
3K
433.6K
Niels Rogge
Niels Rogge@NielsRogge·
Since everyone and their mum is building with Claude Code, the "AI deadlines" app is now also powered by it ... except I use @MiniMax_AI for 10% of the cost 😎 All conference data is now regularly fetched automatically using a bunch of agents + @ExaAILabs's MCP server.
Niels Rogge tweet media
English
4
5
44
9.2K
Jeremy Howard
Jeremy Howard@jeremyphoward·
These 302 people have just had their social credit score slashed. If you are one of these people, and discover you now can't find work, join clubs, or book plane tickets -- remember, you brought this on yourself.
Jeremy Howard tweet media
ekri@ekrii3

The best piece of often overlooked advice that I got at rentech (by Jim himself) in the earlier days was : “Always remember, first sort, then regress”. We lived by this rule. If you have data (X,Y) sort X and Y independently, this gives a better fit.

English
9
2
224
43.2K
Petar Veličković
Petar Veličković@PetarV_93·
i just realised my "everything is connected" survey on graph nets was a "2023 paper of the year runner-up" for the cosb journal 🎉 it's not my best piece of writing about gnns (check out ch. 5 of the gdl book!) but it might just be my most concise intro to this magical world ✨
Petar Veličković tweet mediaPetar Veličković tweet media
English
3
14
148
6.6K
Sebastian Raschka
Sebastian Raschka@rasbt·
If you're getting into LLMs, PyTorch is essential. And lot of folks asked for beginner-friendly material, so I put this together: PyTorch in One Hour: From Tensors to Multi-GPU Training (sebastianraschka.com/teaching/pytor…) 📖 ~1h to read through 💡 Maybe the perfect weekend project!? I’ve spent nearly a decade using, building with, and teaching PyTorch. And in this tutorial, I try to distill what I believe are the most essential concepts. Everything you need to know to get started, and but nothing more, since your time is valuable, and you want to get to building things!
English
49
424
2.4K
148.7K
merve
merve@mervenoyann·
I learnt a lot from O'Reilly books, this is surreal I'm writing a book with amazing people @micuelll @andimarafioti @orr_zohar about VLMs with @huggingface 📖💗 Early Access (first two chapters in raw) are available to everyone, we'd love to have your feedback!
merve tweet media
English
98
237
2K
181.7K
Chris Jericho
Chris Jericho@chrisjericho603·
@rasbt Should I read the book chapter first and then watch the video or vice-versa?
English
1
0
0
629
Sebastian Raschka
Sebastian Raschka@rasbt·
My next tutorial on pretraining an LLM from scratch is now out. It starts with a step-by-step walkthrough of understanding, calculating, and optimizing the loss. After training, we update the text generation function with temperature scaling and top-k sampling. And finally, we also load openly available pretrained weights into our scratch-built model architecture. Along with this pretraining tutorial, I also have bonus material on speeding up the LLM training. These apply not just to LLMs but also to other transformer-based models like vision transformers: 1. Instead of saving the causal mask, this creates the causal mask on the fly to reduce memory usage (here it has minimal effect, but it can add up in long-context size models like Llama 3.2 with 131k-input-tokens support) 2. Use tensor cores (only works for Ampere GPUs like A100 and newer) 3. Use the fused CUDA kernels for `AdamW` by setting 4. Pre-allocate and re-use GPU memory via the pinned memory setting in the data loader 5. Switch from 32-bit float to 16-bit brain float (bfloat16) precision 6. Replace from-scratch implementations of attention mechanisms, layer normalizations, and activation functions with PyTorch counterparts that have optimized CUDA kernels 7. Use FlashAttention for more efficient memory read and write operations 8. Compile the model 9. Optimize the vocabulary size 10. After saving memory with the steps above, increase the batch size
Sebastian Raschka tweet media
English
14
164
959
44.2K
Chris Jericho
Chris Jericho@chrisjericho603·
@akarshkumar0101 Watched 35 mins so far, very interesting. What introductory book would you recommend for artificial life?
English
0
0
0
30
Akarsh Kumar
Akarsh Kumar@akarshkumar0101·
I recently gave a talk at the MIT Embodied Intelligence Seminar on Automating the Search for Artificial Life with Foundation Models, check it out! youtube.com/watch?v=S6uLxS…
YouTube video
YouTube
English
5
31
224
49.3K
Chris Jericho
Chris Jericho@chrisjericho603·
@goodside Someone please explain this meme to me? Who is the dude behind the mic and how is it linked to the tweet?
English
0
0
0
48
Riley Goodside
Riley Goodside@goodside·
GPT-3’s score on the MMLU benchmark was 40%. First release of GPT-4 scored 86%, and today GPT-4o is 89%. An increase of just 3% — that’s a full year of progress. If you plot the prior trend we were supposed to be at 100, maybe 120% by now. AI is hitting a wall.
Riley Goodside tweet media
English
53
16
563
108.1K
Not Me
Not Me@davidschlangen·
Another class I'm teaching this semester is "Programming w/ LLMs". This sidesteps the whole chatbot / assistant / "an AI" theme and looks at LLMs as function approximators -- where, weirdly, the function needs to be "found" first. (Yes, DSPy will feature heavily.)
Not Me tweet mediaNot Me tweet mediaNot Me tweet media
English
11
46
359
76.1K
Chris Jericho
Chris Jericho@chrisjericho603·
@rasbt Thanks and ML knowledge pre reqs? Neuralnets, embeddings etc?
English
1
0
1
121
Sebastian Raschka
Sebastian Raschka@rasbt·
@chrisjericho603 It'd say some experience with PyTorch would be good but there's also a ~40 page chapter (Appendix A) introducing PyTorch.
English
1
0
6
836
Sebastian Raschka
Sebastian Raschka@rasbt·
"Building a Large Language Model from Scratch" is in the final stages. I'm currently working on Chapter 6, finetuning a GPT-like LLM to classify SPAM messages. And while I'm polishing up the main content, I wanted to share some bonus experiments along with some interesting takeaways: 1. Training the Last vs. First Output Token (row 1 vs 2): Training the last output token results in significantly better performance compared to the first. This improvement is expected due to the causal self-attention mask. 2. Training the Last Transformer Block vs. Last Layer (row 1 vs 3): Training the entire last transformer block is much more effective than training only the last layer. 3. Training All Layers vs. Last Transformer Block (row 1 vs 4): Training all layers shows a modest improvement of 2% over just training the last transformer block, but it requires almost three times longer in terms of training duration. 4. Using Larger Pretrained Models (row 1 vs 5, and row 1 vs 6): Employing a 3x larger pretrained model leads to worse results. However, using a 5x larger model improves performance compared to the initial model, as was anticipated. 5. Using a Model with Random Weights vs. Pretrained Weights (row 1 vs 7): Utilizing a model with random weights yields results that are only slightly worse by 1.3% compared to using pretrained 6. Padding Input to Full Context Length vs. Longest Training Example (row 1 vs 8): Padding the input to the full supported context length results in significantly worse I've kept the LLM and dataset small on purpose, so you can run the training on a regular laptop like a MacBook Air M3 in about 15 minutes in case you don't have access to a GPU.
Sebastian Raschka tweet media
English
21
183
1.2K
126.5K
Sasha Rush
Sasha Rush@srush_nlp·
Not sure 100% sure if I'm allowed to joke about Brexit?
English
3
0
12
4.4K