Martino Andrea Scarpolini

138 posts

Martino Andrea Scarpolini banner
Martino Andrea Scarpolini

Martino Andrea Scarpolini

@scarpma

Theoretical physicist. PhD student in fluid dynamics, bio-medical engineering and deep learning. I love tech, computer science, climbing and skiing. 🍂

Katılım Temmuz 2013
191 Takip Edilen25 Takipçiler
Matteo Salvini
Matteo Salvini@matteosalvinimi·
A Roma oggi scontri, violenze, vetrine spaccate, teppisti incappucciati contro Polizia e Carabinieri, lancio di bottiglie e bombe carta, auto in fiamme. Questi non sono manifestanti o pacifisti, sono criminali! GRAZIE alle donne e agli uomini in divisa, siamo sempre al vostro fianco.
Italiano
555
292
1.9K
62.9K
Massimo
Massimo@Rainmaker1973·
According to Hive Systems your passwords should be at least 18 mixed up characters for maximum security. Here's how long it would take a hacker to come up with as many combinations or passphrases before guessing your password [read more, methodology: buff.ly/3Mw5nDM]
Massimo tweet media
English
95
301
1.2K
409.5K
Martino Andrea Scarpolini
The training cost (estimated around 60M $) is quite shocking
Aleksa Gordić (水平问题)@gordic_aleksa

OpenAI's GPT-4 details have apparently been leaked! Looks very detailed and I suspect it's the real deal - given all I know about how these systems work. Here is a summary (extractive+abstractive) I made based on the original thread (see bottom of the post) + some additional pointers from me. (Disclaimer there are probably errors here) ---ARCH--- * GPT-4 is more than 10x the size of GPT-3 (175 B). We believe it has a total of ~1.8 trillion parameters across 120 layers. Mixture Of Experts (16 experts, each ~111B). Not a dense transformer like e.g. PaLM (or GPT-3). They use MQA instead of MHA (classic at this point). * Each forward pass (generation of 1 token) only utilizes ~280B parameters and ~560 TFLOPs. This contrasts with the ~1.8 trillion parameters and ~3,700 TFLOP that would be required per forward pass of a purely dense model. ---DISTRIBUTED--- * To parallelize across all their A100s GPUs They utilized 8-way tensor parallelism. Beyond that, they are using 15-way pipeline parallelism. Also apparently they used DeepSpeed ZeRo Stage 1 or block-level FSDP. (You can check out my video on all of these strategies here: youtube.com/watch?v=hc0u4a… 3D parallelism is what you're looking for & ZeRo) ---VISION--- They have a separate vision encoder from the text encoder, with cross-attention. The architecture is similar to Google DeepMind's Flamingo (I used to work on this project :) ). This adds more parameters on top of the 1.8T of GPT-4. It is fine-tuned with another ~2 trillion tokens, after the text-only pre-training. ---DATA--- * Trained on ~13T tokens (multiple epochs, not unique). Plus millions of rows of instruction fine-tuning data from ScaleAI & internally (I guess acquired through ChatGPT + their API before they changed the policy). * 8k context length for the pre-training phase. The 32k seqlen version of GPT-4 is based on fine-tuning of the 8k after the pre-training. See e.g. MosaicML's blog on how to achieve this: github.com/mosaicml/llm-f…) ---COST--- * OpenAI’s training FLOPS for GPT-4 is ~2.15e25, on ~25,000 A100s for 90 to 100 days at about 32% to 36% MFU. Part of this extremely low utilization is due to an absurd number of failures requiring checkpoints that needed to be restarted from. If their cost in the cloud was about $1 per A100 hour, the training costs for this run alone would be about $63 million. (Today, the pre-training could be done with ~8,192 H100 in ~55 days for $21.5 million at $2 per H100 hour) ---INFERENCE--- OpenAI might be using speculative decoding on GPT-4's inference. See this paper: arxiv.org/abs/2211.17192 The inference runs on a cluster of 128 GPUs. There are multiple of these clusters in multiple datacenters in different locations (it'll be hard for Elizier to nuke these xD). 8-way tensor parallelism and 16-way pipeline parallelism. Original thread: archive.is/2RQ8X (strictly speaking the original one has been removed).

English
0
0
0
18
Martino Andrea Scarpolini retweetledi
Massimo
Massimo@Rainmaker1973·
According to the Guinness Book of Records, the anechoic chamber at Orfield Laboratories in Minneapolis is the quietest place in the world, with a background noise reading of –9.4 decibels. Very few people have been able to bear it for more than 45 minutes without experiencing discomfort, claustrophobia, nausea and panic. If you chatted with someone, your speech would measure around 60 decibels on a sound-level meter. If you stood quietly on your own in a concert hall, the meter would drop down to a level of about 15 decibels. The threshold of hearing, the quietest sound a young adult can hear, is about 0 decibels. The test room at Orfield Laboratories, like the chamber at Salford University, is far quieter than that. You can hear each grind, crunch and squeak your skeleton makes. An anechoic chamber has an impressive silence because it simultaneously presents two unusual sensations: not only is there no external sound, but the room puts your senses out of kilter. An anechoic chamber is often isolated from energy entering from their surroundings. This combination means that a person or detector exclusively hears direct sounds (no reflected sounds), in effect simulating being outside in a free field. Through their eyes, visitors obviously see a room, but their ears hear nothing that indicates a room. Add the claustrophobic drama of being enclosed behind three heavy doors, and some begin to feel uneasy and ask to leave. And what this kind of room is useful for? Anechoic chambers are designed to stop reflections or echoes of sound. They are commonly used in acoustics to conduct experiments in nominally "free field" conditions, free field meaning that there are no reflected signals. All sound energy will be traveling away from the source with almost none reflected back.
Massimo tweet media
English
453
1.1K
10K
3M
Martino Andrea Scarpolini retweetledi
elvis
elvis@omarsar0·
One of the more interesting computer vision papers I read this week: They propose applying the Segment Anything Model (SAM) to medical 2D images. The challenge here is taking a model pretrained on natural images to work on medical images. There is an obvious domain gap. So the authors proposed a large-scale dataset medical image segmentation dataset containing 4.6M images and 19.7M masks, including various modalities and objects. This is huge for the research community. The SAM model was then fine-tuned on this dataset and evaluated on medical image segmentation across modalities and anatomical structures. Finetuning is obviously a big deal here and that's shown through the really strong performance obtained on several datasets. Fun paper to read. paper: arxiv.org/abs/2308.16184… code: github.com/uni-medical/SA…
elvis tweet media
English
6
109
511
120K
Martino Andrea Scarpolini retweetledi
Jim Fan
Jim Fan@DrJimFan·
The famed Stanford Smallville is officially open-source! 25 AI agents inhabit a digital Westworld, unaware that they are living in a simulation. They go to work, gossip, organize socials, make new friends, and even fall in love. Each has unique personality and backstory. Smallville is among the most inspiring AI agent experiments in 2023. We often talk about a single LLM's emergent abilities, but multi-agent emergence could be way more complex and fascinating at scale. A population of AI can play out the evolution of an entire civilization. Endless new possibilities ahead. Gaming will be the first to feel the impact. Github: github.com/joonspk-resear… Paper: arxiv.org/abs/2304.03442 Authors: @joon_s_pk @joseph_c_obrien @carriejcai @merrierm @percyliang @msbernst
Jim Fan tweet media
English
273
2.2K
9.4K
4M
Martino Andrea Scarpolini retweetledi
Emmet Halm
Emmet Halm@ehalm_·
@sama switching my VPN to Italy from now on
Emmet Halm tweet media
English
8
7
182
19.6K
Martino Andrea Scarpolini
@swbeeton @jamesjansson @sama I think that these people always complaining, actually do not understand the usefulness of chatgpt. They do not understand the conceptual separation between high level and low level work
English
1
0
0
34
Sam Altman
Sam Altman@sama·
heard something like this 3 times this week: "our recent grads are now much more productive than people who have worked here for years because they've really learned how to use ChatGPT".
English
420
793
9K
1.8M
RE Giorgy Napolitano
RE Giorgy Napolitano@ReGiorgio_Primo·
@ultimora_pol Complimenti al fotografo che si trovava lì per caso ed ha immortalato questo frangente di spontaneità.
Italiano
2
2
215
4.7K
Ultimora.net - POLITICS
Ultimora.net - POLITICS@ultimora_pol·
Etiopia, Giorgia #Meloni si rivolge ad alcuni ragazzi che si sono avvicinati per presentarsi durante la visita alla scuola Galileo Galilei di Addis Abeba: "Ciao ragazzi, studiate sempre" @ultimora_pol
Ultimora.net - POLITICS tweet media
Italiano
365
47
1.1K
868.7K
Martino Andrea Scarpolini
First He publicly demands for a 6 months stop for new AI big models. Now he starts a new corp in AI... 🙊
English
0
0
0
8
Martino Andrea Scarpolini retweetledi
GREG ISENBERG
GREG ISENBERG@gregisenberg·
Btw, ChatGPT is old news AutoGPT is the new kid on the block Read this if you want to stay ahead:
GREG ISENBERG tweet media
English
450
2.7K
18.1K
6.5M