Mehrdad Farahmand

161 posts

Mehrdad Farahmand banner
Mehrdad Farahmand

Mehrdad Farahmand

@mifarion

Building infrastructure for LLM data quality & synthetic data | Founder @ Calibrion | PhD (NLP), University of Geneva

Berlin, Germany Katılım Ağustos 2017
258 Takip Edilen58 Takipçiler
Sabitlenmiş Tweet
Mehrdad Farahmand
Mehrdad Farahmand@mifarion·
Must-read for anyone post-training. Data quality is important. As one case study here shows, DPO on a popular alignment dataset increases the score on the benchmarks but it actually breaks safety guardrails.
Goodfire@GoodfireAI

Have you debugged your training data? You might not like what you find. Introducing predictive data debugging: reveal and shape what your model will learn before training. In DPO datasets, we found broken guardrails, hallucinations, and fish fart fan fiction (seriously). (1/9)

English
0
0
0
338
Yann LeCun
Yann LeCun@ylecun·
@hosseeb @deanwball Soooo, releasing Linux was dumping? Apache, MySQL, PHP? HTTP, TCP/IP, OpenSSL, OpenSSH? Libjpeg, VLC? The open source software stack of the mobile communication network? Signal? PyTorch? Llama?
English
115
251
3.7K
327.9K
Haseeb >|<
Haseeb >|<@hosseeb·
This argument by @deanwball is being badly misunderstood. It's OK to disagree with it, but first you have to actually understand what he's saying. He's saying: releasing the weights for a frontier-level model is effectively dumping. Dumping is when you sell a product at significantly below cost in order to corner market share. It's illegal. The reason: dumping results in short-term consumer surplus, but long-term it prevents the formation of a competitive market and discourages capex outside of the dumper. Standard Oil famously did this in order to consolidate the oil market before it was broken up. So why is he claiming releasing the weights of a frontier level model is basically dumping? Isn't he just describing open source? His argument: it's not financially sustainable to train a frontier model and release the weights. In the long run, you will not be able to internalize enough of the gains given the cost of training a frontier model, because neoclouds and other inference providers will be able to outcompete you at actually serving the model. It costs an astronomical amount of money to train frontier models, and if everyone else can serve them, you don't capture enough of the surplus to pay for the training and R&D. It's not like normal open source when you build some software and then release it and sell services on top of it. The amount of capex required for frontier-level models is an order of magnitude higher than normal software, which is why doing this at frontier level is so economically irrational. Right now the Hong Kong stock market is ebullient enough that Chinese AI companies are not getting punished for the fact that they're all deeply, deeply unprofitable. Releasing model weights is great marketing, intellectually appealing, and strikes fear into the hearts of their opponents. We can assume the status quo continues for a while because of the AI supercycle. But eventually the AI market will correct, the Hong Kong market will dump, and suddenly these Chinese labs won't be able to afford to training super expensive models without internalizing more of the gains. But what if China, seeing that this strategy is successfully kneecapping the US lead (by discouraging further capex and lowering valuations), says no--don't stop. And so the Chinese government starts buying up the shares of these companies and demanding that they continue releasing frontier-level weights, profitable or not. In that case, it becomes a genuine space race. For-profit companies cannot continue to compete on either side. US labs valuations fall, and the White House realizes that to keep their advantage in the AI race, they cannot rely on the free market to maintain their lead. They nationalize the labs and fund them off government subsidies. Now you have government-controlled and distributed models on both sides. That's what Dean is calling the "dystopian hellscape." The best analogy is drug development: if China were to sell American drugs back to us really cheaply, that would result in a large short-term consumer surplus. Cheap Viagra and Ozempic is obviously great. But in the long run, this would discourage investment in developing new drugs. That's the sense that Dean is saying it's long-term "decel." Now, I happen to disagree with Dean. I think the consumer surplus of having frontier-level open weight models is huge, even at the current capabilities. I also think China is going to defect from this strategy soon (there's been reporting along these lines, that Beijing will stop allowing large models to be open-weight; I think there are other reasons for this aside from competition). I also suspect that nationalization of labs is inevitable as they take on more geopolitical and cyber capabilities. But he's not wrong--releasing frontier-level weight models is weird. The question of how long this market will remain profit-driven is a very coherent question to ask.
martin_casado@martin_casado

"Open-weight models are inherently decelerationist" .... this is a grossly incorrect statement with no supporting arguments or logic that is counter to the long arc of learnings of the industry over the last 50 years. What a stupid thing to say.

English
238
62
636
340.4K
Mehrdad Farahmand
Mehrdad Farahmand@mifarion·
@eisokant Our synthetic data platform isn't public yet, but it's running in a few pilots. If it's useful for you too, would love to chat and showcase it. I can share docs and features (released and planned).
English
0
0
0
171
Eiso Kant
Eiso Kant@eisokant·
@mifarion Want to drop some visual samples here? Will be cool for people to see what it looks like
English
3
0
0
517
Eiso Kant
Eiso Kant@eisokant·
Starting to think/work on visual understanding for future Poolside models. Who would be the best data company to partner with for user interface data? Any YC companies out there?
English
18
4
95
12.7K
Mehrdad Farahmand
Mehrdad Farahmand@mifarion·
We don't have UI-visual understanding annotated or synthetic data yet, we build datasets/workflows as requests come in, happy to scope that with you. In the meantime, here's a look at our HITL annotation + QC platform (quality scoring, filtering, etc.) to give a sense of how the platform looks like:
English
0
0
0
75
Bela Wiertz
Bela Wiertz@blwiertz·
Berlin can build cool stuff, trust me :)
Max Meuer@meuer_max

We're hosting a Continual Learning Hackathon in Berlin on August 1st. Together with @raphaelmosaic are bringing together ~20 of Europe's strongest ML researchers and engineers to work on models that keep learning during deployment. Topics: synthetic RL environments, on-policy distillation, KV-cache compaction … We will provide dataset and compute together with our friends at Lyceum. Looking forward to share more on what we are up to with everyone there. You can find the application link in the comments to this post. We will organise travel stipends for strong student candidates outside of Germany that wouldn’t be able to make it otherwise. If you are working on similar topics reach out directly via PM for those. cc: @blwiertz

English
2
3
16
1.8K
Mehrdad Farahmand
Mehrdad Farahmand@mifarion·
@ajratner Sure. If you had to make a similar prediction today, where would you place your bet: customization through training/fine-tuning (or changing model weights in general), harness/context engineering, or something else entirely?
English
0
0
0
21
Mehrdad Farahmand
Mehrdad Farahmand@mifarion·
Given the current direction of AI, this post from over two years ago has aged remarkably well. Despite the constant hype cycles and changing narratives over the past two years, it turned out to be one of the best predictions of where LLMs and AI agents were headed. Well predicted, Alex Ratner @ajratner
Alex Ratner@ajratner

1/ Prediction: Everyone will soon be using foundation models (FMs) like GPT-4. However, they'll be using FMs trained on their own data & workloads: "GPT-You", not GPT-X Tl/dr: - Closed APIs aren't defensible - The durable moat is data - The last mile generates the real value

English
1
0
3
823
Fireworks
Fireworks@FireworksAI_HQ·
Every company must own its intelligence. We've raised $1.5B Series D at a $17.5B valuation. We’ve surpassed $1B ARR and serve over 40 trillion tokens daily, with 95%+ coming from models specialized on customer data. We’re just getting started. More: fireworks.ai/blog/series-d-…
English
54
53
912
313K
Mehrdad Farahmand
Mehrdad Farahmand@mifarion·
@lunwang1996 Where do you place data quality in this equation? Particularly, how important do you think high-quality training and evaluation data are relative to team, process, and compute?
English
1
0
0
61
Lun Wang
Lun Wang@lunwang1996·
LLM training = management problem cosplaying as technical problem. The recipe: small team, high intellectual intensity, frictionless process, enough compute. Nothing else matters as much.
English
1
2
21
2.5K
Mehrdad Farahmand
Mehrdad Farahmand@mifarion·
@alexandr_wang Congrats on the launch! This looks like a remarkable launch. Could you share more about the privacy settings? Are uploaded images and prompts used for training, how long are they retained, and can users opt out?
English
0
0
0
2.2K
Alexandr Wang
Alexandr Wang@alexandr_wang·
1/ releasing muse image today — the first image generation model from MSL. it's agentic: pairs with muse spark to reason through your prompt, search the web, and plan before it generates. people get what they meant on the first try. live now in the Meta AI app.
Alexandr Wang tweet mediaAlexandr Wang tweet mediaAlexandr Wang tweet mediaAlexandr Wang tweet media
English
141
247
2.1K
605.1K
Mehrdad Farahmand
Mehrdad Farahmand@mifarion·
This is also how we think about it at Calibrion: synthetic data only becomes useful when the quality layer is solid. A blend of distributions, drift, entropy, predictive analysis, llm-as-a-judge, and human review where needed. Our short follow-up here: calibrion.ai/blog/synthetic…
English
0
0
0
20
Mehrdad Farahmand
Mehrdad Farahmand@mifarion·
A revisit of our blog from early 2025 about synthetic data. Back then, synthetic training data looked risky. Model collapse was the main concern and research supported this. Though I still think naive synthetic data is risky, the picture changed quite a bit.
English
4
0
0
35
Mehrdad Farahmand
Mehrdad Farahmand@mifarion·
But the crucial step now is data quality. Creating and annotating data at scale is becoming easier and cheaper with agents and synthesis. Knowing which examples are good enough to train on is the harder problem.
English
0
0
0
17
Mehrdad Farahmand
Mehrdad Farahmand@mifarion·
NVIDIA has done a lot here recently (Llama-Nemotron, OpenMathReasoning, , Cosmos). Microsoft too (Phi-4-reasoning). It is not only research anymore. The market around synthetic data is also moving fast, which means it's being treated as infrastructure.
English
0
0
0
14
Mehrdad Farahmand
Mehrdad Farahmand@mifarion·
Both research and industry have shown that synthetic data can work when it is targeted, filtered, and verified. Synthetic data that then looked like a path to model collapse is now part of the training data preparation for a growing number of strong models.
English
0
0
0
8
Mehrdad Farahmand
Mehrdad Farahmand@mifarion·
@ClementDelangue Well data scale is easy at this point now that data synthesis via llms is becoming more popular and widespread, but ensuring data quality at this scale of data is the critical part.
English
0
0
0
39
clem 🤗
clem 🤗@ClementDelangue·
This is how how much data AI builders are storing on HF Xet (replaced git storage fully in ~Nov 25). Feels like this is just the beginning and should get to exabytes soon!
clem 🤗 tweet media
will depue@willdepue

A Stargate for Data Labs are on a trajectory towards >$100B/year of data spend by 2030. As we begin the trillion-dollar compute project, we need to think about the equivalent civilizational-scale effort for the other core ingredient: data. At the foundation of the scaling revolution is a simple empirical law: deep neural networks improve smoothly, near magically, as you scale two things in proportion — (1) the size of the model and (2) the amount of data you train on. And despite the scaling laws being brutally diminishing, we’ve successfully bitten the bullet of logarithmic scaling with exponentially larger clusters and datasets, and received incredible new capabilities in return. But this exponential scaling is bound to hit some limits. Oddly enough, compute has compounded fairly smoothly without limit, with trillions flowing into hypercluster buildout. Instead, we’re starting to hit the limits of an exponential demand for data. Gone are the days of being purely in the compute-limited regime, where we had effectively infinite internet data but never enough GPUs, we’re now entering a data-limited regime. Luckily, this limitation is coinciding with staggering improvements in AI capabilities. Incredibly, we seem to have a real line of sight towards automating a majority of knowledge work with the methods we have today. RL + pretraining, and the data for each, will be generally sufficient to achieve most economically valuable tasks, given some minimal algorithmic progress and continued compute scaling. In a data-limited world, economic progress & scientific acceleration will be directly bottlenecked by our coverage in each domain. We need to see data collection as imperative, deserving the same civilizational ambition we’ve given compute. The internet as a one-time subsidy It’s underrated how much all progress in AI owes everything to the blessing of the internet, this one-time civilizational subsidy to deep learning, decades of unintentional accumulation of a perfect dataset: every book, blog post, image, video, paper, discussion, etc. all digitized and freely available. Without the internet, we’d likely see comparably minimal progress in AI today, and in fact, if you notice where systems currently underperform, it’s almost always a domain where web coverage is limited and data is private, expensive, non-digitized, or non-existent. But we’re running out of it. There are only about 300 trillion tokens of useful public human text, and the internet doesn’t produce nearly enough new high-quality data to match what scaling demands — we’re soon to hit the limits of public data for pretraining. And though the advent of RL bought us reprieve — chain-of-thought RL needed a new form of untapped data, gradable math & coding tasks, also available online — we’re quickly running dry of hard tasks for RL as well. Why do we need so much data anyways? Humans learn comparably in far less time, needing just one textbook where language models might need the equivalent of hundreds to learn a new topic. It’s possible we discover methods that are massively more data efficient — synthetic data, data efficient architectures, other exotic algorithms — but fundamental progress is slow and highly unpredictable, and the recipe we have just works today. And, while I’m wary of getting too deep here, even arbitrary data efficiency can’t replace data that just doesn’t exist in the first place. There’s a massive amount of missing information on the web: the dark matter of the internet — tacit knowledge, undocumented processes, etc. — most of which was never published and lives only inside organizations, the physical world, or just in people’s heads. I’ll leave it here and say, for reasons far longer than I can fit in this post [1], it’s best to operate on the assumption that our insatiable desire for data will continue as it has for the last decade. There will be >$100B/year in data spend by 2030 We’re not screwed yet, of course. Only a fraction of useful data in the world is on the public internet, the rest is stored inside private datasets, corporations, personal archives, universities, governments, and otherwise. Labs can and will continue to license these private datasets, or create them from scratch, like Anthropic’s book scanning project. And we’ll increasingly task human experts to manufacture new high-quality data, with a large fraction of hard RL training tasks already being sourced this way. But collecting this data, unlike before, will be expensive. As the free internet dries up and demand for data rises, we should see labs investing equally in data as compute, likely spending a significant fraction of their compute budgets on data. As we see trillions spent on compute, we should also expect hundreds of billions spent on data (human data & collection budgets), given their equivalent importance. And, notably, data spend is already tracking this way: total data spend across vendors, not counting internal lab efforts, is already roughly $7 billion per year. It’s quite reasonable we’ll see >10x by 2030. Data is the moat Data becoming increasingly private will also majorly shift the competitive landscape. While compute is a commodity — everyone buys the same chips and builds the same clusters — data really isn’t. The big reason why frontier models have felt eerily similar to one another, until now, is they were trained on substantially the same internet (pretraining data variability across labs seems pretty low). As labs diverge onto more exclusive, manually collected corpora, I think models will begin to increasingly diverge. OpenAI pulling ahead in mathematics and Anthropic in cybersecurity isn’t an accident. I really think laser-focused collection of high-quality midtraining tokens, custom RL tasks, environments, with dedicated research effort, has driven much of the visible progress in the last year. James Betker has an excellent blog about “the ‘it’ in a model is the dataset”: model architecture and compute buy you efficiency and order-of-magnitude performance, but ultimately, models, of any architecture, are such incredible approximators of their dataset that the core meat of a model boils down to just that, nothing else. Data is a major moat. AGI long, ASI short As I’ve tweeted before, I’m confident that, despite the narrative, the data labeling industry will continue to fuel great businesses and be an excellent AGI long, ASI short. The argument is just: By the time the AGI labs no longer need data, it’s probably over for everything else too [2]. In this frame, the last companies left should be the data companies, as the last speck of economically relevant data is sucked in. And these companies are already among some of the fastest-growing companies in history: Mercor, founded three years ago, is rumored to be doing $2 billion in revenue with something like a few million expert labelers under contract. While these businesses are very non-stationary, what type of data is needed shifts constantly, I don’t think that diminishes their value. The long-tail of the economy is long, and the value isn’t diminishing as you extend farther into more obscure information: as models get more capable, the value of the marginal dataset goes up, not down. Automating a full job means covering its full distribution of tasks, tools, edge-cases, and long-horizon loops. There’s some O-ring logic to it: a dataset that buys a 1% bump can justify a previously unjustifiable collection cost when it’s the difference between a system that does 99% of a job and one that does all of it [3]. The competitive dynamics of the data industry are still evolving but as demand for data is increasingly niche, ultra high-quality, expert-generated, I think we’ll see real consolidation. Again, contra-narrative, we’ll probably see true competitive differentiation built on brand, quality control of data (which, from personal experience, can vary massively), as well as in network effects from the talent networks themselves over time. We’ve already seen rapidly shifting data type demand work in favor of incumbents, benefiting those with early knowledge of where the market is headed. The binding constraint It’s truly remarkable that we seem to have the recipe — pretraining + RL — to absorb most economically valuable work, despite being far from a lot of what we expected from “AGI”. The same way chess engines revealed we never needed general intelligence to solve chess, as we originally thought, we’ll soon realize that software, mathematics, and the vast majority of the economy (including physical, just running ~3 years behind!) are the same. If recursive self-improvement or some other algorithmic breakthrough arrives, that’s wonderful, but we really don’t have to wait for it. The binding constraint between here and an automated economy isn’t that, it’s data coverage: every app, workflow, edge case, process, etc. sitting in private stores or someone’s head. Ultimately, while we make tremendous strides in more efficient model architectures, and clusters like Stargate equip us with zettaflop-scale compute, we really aren’t making rapid progress collecting the data we lack. We’ll soon live in a world where we have the methods & compute to accelerate scientific progress or economic growth, but not the data. And we’re already there today: frontier models would surely be as good at accounting/many medical tasks/legal advice as they are at software engineering if we only had the same pretraining & RL coverage as we did for code. I really want to drill this in: The speed at which we automate the economy is going to be directly rate-limited by our ability to collect data about it. Worth noting that under this assumption, with data as defensible and directly proportional to economic & scientific progress, data should also be considered a national strategic asset like compute. Imagine what we’d do in a world where we had a Manhattan Project-effort for AI and needed to mobilize data collection as a limiting factor. We should be concerned about China, with greater state capacity and authoritarian economic control, being capable of mobilizing data collection at national scale, potentially compounding their economy and scientific output faster than us down the line. A Stargate for data I’m leaving my complete ideas for a future post, as this one is already far too long, so I’d really like to pose the question here. Stargate exists because we organized trillions of dollars, international strategy, gigawatts around compute as a fundamental ingredient. What would equivalent ambition look like for data? Obviously, scaling data collection, a heterogeneous mass of information across the economy, isn’t going to be as clear as scaling compute, as a homogenous infrastructural effort. A core division will be first, coverage — all uncaptured knowledge sitting across the economy/science/physical world and all that simply isn’t recorded — and, secondly, sheer volume in the domains we already train on: more hard math tasks, more high-quality web text, way more coding data, more legal drafts, etc. I have a post coming soon which breaks down my proposals. There’s a lot of room for creativity. Quickly, we’ll probably want to start with a deep census of what we have and what we’re missing, predict what the 2030 model will still be bad at and work backward to what we should be collecting today. You can probably license a large amount, leveraging high lab valuations to buy datasets or companies altogether. There’s an adversarial nature to a lot of this collection with firms, so there’s lots of engineering to do this correctly. We should go convince important companies to turn off deletion policies, even if we’re not buying from them yet. Data flywheels in consumer products will be massive. Confidential training, government legislation for grant-funded research, running companies at a loss for their data, etc. We’re headed towards hundreds of billions in expenditure, national prioritization, and major data limitation on the horizon. We have a great opportunity to think creatively about what a megaproject for data would look like: How do we, deliberately this time, construct the next internet’s worth of data? Footnotes: [1]: I’ll probably soon publish my much longer post explaining my position on data efficiency and why the value of this data is still pretty high in most worlds regardless of new algorithms. [2]: The “AGI freeroll” bet: heads you win, tails ASI flips the world upside down anyways. [3]: We already see a glint of validation of this point, given the data market is strongly tilting towards ultra-high-quality agentic data, rather than unskilled labeling — niche expert workflows, live environments, and evaluations requiring increasingly obscure talent & knowledge — yet shows increasing, not decreasing, revenues.

English
10
15
121
20.9K
Mehrdad Farahmand
Mehrdad Farahmand@mifarion·
@butshaunn Truth. And if you’re in AI, the most brutal boss of them all: the shifting market.
English
0
0
1
5
Shaun Gold | Venture Comedy
Founders have a very strange definition of freedom. Everyone starts a company because they don't want a boss. Then they get customers. Now the customers are the boss. Then investors. Then payroll. Then enterprise procurement. Then legal. Then compliance. Then the IRS. The average founder escapes one boss and somehow ends up with thirty.
English
2
4
20
265
Cohere
Cohere@cohere·
Cohere and there
Cohere tweet mediaCohere tweet mediaCohere tweet mediaCohere tweet media
English
9
7
268
16.7K