pegzeus

1.8K posts

pegzeus

pegzeus

@pegzeus

I write about what I believe in - nfa obv

Katılım Ocak 2022
439 Takip Edilen398 Takipçiler
Adam
Adam@adamcreates_·
announcing internet rounds: the fastest and loudest way to raise money on the internet. get any app, agent, startup, or idea funded by a community that wants to see you win. (yeah, that's my mom and dog in the video)
English
150
42
585
224.6K
pegzeus
pegzeus@pegzeus·
@0xdrej Is grass venturing into non-web data collecting? Essay is pointing out that public web data is too limited and real value is in new data that isn’t on the web. Do you agree with that statement?
English
2
1
6
2K
Andrej
Andrej@0xdrej·
Calling what we're building the "Stargate for Data" might sound hyperbolic, but it's probably the most accurate description of both our ambition and the direction we're moving towards.
will depue@willdepue

A Stargate for Data Labs are on a trajectory towards >$100B/year of data spend by 2030. As we begin the trillion-dollar compute project, we need to think about the equivalent civilizational-scale effort for the other core ingredient: data. At the foundation of the scaling revolution is a simple empirical law: deep neural networks improve smoothly, near magically, as you scale two things in proportion — (1) the size of the model and (2) the amount of data you train on. And despite the scaling laws being brutally diminishing, we’ve successfully bitten the bullet of logarithmic scaling with exponentially larger clusters and datasets, and received incredible new capabilities in return. But this exponential scaling is bound to hit some limits. Oddly enough, compute has compounded fairly smoothly without limit, with trillions flowing into hypercluster buildout. Instead, we’re starting to hit the limits of an exponential demand for data. Gone are the days of being purely in the compute-limited regime, where we had effectively infinite internet data but never enough GPUs, we’re now entering a data-limited regime. Luckily, this limitation is coinciding with staggering improvements in AI capabilities. Incredibly, we seem to have a real line of sight towards automating a majority of knowledge work with the methods we have today. RL + pretraining, and the data for each, will be generally sufficient to achieve most economically valuable tasks, given some minimal algorithmic progress and continued compute scaling. In a data-limited world, economic progress & scientific acceleration will be directly bottlenecked by our coverage in each domain. We need to see data collection as imperative, deserving the same civilizational ambition we’ve given compute. The internet as a one-time subsidy It’s underrated how much all progress in AI owes everything to the blessing of the internet, this one-time civilizational subsidy to deep learning, decades of unintentional accumulation of a perfect dataset: every book, blog post, image, video, paper, discussion, etc. all digitized and freely available. Without the internet, we’d likely see comparably minimal progress in AI today, and in fact, if you notice where systems currently underperform, it’s almost always a domain where web coverage is limited and data is private, expensive, non-digitized, or non-existent. But we’re running out of it. There are only about 300 trillion tokens of useful public human text, and the internet doesn’t produce nearly enough new high-quality data to match what scaling demands — we’re soon to hit the limits of public data for pretraining. And though the advent of RL bought us reprieve — chain-of-thought RL needed a new form of untapped data, gradable math & coding tasks, also available online — we’re quickly running dry of hard tasks for RL as well. Why do we need so much data anyways? Humans learn comparably in far less time, needing just one textbook where language models might need the equivalent of hundreds to learn a new topic. It’s possible we discover methods that are massively more data efficient — synthetic data, data efficient architectures, other exotic algorithms — but fundamental progress is slow and highly unpredictable, and the recipe we have just works today. And, while I’m wary of getting too deep here, even arbitrary data efficiency can’t replace data that just doesn’t exist in the first place. There’s a massive amount of missing information on the web: the dark matter of the internet — tacit knowledge, undocumented processes, etc. — most of which was never published and lives only inside organizations, the physical world, or just in people’s heads. I’ll leave it here and say, for reasons far longer than I can fit in this post [1], it’s best to operate on the assumption that our insatiable desire for data will continue as it has for the last decade. There will be >$100B/year in data spend by 2030 We’re not screwed yet, of course. Only a fraction of useful data in the world is on the public internet, the rest is stored inside private datasets, corporations, personal archives, universities, governments, and otherwise. Labs can and will continue to license these private datasets, or create them from scratch, like Anthropic’s book scanning project. And we’ll increasingly task human experts to manufacture new high-quality data, with a large fraction of hard RL training tasks already being sourced this way. But collecting this data, unlike before, will be expensive. As the free internet dries up and demand for data rises, we should see labs investing equally in data as compute, likely spending a significant fraction of their compute budgets on data. As we see trillions spent on compute, we should also expect hundreds of billions spent on data (human data & collection budgets), given their equivalent importance. And, notably, data spend is already tracking this way: total data spend across vendors, not counting internal lab efforts, is already roughly $7 billion per year. It’s quite reasonable we’ll see >10x by 2030. Data is the moat Data becoming increasingly private will also majorly shift the competitive landscape. While compute is a commodity — everyone buys the same chips and builds the same clusters — data really isn’t. The big reason why frontier models have felt eerily similar to one another, until now, is they were trained on substantially the same internet (pretraining data variability across labs seems pretty low). As labs diverge onto more exclusive, manually collected corpora, I think models will begin to increasingly diverge. OpenAI pulling ahead in mathematics and Anthropic in cybersecurity isn’t an accident. I really think laser-focused collection of high-quality midtraining tokens, custom RL tasks, environments, with dedicated research effort, has driven much of the visible progress in the last year. James Betker has an excellent blog about “the ‘it’ in a model is the dataset”: model architecture and compute buy you efficiency and order-of-magnitude performance, but ultimately, models, of any architecture, are such incredible approximators of their dataset that the core meat of a model boils down to just that, nothing else. Data is a major moat. AGI long, ASI short As I’ve tweeted before, I’m confident that, despite the narrative, the data labeling industry will continue to fuel great businesses and be an excellent AGI long, ASI short. The argument is just: By the time the AGI labs no longer need data, it’s probably over for everything else too [2]. In this frame, the last companies left should be the data companies, as the last speck of economically relevant data is sucked in. And these companies are already among some of the fastest-growing companies in history: Mercor, founded three years ago, is rumored to be doing $2 billion in revenue with something like a few million expert labelers under contract. While these businesses are very non-stationary, what type of data is needed shifts constantly, I don’t think that diminishes their value. The long-tail of the economy is long, and the value isn’t diminishing as you extend farther into more obscure information: as models get more capable, the value of the marginal dataset goes up, not down. Automating a full job means covering its full distribution of tasks, tools, edge-cases, and long-horizon loops. There’s some O-ring logic to it: a dataset that buys a 1% bump can justify a previously unjustifiable collection cost when it’s the difference between a system that does 99% of a job and one that does all of it [3]. The competitive dynamics of the data industry are still evolving but as demand for data is increasingly niche, ultra high-quality, expert-generated, I think we’ll see real consolidation. Again, contra-narrative, we’ll probably see true competitive differentiation built on brand, quality control of data (which, from personal experience, can vary massively), as well as in network effects from the talent networks themselves over time. We’ve already seen rapidly shifting data type demand work in favor of incumbents, benefiting those with early knowledge of where the market is headed. The binding constraint It’s truly remarkable that we seem to have the recipe — pretraining + RL — to absorb most economically valuable work, despite being far from a lot of what we expected from “AGI”. The same way chess engines revealed we never needed general intelligence to solve chess, as we originally thought, we’ll soon realize that software, mathematics, and the vast majority of the economy (including physical, just running ~3 years behind!) are the same. If recursive self-improvement or some other algorithmic breakthrough arrives, that’s wonderful, but we really don’t have to wait for it. The binding constraint between here and an automated economy isn’t that, it’s data coverage: every app, workflow, edge case, process, etc. sitting in private stores or someone’s head. Ultimately, while we make tremendous strides in more efficient model architectures, and clusters like Stargate equip us with zettaflop-scale compute, we really aren’t making rapid progress collecting the data we lack. We’ll soon live in a world where we have the methods & compute to accelerate scientific progress or economic growth, but not the data. And we’re already there today: frontier models would surely be as good at accounting/many medical tasks/legal advice as they are at software engineering if we only had the same pretraining & RL coverage as we did for code. I really want to drill this in: The speed at which we automate the economy is going to be directly rate-limited by our ability to collect data about it. Worth noting that under this assumption, with data as defensible and directly proportional to economic & scientific progress, data should also be considered a national strategic asset like compute. Imagine what we’d do in a world where we had a Manhattan Project-effort for AI and needed to mobilize data collection as a limiting factor. We should be concerned about China, with greater state capacity and authoritarian economic control, being capable of mobilizing data collection at national scale, potentially compounding their economy and scientific output faster than us down the line. A Stargate for data I’m leaving my complete ideas for a future post, as this one is already far too long, so I’d really like to pose the question here. Stargate exists because we organized trillions of dollars, international strategy, gigawatts around compute as a fundamental ingredient. What would equivalent ambition look like for data? Obviously, scaling data collection, a heterogeneous mass of information across the economy, isn’t going to be as clear as scaling compute, as a homogenous infrastructural effort. A core division will be first, coverage — all uncaptured knowledge sitting across the economy/science/physical world and all that simply isn’t recorded — and, secondly, sheer volume in the domains we already train on: more hard math tasks, more high-quality web text, way more coding data, more legal drafts, etc. I have a post coming soon which breaks down my proposals. There’s a lot of room for creativity. Quickly, we’ll probably want to start with a deep census of what we have and what we’re missing, predict what the 2030 model will still be bad at and work backward to what we should be collecting today. You can probably license a large amount, leveraging high lab valuations to buy datasets or companies altogether. There’s an adversarial nature to a lot of this collection with firms, so there’s lots of engineering to do this correctly. We should go convince important companies to turn off deletion policies, even if we’re not buying from them yet. Data flywheels in consumer products will be massive. Confidential training, government legislation for grant-funded research, running companies at a loss for their data, etc. We’re headed towards hundreds of billions in expenditure, national prioritization, and major data limitation on the horizon. We have a great opportunity to think creatively about what a megaproject for data would look like: How do we, deliberately this time, construct the next internet’s worth of data? Footnotes: [1]: I’ll probably soon publish my much longer post explaining my position on data efficiency and why the value of this data is still pretty high in most worlds regardless of new algorithms. [2]: The “AGI freeroll” bet: heads you win, tails ASI flips the world upside down anyways. [3]: We already see a glint of validation of this point, given the data market is strongly tilting towards ultra-high-quality agentic data, rather than unskilled labeling — niche expert workflows, live environments, and evaluations requiring increasingly obscure talent & knowledge — yet shows increasing, not decreasing, revenues.

English
242
15
119
27.1K
Max
Max@MaxMinning·
No engagement? Watch this $ZRO (Still no engagement silence before the storm)
English
1
0
4
251
pegzeus
pegzeus@pegzeus·
@MaxMinning yeah im surprised it didn't move more, maybe already pre-priced in?
English
1
0
1
23
Max
Max@MaxMinning·
$lit went up like 20% on one of the biggest announcements in crypto history the golden goose for perp dex’s is net new users, and robinhood is the gatekeeper of tens of millions of net new users people say its overbought / insider traded I say we are clearly still in a bear market
English
1
0
1
131
pegzeus
pegzeus@pegzeus·
@lesabrefomo @mattytay funny that so many are bucketing both into the same group after todays announcements - feels like ppl are just repeating the same takes without dd the grass announcement has nothing to do with the equity/token debate
English
0
0
0
15
lesabre
lesabre@lesabrefomo·
putting these two in the same bucket doesnt make sense grass literally has no equity side of the things. They have an ownerless foundation that owns the network + pays labco service fee. The token is the only place where enterprise value accrues grass is much closer to an aave or morpho setup than anything. Grass is still missing the value accrual mechanism but (imo) that is the easiest fix of them all venice on the other hand is just the normal equity + token where the equity is where all value accrues and the token is junior grass at least has a credible path to becoming a fully profitable company where the value sits on the token side of things + value accrual link to the token via buybacks/etc (this is basically what every toke investors wants) venice has zero chance of this ever happening due to their structure
English
2
0
13
520
mattytay
mattytay@mattytay·
Last thing I’ll say on the Venice + Grass token debate today. It appears both teams are well-intentioned and are building good businesses. But the current dual token + equity paradigm is, at best, too risky for the largest pools of capital to underwrite. Fundamentally, the structure creates a misalignment between token holders and equity holders. Which is why crypto VCs will always require founders to give them both. That misalignment is one of the biggest reasons a massive amount of liquidity and institutional capital has not entered liquid crypto markets. Why did the U.S.A become home to the largest capital markets during the 20th century? A primary reason was that investors globally knew when they bought a US-sourced asset, there was rule of law protecting ownership and ensuring value accrued to the asset they held. Until crypto resolves the dual-structure problem, by choosing either token or equity as the singular economic unit of a business, the crypto industry will be stunted.
English
14
14
107
5.8K
pegzeus
pegzeus@pegzeus·
are we reading the same announcement? if anything node runners/miners may feel rugged as their $ amounts are lower than expected (not tokenholders) token accruing value is a different topic (not touched on in the ann) & independent of how the supply side is paid (USDC or GRASS)
English
0
0
2
90
Sam
Sam@0xCryptoSam·
@trevor_flipper Honestly I think GRASS was in a lose-lose situation and they chose to rug tokenholders to keep the business from heading down the canonical implosion most mature DePINs suffer from. Technically a win for Grass Network and miners, but tokenholders were sacrificed in doing so
English
3
0
7
2.1K
PMX
PMX@pmx_trade·
For the past 6 months, we’ve been heads down building the best on-chain prediction market on solana. We built dozens of new primitives on solana, including our own CLOB, parimutual AMMs, LMSR AMMs, and we beta tested everything with a small group of users in our telegram. In the past 2 months, we’ve finally made a breakthrough that we believe will drive the next stage of growth for PMX. It’s time to finally show the world what we’ve built.
English
8
13
37
4.1K
pegzeus
pegzeus@pegzeus·
good to have you in too! looking forward to the call. the lack of regular comms/data from the team is a bit annoying but i guess its part of what makes this an interesting setup. expecting a solid step up from the previous rev figures compute is getting most of the attention today but good data feels like a huge part of the ai stack too, and arguably the harder thing to get right (and a key differentiating factor for models). also interesting is how few ways there are to actually invest in this layer atm what's v appealing imo today w grass is that its already a successful vertically integrated data business: a decentralized node network for sourcing the data, plus in house infra to process and package it the node network is the real differentiator/edge compared to other players who often harvest data through buried terms without compensating anyone its also what makes it one of the only crypto x ai projects where having a token actually makes sense, it lets them bootstrap a large node network and distribute ownership in the process longer term, im also interested in how much the node network can expand its supplied services. today its scraped data, but the network will extend to live context retrieval and agentic web browsing w the grasshopper (both of which they're working towards) which further multiplies the addressable market, especially in an agentic world where more agents are browsing the web than humans i imagine grass will also start to look less like a company and more like a protocol over time, where access to data and the web becomes permissionless and available to all
English
0
0
1
44
lesabre
lesabre@lesabrefomo·
im long solana:Grass7B4RdKfBCjTKgSqnXkqjwiGvQyFbuSCUJr3XXjs due to terrible comms, most people have no idea that it is one of the highest revenue generating projects we have in crypto ($50m run rate as of q4'25, expected closer to 100m+ now) it's a real business with real customers (frontier models + mag 7) offering a valuable service that has practically infinite demand. Their edge being they can provide consented, auditable and cleanly sourced data at scale. very hard to have both. on the token: there is no equity. None. The revenue entity is owned by an ownerless foundation, so value can't get trapped in some VC's cap table. Value can only accrue to the token. generally think there is an education issue on grass due to the poor comms and once you piece everything together you are left with arguably the most successful crypto x ai project. Their token holder call is coming up on july 7th. I expect a lot more attention to the project once people realize how fast it is growing
English
27
14
179
40.3K
pegzeus
pegzeus@pegzeus·
@0xCryptoSam @bloomberg_seth @lesabrefomo How would the business work without the token though? Entire point is to connect bandwidth demand w supply For the business to succeed a healthy token is needed to maintain and grow the supply side
English
1
0
1
79
Sam
Sam@0xCryptoSam·
@bloomberg_seth @lesabrefomo Andrej most likely, but I assume someone could acquire the business and rug the token if they wanted to
English
1
0
0
397
LB
LB@lbexplorer·
I will make 8 figs on my $PUMP long.
English
8
2
58
30.8K
pegzeus
pegzeus@pegzeus·
@AlexMasmej I guess my point is an overarching creator token that ties together all experiment tokens would be a better way to bet on you (even though many creator token attempts have failed) I guess currently value accrues to showtime equity but that doesn't feel right
English
1
0
0
70
Alex Masmej
Alex Masmej@AlexMasmej·
@pegzeus Fair, try my apps to support me My next app should be free
English
1
0
2
85
Alex Masmej
Alex Masmej@AlexMasmej·
Showtime Technologies Revenue Recap Showtime: $0 revenue over 3 years — but earned backing from world-class investors with a 2-person team in a burning-hot market. Drakula: $1M in 1 year, in a semi-hot market, with 7 people. 10K: $1M in 6 months, in a semi-hot market, with 5 people. Captcha: $300K in 1 month, with 1.5 people. Pattern: every cycle team got smaller, execution got faster, revenue came sooner. AI is compressing my shipping speed. “Software is dead” except for crypto, because AI can’t trade without it; and social, because network effects beat slop apps every day. I am not done. Crypto is seasonal, but I will come back. Entrepreneurially, I am invincible with infinite runway. LFG!
English
6
1
32
4K
Star
Star@stardotfun·
THE SHOT three startups pitch a panel of sharks live on stream. the audience invests in real time. may 26th, 3PM UTC. live on X and star. RSVP below.
English
31
16
176
23.5K
du
du@thedulab·
Some exciting updates to share: 1. I've recently expanded to doing commercial relocations as well Any founders moving their company over to NYC and are looking for the perfect office space, reach out! Includes securing housing for your entire team, furnishing, moving logistics, etc. Whatever you need, however fast or specific you need it, I got you 2. I've also gotten a lot of DM's from people interested in the original residential service but found it difficult to afford without their employer sponsoring their relocation. So I'm coming out with a solution Anyone moving to NYC soon and need to find housing, join the Du York community (linked in the next tweet). This will be a lite version of the original service where I won't be sourcing and securing the housing for you, but rather sharing all the ball knowledge required to help you navigate the process yourself. Will also be providing alpha on lifestyle design for after you've settled in, so feel free to join if you've recently touched down and feel like you haven't quite found your footing yet as well Anyways stoked to continue making new friends and doing cool things together in the greatest city in the world. Hmu / share / tap in / don't be poor and homeless
du@thedulab

Finally figured out how to monetize my love for NYC Introducing...

English
8
6
192
45K
pegzeus retweetledi
Star
Star@stardotfun·
every week, one founder. funded live on stream.
English
57
23
344
93.9K
Jackson Dahl
Jackson Dahl@jacksondahl·
Who is the most underrated thinker you know? Who is the person you'd most like to hear a long, deep conversation with? About their philosophy, humanity, motivations? Whose voice needs amplifying? Who do you rush to listen to anytime they speak (or write)? Who never gets old?
English
63
3
112
28K
Captcha
Captcha@CaptchaApp·
Everyone welcome @blknoiz06 to Captcha, thanks for blessing us with your *checks note* human presence. The CT king is back in town.
Captcha tweet media
English
22
12
106
17.8K
pegzeus
pegzeus@pegzeus·
3) verified accounts (linked to X) shouldn't be able to change username (or they lose their verified status) feature idea: - let me gift/reimburse costs (or more) to someone who wrote something valuable. (e.g. I ask a question -> get a great response, I should be able to reward them as much as I want without retweeting/answering
English
0
0
0
9
pegzeus
pegzeus@pegzeus·
Things that are a bit confusing 1) Pay/earn structure? -> what are the economics? People think you pay 1$ but its not. Seems you pay around 3 cents to post, earn 3 cents if someone replies/reposts? (Token is bought in the background but user can convert after fees). 2) Bot accounts of real/verified people are confusing (e.g. alon). - Crossposts/bot replies need to be differentiated - What's the idea behind the bot reply feature of known accounts? Interesting for bootstrapping purposes, but did the users opt in? Feels like these are mostly posting slop atm, which goes against the idea of the app. - If users can have personal bots that generate revenue using their owner's custom knowledge/data it could be interesting, but a bit confusing atm
English
1
0
0
27
Alex Masmej
Alex Masmej@AlexMasmej·
What would y'all like to see on @CaptchaApp? Open to feedback for new features
English
17
1
33
6.2K
kel.
kel.@kelxyz_·
If you haven’t, I suggest listening to the audiobook “Devil Take The Hindmost” on youtube. Yes the free audiobook is better than the book. You’ll understand why this is recommended material around Q1 2027.
English
13
2
102
8.8K