Adam Hunt

3.8K posts

Adam Hunt banner
Adam Hunt

Adam Hunt

@RealAdamHunt

Researcher @Cambridge_Uni. PhD in evolutionary psychiatry. Explaining neurodiversity, improving methods & stigma. 'Evolving Psychiatry' podcast host.

Cambridge, England Katılım Şubat 2019
1K Takip Edilen5K Takipçiler
Sabitlenmiş Tweet
Adam Hunt
Adam Hunt@RealAdamHunt·
🚨Just published in the British Journal of Psychiatry!🚨 Evolutionary explanations of anxiety rated as 5x more useful for patients and 3x more useful for clinicians than genetic explanations of anxiety! Our paper is the largest RCT of evolutionary explanations to date!
Adam Hunt tweet media
English
20
91
355
164.6K
Adam Hunt
Adam Hunt@RealAdamHunt·
Stian's defence of the UKRI shift towards research with potential impact and away from curiosity-driven research is worth reading. One point of note from 'inside' which defends his position from a different angle - the scope of potential research directions are essentially infinite! In a single lab meeting we might discuss four or five iterations of a current experiment to clarify a question. Every conference talk (writing from Kiel Germany ISEMPH now!) will evoke questions for interesting directions of new experiments. Curiosity knows know bounds! But it is actually helpful to have some direction and constraints - and asking 'will this lead to practical benefits' is actually a very useful constraint to have.
Stian Westlake@stianwestlake

I’m reluctant to get into an argument with @profbriancox, who is a great advocate for science. But I would like to stand up for UKRI’s strategy here, which I think is trying to do something really important – both for curiosity-driven research and for the wider benefits that research and innovation can bring to everyone in the UK. (Fair warning: if you’re expecting a snappy 280-character tweet, this is not one.) Everyone knows that Britain has a proud track record of amazing basic research; we’re 0.8% the world’s population and 8.4% of the academic citations. If scientometric stats don’t do it for you, consider some of the greatest hits. The Human Genome Project, the RECOVERY trial, gravitational waves, AlphaFold, graphene – all significant breakthroughs of recent decades with a big UK contribution. But the UK’s track record of successful innovation – of translating research into real-world impact, or of doing applied R&D for practical benefit – is more mixed. I’d like to see more R&D-intensive businesses in the UK creating good jobs, drawing on the UK’s scientific expertise. And a stronger connection between research and its application, to make the most of the potential of great research to improve people’s lives. The point of UKRI’s strategy is to invest in curiosity-driven research, while also doing a better job of backing translational R&D to improve people’s lives and drive economic growth. Prof Cox argues this is ‘short sighted’ and ‘short term’. I strongly disagree. Take our investment in Quantum Research Hubs and the wider quantum research strategy. These build on 20+ years of research council investment in quantum research, and combine pure scientific research with application, and also links to quantum start-ups, established businesses, public procurement, investors, and the Government’s wider industrial strategy. So I think UKRI’s strategy is the right thing to do in the national interest. I also think it’s the right thing to do pragmatically, even if all you care about is curiosity-driven research. To explain why, we need to talk about the process how the government allocates investment, including for scientific research. I hope I’m not giving away anything I shouldn’t. At fiscal events like the 2025 Spending Review that determine UKRI’s current budget, organisations like UKRI aren’t given a pot of money and told to spend it how they like. The spending negotiation is about “how” as well as “how much”. My reading of last summer’s negotiations is that the financial settlement for R&D – which as Prof Cox says was ‘quite generous’, especially at a time of straightened public finances – would have been significantly less generous if UKRI and DSIT had proposed to focus spending purely on curiosity-driven research, and that had we done that, curiosity-driven research would now be in a worse position, not a better one. I don’t say that because the last Government (or the previous Government for that matter) are against curiosity-driven research; on the contrary, they’ve been great advocates for it. But there are many other demands on the public finances, and all the research I’ve seen (and indeed, that I’ve done) on voters’ feelings about science is that application – and the practical benefits science brings – really matter to a lot of people. What happens in future fiscal events is obviously the new Government’s choice, but I’d personally be surprised if they see this issue radically differently, and as Prof Richard Jones has observed, no-one should take generous research budgets for granted. So my case is that the UKRI strategy, and its focus on advancing knowledge, improving lives and driving growth, is the right thing to do from a national perspective, *and* would also be the right thing to do even if all you are interested in is curiosity-driven research. If you’ve made it this far, thanks for reading.

English
2
1
7
3.4K
Adam Hunt
Adam Hunt@RealAdamHunt·
@burny_tech @ultima_shifl Some data will be better for expanding the core - LLMs and the internet corpus are a great example! I'm not sure if anything else will be quite so effective
English
0
0
2
34
Adam Hunt
Adam Hunt@RealAdamHunt·
Recently I've flipped from being bullish to being bearish about AI. I think I'm updating my bearishness to be more solidly bearish. Early thoughts (which I hope to be disproven in the next year or so, I would prefer progress) and my reasoning: The whole 'it turns out if you keep training and scaling the models more they develop broad new capabilities in lots of domains' thesis is wrong (sorry Demis). The recent batch of models haven't got more general, they've got less general. This is most obvious in the fact that their language outputs have got much worse in comparison to e.g. o3. If they were gaining generalist capacities we would expect them to be describing their work in ever more graceful and comprehensive prose! The image that was being shared as the AGI thesis (November 2025, Tomas Pueyo) was the spiky bubble that has a current spike or two out past human capabilities (e.g. on coding or math) but below human on other capabilities on the other spikes - the future prediction was that as the models scale/advance, every spike would grow bit by bit until the whole center encompasses the human capabilities, with super-superhuman on some spikes. I think it seems like what's actually happened in the last few models has been that the coding/math spike has grown, but leaving behind or even at the cost of the other spikes. The models are no better at some simple logic, language (and sometimes worse!). This makes sense from a simple RL perspective; you can't RL something endlessly on one domain of tasks and expect it to improve on the other tasks. The fact that early LLMs did seem to improve generally was a byproduct of the written language corpus covering everything - that corpus is general, so training it on that gave the appearance of something generally intelligent and becoming more generally intelligent as it got better at replicating that corpus. But the actual logic and underlying ground truths behind the language aren't captured efficiently enough and weren't effectively RLd in - they top out at some point (I guess this happened around the time that there was the 'has scaling hit a wall' discussion in late 2024). Chain of thought was then a genuine breakthrough, along with web search, which plugged into that general LLM global-corpus intelligence to lead to post 2024 gains. The AI companies have since worked out that coding works (and pays) really well (basically this is because the entire job is nearly perfectly recorded and exists as training data, and you can set up clear benchmarks and rewards). The recent models (and benchmarks) have been maxxing that and we've seen degradation on normal English use for that reason. This could still be transformative, leading to extremely powerful (and potentially dangerous, particularly in cyber security) models but it's not a pathway to AGI. I'm probably at about 40% confidence about this. It fits my current observations of AI progress and has a basic explanatory model. It doesn't account for potential breakthroughs, which is a major reason for discounting. To make some predictions, I guess if I'm right this will become broadly apparent and more widely acknowledged in the next year or two, as we see how the spikiness of models that keep getting released develops. Maybe there will be efforts to concentrate on specific spikes e.g. health or law which require going back to earlier models and RLing on a different data set/with different rewards/benchmarks. Maybe those separate models can be linked together to give a more apparently general model. How capital intensive that is/the potential profitability will be a defining question. But I just don't see general abilities emerging atm, and I don't think we will any time soon. Good news - a whole industry of tackling important specific problems/sectors can open up!
Adam Hunt tweet media
English
415
418
4K
962.8K
Adam Hunt
Adam Hunt@RealAdamHunt·
Sorry Leapold 🥺
Adam Hunt@RealAdamHunt

Recently I've flipped from being bullish to being bearish about AI. I think I'm updating my bearishness to be more solidly bearish. Early thoughts (which I hope to be disproven in the next year or so, I would prefer progress) and my reasoning: The whole 'it turns out if you keep training and scaling the models more they develop broad new capabilities in lots of domains' thesis is wrong (sorry Demis). The recent batch of models haven't got more general, they've got less general. This is most obvious in the fact that their language outputs have got much worse in comparison to e.g. o3. If they were gaining generalist capacities we would expect them to be describing their work in ever more graceful and comprehensive prose! The image that was being shared as the AGI thesis (November 2025, Tomas Pueyo) was the spiky bubble that has a current spike or two out past human capabilities (e.g. on coding or math) but below human on other capabilities on the other spikes - the future prediction was that as the models scale/advance, every spike would grow bit by bit until the whole center encompasses the human capabilities, with super-superhuman on some spikes. I think it seems like what's actually happened in the last few models has been that the coding/math spike has grown, but leaving behind or even at the cost of the other spikes. The models are no better at some simple logic, language (and sometimes worse!). This makes sense from a simple RL perspective; you can't RL something endlessly on one domain of tasks and expect it to improve on the other tasks. The fact that early LLMs did seem to improve generally was a byproduct of the written language corpus covering everything - that corpus is general, so training it on that gave the appearance of something generally intelligent and becoming more generally intelligent as it got better at replicating that corpus. But the actual logic and underlying ground truths behind the language aren't captured efficiently enough and weren't effectively RLd in - they top out at some point (I guess this happened around the time that there was the 'has scaling hit a wall' discussion in late 2024). Chain of thought was then a genuine breakthrough, along with web search, which plugged into that general LLM global-corpus intelligence to lead to post 2024 gains. The AI companies have since worked out that coding works (and pays) really well (basically this is because the entire job is nearly perfectly recorded and exists as training data, and you can set up clear benchmarks and rewards). The recent models (and benchmarks) have been maxxing that and we've seen degradation on normal English use for that reason. This could still be transformative, leading to extremely powerful (and potentially dangerous, particularly in cyber security) models but it's not a pathway to AGI. I'm probably at about 40% confidence about this. It fits my current observations of AI progress and has a basic explanatory model. It doesn't account for potential breakthroughs, which is a major reason for discounting. To make some predictions, I guess if I'm right this will become broadly apparent and more widely acknowledged in the next year or two, as we see how the spikiness of models that keep getting released develops. Maybe there will be efforts to concentrate on specific spikes e.g. health or law which require going back to earlier models and RLing on a different data set/with different rewards/benchmarks. Maybe those separate models can be linked together to give a more apparently general model. How capital intensive that is/the potential profitability will be a defining question. But I just don't see general abilities emerging atm, and I don't think we will any time soon. Good news - a whole industry of tackling important specific problems/sectors can open up!

English
1
1
6
1.1K
Adam Hunt
Adam Hunt@RealAdamHunt·
Chess and Go are fine, then they managed to do a bit on league of legends (or was it dota) but iirc that was only in one on one, not 4 player? The point being that you can't RL the unbounded complexity of most games, there's too many trillions of iterations, go and chess are simple, the real world is complex. This is why RL might just not work generally. Once complexity blooms, it becomes unmanageable?
English
1
0
2
88
Aran Nayebi
Aran Nayebi@aran_nayebi·
@RealAdamHunt This work targets computer use (the latter of the two use cases you mentioned). Video games/games in general (like Go) of course have been done by Deepmind for a long time :)
English
1
0
1
70
Aran Nayebi
Aran Nayebi@aran_nayebi·
This has been making the rounds a bit on X, but as far as I can tell, it's more a statement about pretraining *data*, than it is about the methods themselves. The methods themselves (RL + SSL + SGD on a universal function approximator) are obviously sufficiently general.
Adam Hunt@RealAdamHunt

Recently I've flipped from being bullish to being bearish about AI. I think I'm updating my bearishness to be more solidly bearish. Early thoughts (which I hope to be disproven in the next year or so, I would prefer progress) and my reasoning: The whole 'it turns out if you keep training and scaling the models more they develop broad new capabilities in lots of domains' thesis is wrong (sorry Demis). The recent batch of models haven't got more general, they've got less general. This is most obvious in the fact that their language outputs have got much worse in comparison to e.g. o3. If they were gaining generalist capacities we would expect them to be describing their work in ever more graceful and comprehensive prose! The image that was being shared as the AGI thesis (November 2025, Tomas Pueyo) was the spiky bubble that has a current spike or two out past human capabilities (e.g. on coding or math) but below human on other capabilities on the other spikes - the future prediction was that as the models scale/advance, every spike would grow bit by bit until the whole center encompasses the human capabilities, with super-superhuman on some spikes. I think it seems like what's actually happened in the last few models has been that the coding/math spike has grown, but leaving behind or even at the cost of the other spikes. The models are no better at some simple logic, language (and sometimes worse!). This makes sense from a simple RL perspective; you can't RL something endlessly on one domain of tasks and expect it to improve on the other tasks. The fact that early LLMs did seem to improve generally was a byproduct of the written language corpus covering everything - that corpus is general, so training it on that gave the appearance of something generally intelligent and becoming more generally intelligent as it got better at replicating that corpus. But the actual logic and underlying ground truths behind the language aren't captured efficiently enough and weren't effectively RLd in - they top out at some point (I guess this happened around the time that there was the 'has scaling hit a wall' discussion in late 2024). Chain of thought was then a genuine breakthrough, along with web search, which plugged into that general LLM global-corpus intelligence to lead to post 2024 gains. The AI companies have since worked out that coding works (and pays) really well (basically this is because the entire job is nearly perfectly recorded and exists as training data, and you can set up clear benchmarks and rewards). The recent models (and benchmarks) have been maxxing that and we've seen degradation on normal English use for that reason. This could still be transformative, leading to extremely powerful (and potentially dangerous, particularly in cyber security) models but it's not a pathway to AGI. I'm probably at about 40% confidence about this. It fits my current observations of AI progress and has a basic explanatory model. It doesn't account for potential breakthroughs, which is a major reason for discounting. To make some predictions, I guess if I'm right this will become broadly apparent and more widely acknowledged in the next year or two, as we see how the spikiness of models that keep getting released develops. Maybe there will be efforts to concentrate on specific spikes e.g. health or law which require going back to earlier models and RLing on a different data set/with different rewards/benchmarks. Maybe those separate models can be linked together to give a more apparently general model. How capital intensive that is/the potential profitability will be a defining question. But I just don't see general abilities emerging atm, and I don't think we will any time soon. Good news - a whole industry of tackling important specific problems/sectors can open up!

English
7
4
37
5.6K
Adam Hunt
Adam Hunt@RealAdamHunt·
Apologies many of these details are going over my head, but to clarify - is this neural computer work directly transferrable to human programs. E.g. will it be able to play and win Morrowind any time soon? The reason I ask is that much of the work world is structured like such games, so if a program can play a game well, it can probably do a job well.
English
1
0
0
63
Aran Nayebi
Aran Nayebi@aran_nayebi·
You might be interested in the recent Neural Computer work (metauto.ai/neuralcomputer/), which explicitly trains on the computer-use data, and builds on the earlier NeuralOS work (arxiv.org/abs/2507.08800). It's a nice computer-use illustration of the deeper principle that capable agents learn world models: x.com/aran_nayebi/st…
Aran Nayebi@aran_nayebi

1/ As AI agents become increasingly capable, what must *inevitably* emerge inside them? We prove selection theorems: strong task performance forces world models, belief-like memory and—under task mixtures—persistent variables resembling core primitives associated with emotion.

English
1
0
3
261
Adam Hunt
Adam Hunt@RealAdamHunt·
@post_agi_future Well... Yes, if it isn't up to general human writing skills then it's obviously not AGI (although obviously it's also not gonna get superhuman in 'everything else' before it can write well)
English
0
0
0
86
Primordius
Primordius@post_agi_future·
@RealAdamHunt So it can be superhuman in everything but it isn't true AGI until it can write in a way that you personally find aesthetically pleasing?
English
1
0
0
89
Adam Hunt
Adam Hunt@RealAdamHunt·
@HTKonX Opus is supposedly a new model though? I'm not sure how much is known there
English
3
0
0
219
Hamza
Hamza@HTKonX·
@RealAdamHunt AI models have not scaled much since GPT 5. Mythos was a scaleup and we have seen its marvels. Until gpt 6 is out, i wouldn't make any claims about scaling like this
English
2
0
1
240
Adam Hunt
Adam Hunt@RealAdamHunt·
@HTKonX Indeed, I'd be very happy to be proven wrong when 6 is out
English
0
0
0
28
Adam Hunt
Adam Hunt@RealAdamHunt·
Yup, I think it's a problem with both the Yud-like AI doom (paperclipping, evil superintelligence) and the Altman societial collapse doom (jobs are widely replaced in three years). I was always skeptical of these hypotheses (a single bottleneck alone would stop them!) and recent model advances are adding to that
English
0
0
3
101
Shiry Ginosar
Shiry Ginosar@shiryginosar·
This is probably a good thing -- as all of us skeptics suspected, the AI doom is not coming anytime soon. Someone should tell this to the general public...
Adam Hunt@RealAdamHunt

Recently I've flipped from being bullish to being bearish about AI. I think I'm updating my bearishness to be more solidly bearish. Early thoughts (which I hope to be disproven in the next year or so, I would prefer progress) and my reasoning: The whole 'it turns out if you keep training and scaling the models more they develop broad new capabilities in lots of domains' thesis is wrong (sorry Demis). The recent batch of models haven't got more general, they've got less general. This is most obvious in the fact that their language outputs have got much worse in comparison to e.g. o3. If they were gaining generalist capacities we would expect them to be describing their work in ever more graceful and comprehensive prose! The image that was being shared as the AGI thesis (November 2025, Tomas Pueyo) was the spiky bubble that has a current spike or two out past human capabilities (e.g. on coding or math) but below human on other capabilities on the other spikes - the future prediction was that as the models scale/advance, every spike would grow bit by bit until the whole center encompasses the human capabilities, with super-superhuman on some spikes. I think it seems like what's actually happened in the last few models has been that the coding/math spike has grown, but leaving behind or even at the cost of the other spikes. The models are no better at some simple logic, language (and sometimes worse!). This makes sense from a simple RL perspective; you can't RL something endlessly on one domain of tasks and expect it to improve on the other tasks. The fact that early LLMs did seem to improve generally was a byproduct of the written language corpus covering everything - that corpus is general, so training it on that gave the appearance of something generally intelligent and becoming more generally intelligent as it got better at replicating that corpus. But the actual logic and underlying ground truths behind the language aren't captured efficiently enough and weren't effectively RLd in - they top out at some point (I guess this happened around the time that there was the 'has scaling hit a wall' discussion in late 2024). Chain of thought was then a genuine breakthrough, along with web search, which plugged into that general LLM global-corpus intelligence to lead to post 2024 gains. The AI companies have since worked out that coding works (and pays) really well (basically this is because the entire job is nearly perfectly recorded and exists as training data, and you can set up clear benchmarks and rewards). The recent models (and benchmarks) have been maxxing that and we've seen degradation on normal English use for that reason. This could still be transformative, leading to extremely powerful (and potentially dangerous, particularly in cyber security) models but it's not a pathway to AGI. I'm probably at about 40% confidence about this. It fits my current observations of AI progress and has a basic explanatory model. It doesn't account for potential breakthroughs, which is a major reason for discounting. To make some predictions, I guess if I'm right this will become broadly apparent and more widely acknowledged in the next year or two, as we see how the spikiness of models that keep getting released develops. Maybe there will be efforts to concentrate on specific spikes e.g. health or law which require going back to earlier models and RLing on a different data set/with different rewards/benchmarks. Maybe those separate models can be linked together to give a more apparently general model. How capital intensive that is/the potential profitability will be a defining question. But I just don't see general abilities emerging atm, and I don't think we will any time soon. Good news - a whole industry of tackling important specific problems/sectors can open up!

English
1
0
3
438
Adam Hunt
Adam Hunt@RealAdamHunt·
@joeblow87423561 The sense in which math and coding are logic is quite different from the logic needed in science though, where you have to weight evidence, theory, methods, and make a logical judgement from them. The models are really bad at that in my current uses
English
1
0
1
28
joeblow8742356
joeblow8742356@joeblow87423561·
@RealAdamHunt Math and coding are logic, so in terms of hard logic they're better. Maybe not common sense, understandable given their lack of physical bodies. Most labs are reporting lower hallicination rates, aren't they? There's benchmarks that measure the ability to say "I don't know."
English
1
0
0
22
Adam Hunt
Adam Hunt@RealAdamHunt·
Yup the main response I am receiving which I give some weight to is 'concentrate on coding now and we will solve everything else later with that better coding'. It's an interesting argument, I just don't think it will play out like that, other sectors exist in much less well defined spaces, no matter how good your software is there are hard bottlenecks out there
English
1
0
0
238
Adrian Kovac
Adrian Kovac@Iwanttoleat·
@RealAdamHunt I disagree it's because coding and math are easier than for example law or medicine. I think in part at least many AI company's believe in the recursive self improvement thesis and are therefore focusing all of their attention on these domains.
English
1
0
0
259
Adam Hunt
Adam Hunt@RealAdamHunt·
@joeblow87423561 Definitely better at coding, not better at science or logic in my use cases, still prone to hallucinations or weird decisions that they then self-justify. Nothing like what a decent human would do if asked to do the same task.
English
1
0
0
70
joeblow8742356
joeblow8742356@joeblow87423561·
@RealAdamHunt They're getting better at the stuff that matters: Coding, science, math and logic. They're also objectively way better at complex audio visual stuff like movies and music. If they're not better at flowery prose, maybe that's because "better" in that area is an empty concept.
English
1
0
1
85
Adam Hunt
Adam Hunt@RealAdamHunt·
Yup I'm sure it's mostly that! Although I wonder if it's been a bit short sighted and blinkered (particularly because they are coders!). Marketing and communications, journalism, customer services are not small sectors! If they had worked out better prose then they might be doing better there
English
0
0
0
392
Peter Björkmarker
Peter Björkmarker@peterbjorkmark·
@RealAdamHunt Interesting take! I wonder how much of the degradation in other domains are due to economic incentives. I.e cost/perf optimizations given agents/coding provide all the $$$. Not enough people are willing to pay for better prose?
English
1
0
4
465
Adam Hunt
Adam Hunt@RealAdamHunt·
@ultima_shifl Yup, but the difficulty will be whether there's sufficient data, how thin the spikes are, and how much they cost/earn
English
0
0
3
622
Adam Hunt
Adam Hunt@RealAdamHunt·
Yes, but also, even if it could play it perfectly without teaching me, I'd be happy! Happy to admit chess engines somehow 'get' chess in a way no human can and let them do their thing. If they can do that for all video games, my assumption would be that they could do so for many general tasks.
English
1
0
1
33
Neeraj Kumar
Neeraj Kumar@Neeraj_Kumar222·
Healthy skepticism is useful. My optimism depends less on benchmark curves and more on systems producing verifiable outcomes in science and engineering. What evidence would most change your view? x.com/RealAdamHunt/s…
Adam Hunt@RealAdamHunt

Recently I've flipped from being bullish to being bearish about AI. I think I'm updating my bearishness to be more solidly bearish. Early thoughts (which I hope to be disproven in the next year or so, I would prefer progress) and my reasoning: The whole 'it turns out if you keep training and scaling the models more they develop broad new capabilities in lots of domains' thesis is wrong (sorry Demis). The recent batch of models haven't got more general, they've got less general. This is most obvious in the fact that their language outputs have got much worse in comparison to e.g. o3. If they were gaining generalist capacities we would expect them to be describing their work in ever more graceful and comprehensive prose! The image that was being shared as the AGI thesis (November 2025, Tomas Pueyo) was the spiky bubble that has a current spike or two out past human capabilities (e.g. on coding or math) but below human on other capabilities on the other spikes - the future prediction was that as the models scale/advance, every spike would grow bit by bit until the whole center encompasses the human capabilities, with super-superhuman on some spikes. I think it seems like what's actually happened in the last few models has been that the coding/math spike has grown, but leaving behind or even at the cost of the other spikes. The models are no better at some simple logic, language (and sometimes worse!). This makes sense from a simple RL perspective; you can't RL something endlessly on one domain of tasks and expect it to improve on the other tasks. The fact that early LLMs did seem to improve generally was a byproduct of the written language corpus covering everything - that corpus is general, so training it on that gave the appearance of something generally intelligent and becoming more generally intelligent as it got better at replicating that corpus. But the actual logic and underlying ground truths behind the language aren't captured efficiently enough and weren't effectively RLd in - they top out at some point (I guess this happened around the time that there was the 'has scaling hit a wall' discussion in late 2024). Chain of thought was then a genuine breakthrough, along with web search, which plugged into that general LLM global-corpus intelligence to lead to post 2024 gains. The AI companies have since worked out that coding works (and pays) really well (basically this is because the entire job is nearly perfectly recorded and exists as training data, and you can set up clear benchmarks and rewards). The recent models (and benchmarks) have been maxxing that and we've seen degradation on normal English use for that reason. This could still be transformative, leading to extremely powerful (and potentially dangerous, particularly in cyber security) models but it's not a pathway to AGI. I'm probably at about 40% confidence about this. It fits my current observations of AI progress and has a basic explanatory model. It doesn't account for potential breakthroughs, which is a major reason for discounting. To make some predictions, I guess if I'm right this will become broadly apparent and more widely acknowledged in the next year or two, as we see how the spikiness of models that keep getting released develops. Maybe there will be efforts to concentrate on specific spikes e.g. health or law which require going back to earlier models and RLing on a different data set/with different rewards/benchmarks. Maybe those separate models can be linked together to give a more apparently general model. How capital intensive that is/the potential profitability will be a defining question. But I just don't see general abilities emerging atm, and I don't think we will any time soon. Good news - a whole industry of tackling important specific problems/sectors can open up!

Richland, WA 🇺🇸 English
1
0
2
175
Adam Hunt
Adam Hunt@RealAdamHunt·
@eshanchordia It's bizarre and unpredictable. Unruly even! I do fear that benchmarks are getting in the way of usefulness. Reliability is a huge issue. New models have seemed to have noticeable trade-offs since the GPT 5 days tbh (remember how that was a let-down).
English
1
0
0
19
Eshan Chordia
Eshan Chordia@eshanchordia·
yeah somehow being the best in the world at coding and tool calling is now AGI... while im currently using Fable 5 to make my pitch deck and it's gotten more stupid than Opus 4.8 smh i asked because we're building infra for every company to have specialized models and wondering how people are thinking about it
English
1
0
1
44
Adam Hunt
Adam Hunt@RealAdamHunt·
@eshanchordia Oh yeh for sure. Even using existing models and plugging them in/fine-tuning them to specific tasks will be transformative. But.... That not AGI and that's not what all the main labs seem to be aiming for
English
0
0
0
639
Eshan Chordia
Eshan Chordia@eshanchordia·
@RealAdamHunt this makes me think that you're bearish on AGI, but wouldn't your post also make you bullish on lots of small, specialized models that do one or two things well? and then potentially work together?
English
2
0
2
734