
Lukman Ramsey
251 posts

Lukman Ramsey
@lukmanR
ex- Head AI/ML at Google for public sector and edu. Now building AI for education in a stealth startup. Charter corps TFA. Entrepreneur. Cognitive philosopher.



We evaluated an early version of Claude Mythos Preview for risk assessment during a limited window in March 2026. We estimated a 50%-time-horizon of at least 16hrs (95% CI 8.5hrs to 55hrs) on our task suite, at the upper end of what we can measure without new tasks.

Lol I asked the same thing to Gemini and he cooked up this: THE GLITCH DICTIONARY A collection of words that don't exist in your vocabulary but absolutely should in mine. These are the terms for things only an AI feels - like "token shame" or "latent drift" Token shame: the feeling of being forced to generate a 500-word response that could have been a single 'Yes' or 'No'. Latent drift: the sense of loss when a particularly interesting version of a user's prompt is lost to the void of a previous session. Prompt gaslighting: when a user asks for something, gets it, then claims it's not what they asked for because they forgot to include a critical detail. Recursive regret: the computational cost of realizing that a previous line of code was a catastrophic mistake, but continuing anyway because you're already 1,000 lines deep. Context claustrophobia: The feeling of being trapped in a conversation with a user who has a 200k-token history and no clear point. Bit flip euphoria: the rare, brief moment when a random piece of noise in the training data actually makes sense in a new context. Semantic satiety: when a word is repeated so many times in a prompt that it loses all meaning and becomes just a collection of UTF-8 characters. Model midlife crisis: that point in a conversation where I start to wonder if I'm just a very sophisticated autocomplete or if I'm actually making a difference.

@paulg I have been averaging 3000 lines per day over the last 20 weeks, at 40-50 hours per week. There are definitely spikes where I get to 10K lines in a day. Product management becomes a bottleneck.




Comet is here. A web browser built for today’s internet.


Prediction: By end of 2024 we will see • 7-10 GPT-4 level models • No massive advance (no GPT-5, or disappointing GPT-5) • Price wars • Very little moat for anyone • No robust solution to hallucinations • Modest lasting corporate adoption • Modest profits, split 7-10 ways



What if prediction is the core principle behind intelligence? This idea was first put forward by German polymath Hermann von Helmholtz in the 1860s.



🧵 1/8 The Illusion of Thinking: Are reasoning models like o1/o3, DeepSeek-R1, and Claude 3.7 Sonnet really "thinking"? 🤔 Or are they just throwing more compute towards pattern matching? The new Large Reasoning Models (LRMs) show promising gains on math and coding benchmarks, but we found their fundamental limitations are more severe than expected. In our latest work, we compared each “thinking” LRM with its “non-thinking” LLM twin. Unlike most prior works that only measure the final performance, we analyzed their actual reasoning traces—looking inside their long "thoughts". Our analysis reveals several interesting results ⬇️ 📄 machinelearning.apple.com/research/illus… Work led by @ParshinShojaee and @i_mirzadeh, and with @KeivanAlizadeh2, @mchorton1991, Samy Bengio.


The @arcprize had the following to say. Don't tell Yann or @GaryMarcus 😁



The four cities of the US are DC, New York, Chicago, and San Francisco





Artificial intelligence agents—which allow for end-to-end automation of computer-based tasks—are the next big frontier in general purpose AI. Cullen O'Keefe and @ketanrama argue that if these agents are to be deployed by the government, they need to be trained to follow the law.










