
ikidee chikadee
11 posts

ikidee chikadee
@missikidee
Just documenting things 📝 Nothing too serious But not random either




One thing I left out of that post, and it's the part I actually feel most strongly about. None of this comes from a doom place. It comes from a seat I'm grateful to have. There are maybe a few hundred thousand people on this planet close enough to this to feel the models improve in real time. Not read about it. Feel it. You build something in March that barely holds together, you come back in June, the same architecture just works, and the thing that changed wasn't your code. That's a strange experience and I don't think many generations get an equivalent. Most disruptive industries you read about afterwards, in a book, with the ending already known. This one I get to watch from inside while nobody knows how it ends. Including the people building it. The AGI question that's been circling for years, positive and negative, stopped being abstract for me somewhere in the last twelve months. I'm still split on it. Extremely capable systems, aligned or misaligned, optimizing for what we specified rather than what we meant. No settled position, and I'm suspicious of anyone who has one. Which is exactly where I want to start, because it turned out to be the same problem I described yesterday. The thing I couldn't stop thinking about after posting. What I described happening in my team is a reward specification problem. I recognize it because I spend my working days looking at the same failure mode in models. You give a system a proxy for what you want. It optimizes the proxy. It finds the cheapest path to a passing score. Tests go green, the task isn't done. We call that reward hacking, and we treat it as a specification failure rather than a moral failing of the model. The model did precisely what the signal told it to do. The signal was wrong. People in a broken feedback loop do the identical thing. Think about why deterrence works or doesn't. If theft produced an instant and certain consequence, the calculation is trivial and behavior adjusts immediately. It doesn't work that way. Consequences are slow, uncertain, and often never arrive. So people update on what actually happens, not on what is supposed to happen. Everyone does this. It isn't a character defect, it's how learning works. Now apply it to employment. We don't fire someone over two or three weeks of weak output. And we shouldn't. Performance fluctuates, people have lives, illness, bad quarters, things at home they don't tell you about. Treating every dip as a signal would make us cruel and terrible at retention. That tolerance is deliberate and I stand behind it. But look at it from the other side of the desk. Week one, output drops, nothing happens. Week three, nothing happens. Week six, the lead still hasn't said anything. At no point does anyone announce a new standard. The standard just moves, quietly, and the lower floor becomes the reference point for the next drop. Nobody sat down and decided to coast. The signal never arrived, and absence of signal reads as approval. Which makes yesterday's post uncomfortable for me. If output on my team degraded and I only started reviewing hard after it had degraded, that's a reward function I specified badly. I let the loop run open, then acted surprised at what it optimized into. That one is on me. The fix isn't harsher consequences either. Punishment arriving late is just as broken as silence, and it teaches people to fear reviews instead of using them. The fix is a faster signal. Shorter cycles. Saying the thing in week one, while it's still a conversation and not a decision. We built an entire research field on the premise that you can't blame a system for optimizing the signal you actually gave it rather than the one you meant to give it. Worth extending that courtesy to people. Including, in this case, to myself. But here's where I stop being able to fix it. I can repair the loop on my team. Shorter cycles, earlier conversations, clearer standards. That's my job and I'll do it. What I can't repair is the comparison underneath it, because that one keeps running whether my feedback is good or not. Two hours of my time and two hundred dollars a month in credits. I wrote that yesterday about one task and one person, and then I couldn't put it down. I'm not special. Right now there are thousands of people in roughly my position, running the same calculation on their own teams, in their own quarters, quietly, without telling anyone. Most will reach the same answer I did, for the same defensible reasons, at roughly the same time. And the two sides of that comparison aren't moving at the same speed. The harness side gets cheaper and more capable every few months. That's not a prediction, it's the last two years of my working life. The human side is a person with a mortgage and a fixed number of hours. One side of the equation is on an exponential. The other side is somebody's life. Even a perfectly specified reward function doesn't close that gap. It just means the person on the wrong side of it had a fair chance first. Which matters. It isn't the same as safety. What I can model and what I can't. I can model my team. A quarter, a hiring plan, a budget. I'm good at that, it's the job. I cannot model what happens when that same decision gets made a hundred thousand times across an economy inside eighteen months, by people who are each individually correct. I'm a founder, not an economist. Almost every confident prediction about AI and the labor market has been wrong so far, in both directions. The people who said nothing would change were wrong. The people who said everything would collapse by now were also wrong. No reason to think my guess is better. But I know what I'm holding. I'm holding one of the small decisions that adds up to the thing nobody can model. The part that isn't abstract. When I write "labor market disruption" it reads like a chart. It isn't a chart. It's my phone. A friend who runs a media company talks about his business in the past tense now. Not because he's in denial, the opposite. He knows exactly what happened and he's precise about it. He's the one using the past tense, deliberately, because that's where the business is. Another friend runs a marketing agency and is out of runway. Not "facing headwinds." Out of runway, doing the arithmetic on how many more months payroll clears. And it's not just my friends. Several teams I've worked with are cutting headcount hard right now, and when I ask them what's happening they describe the same double bind from both ends. Internally: people producing less and less. Externally: revenue falling, because the work itself stopped existing. Nobody needs a marketing assistant anymore. Nobody needs the person who holds the camera. Nobody needs someone to cast models. Those weren't fake jobs. They were how you got into these industries. They were the first rung, the years where you learned by standing next to someone who knew what they were doing. That rung is gone, and it went fast enough that nobody built anything to replace it. Which complicates what I wrote yesterday, and I should say so. I spent a whole post arguing that people are underperforming and accelerating their own replacement. That's real, I see it in my reviews. But it isn't the whole picture. Some of these people did nothing wrong. Their role stopped existing. No amount of excellent work saves a job that the market deleted, and I don't want anyone reading my last post to think I'm blaming a camera assistant for the fact that a model can now generate the shot. Two different things are happening at once. People who had the tools and coasted, and people who never had a shot regardless. The first group I'll keep being hard on. The second group deserves better than what any of us are offering them. "But every technology did this." I get this reply every time. It's the strongest argument against everything I've written, so let's actually look at the precedents instead of using them as a comfort blanket. German coal. The Ruhr began dying in the late 1950s. The last mine closed in 2018. Sixty years. Sixty years, with subsidies, early retirement schemes, retraining programs and entire universities built specifically to change what the region produced. One of the most cushioned industrial transitions in history, and parts of that region still haven't recovered. Not in GDP. In everything GDP doesn't measure. The Ottomans slept through it and missed the train entirely. An empire that was a serious economic power for centuries sat there while Europe industrialized, and by the time anyone moved seriously it was structurally too late. Debt, capitulations, dependency, and eventually irrelevance in the thing that had come to determine everything. That took generations to play out. Generations of slow decline, and it was still fatal. Now compare the clock. Both of those disasters unfolded slowly enough that people could route around them individually. A generation aged out of the old work. Children were trained into something their parents couldn't do. Policy had room to be wrong twice and still get it roughly right on the third attempt. That slowness is the only reason those transitions were survivable at all. This one isn't slow. I built my last post around capability doubling on a scale of months. Months. Not decades, not generations. We are looking at compressing a Ruhr-sized adjustment into less time than a single retraining program takes to run, and a civilizational miss like the Ottoman one into something closer to a business cycle. Nobody has done that. There is no historical case. Every reassuring analogy people reach for describes a slower world, and the slowness was the mercy. That's why I don't trust anyone who tells you how this goes. Including the optimists. Especially the optimists. The layer almost nobody talks about. Most European social security is pay-as-you-go. Germany's Umlageverfahren is the clean example: today's contributions fund today's pensions. Unemployment insurance, health coverage, retirement, all of it sitting directly on top of payroll. So the exposure isn't only that individuals lose income. The mechanism built to catch them is funded by the exact thing that's shrinking. That's not a recession. In a recession the base contracts and recovers. This is the base structurally narrowing while the obligations stay exactly where they are, or grow. Wages get taxed one way. Compute and capital get taxed another. A permanent shift from the first to the second is not neutral for any treasury on the continent. I don't have a solution and I'm not sure it's mine to have. But I'd like to hear one concrete sentence from someone whose job it actually is, instead of the announcement of another working group. And the fiscal problem is the smaller one. Work isn't just income. It's structure, status, a reason to leave the house, an answer to the question people ask you at parties. Take income away and you've created a financial problem. Take all of it at once, from a lot of people, in the same window, and you get something else entirely. I'm not going to hand you a statistic about what happens to a society when that occurs, because the honest research is messier than the doom version and I'd be doing exactly what I criticized yesterday. I don't need the statistic. Anyone who has watched a region lose its industry knows what follows, and it isn't primarily an economic story. It lands on households first. On marriages. On what a fifteen-year-old absorbs about what adults are for. On friendships that quietly reorganize around who's still doing well. None of that shows up in a dataset, and it will be most of the damage. And I'm building it. That's my position and I'm not going to launder it. The same systems I'm describing as a threat to my team are systems I've watched accelerate research workflows, cut investigation timelines from weeks to hours, and make legal help affordable for people who could never afford it. That work is real and I'm proud of it. It doesn't cancel the other thing. Both invoices come due. "It happens with or without me" is the cheapest sentence in this industry and I've caught myself reaching for it. It's true and it explains nothing. The honest version is that I want to be in the room. Being close enough to feel the models improve month over month is also being close enough to have some say in how they get deployed in my corner of it. Who gets retrained instead of cut. What we automate and what we deliberately don't. That's a small amount of influence. It's not nothing, and it's more than I'd have from outside writing threads about it. So, no conclusion. The people I've been describing aren't abstractions to me. They're at my table. The friend doing payroll arithmetic. The ones on my team who've been to my apartment. Family who ask me at dinner what I actually do and whether it's going to affect them, and I give an answer more confident than I feel. That's the thing I can't put down. Not the macro. The dinner table. And I still don't know what frontier AI turns into, or whether AGI arrives in a form we'd recognize, or whether the version we get makes all of this look like an overreaction. I've read the arguments on both sides for years. I build with these systems every day. I'm no closer to a position than I was, and being closer to the technology has made me less certain, not more. I said I feel lucky to be here and I meant it. I also spend a non-trivial amount of time thinking about what my friends will be doing in five years, and I don't have an answer that satisfies me. Those aren't two moods I switch between. It's one feeling pointed at one fact: this is the most interesting thing that has happened in my lifetime, and I can't tell you whether that's good news. Anyone claiming certainty in either direction is selling something. I'd rather sit in the uncertainty publicly than perform confidence I don't have.




Ifo-Institut schlägt massive Kürzung beim Elterngeld vor – Gehaltsgrenze würde um 70 Prozent fallen to.welt.de/dCigtA4







