Deebz&Sak

31 posts

Deebz&Sak

Deebz&Sak

@eyyeffass

Katılım Mart 2010
33 Takip Edilen13 Takipçiler
Deebz&Sak
Deebz&Sak@eyyeffass·
the context management challenge is underrated. deciding what the system can do autonomously is the hardest part
Steve · AI@stev_builds

I've spent the last months building agentic infrastructure on frontier models and leading the dev team doing it. Concretely: the harness layer that gives a model tools, memory and verification loops, and the orchestration on top of it where a planner delegates to parallel subagents and reconciles what comes back. Most of the real engineering is not in the prompt. It's in context management, failure handling, and deciding what the system is allowed to do without a human watching. I want to write down something that's been bothering me, because the gap between what these systems can do in July 2026 and what the people closest to them actually do with it has gotten absurd. First, where the capability actually is. METR tracks the length of task a frontier agent can complete autonomously. It has been doubling roughly every seven months since 2019, and since late 2023 that compressed to about every four months. Their February 2026 measurement put the leading model at a 50% time horizon of around 14 hours, meaning work that would take a skilled engineer most of a day. Eighteen months earlier the same number was in minutes. AI DigestAgentmarketcap Work that was genuinely painful six months ago now lands near principal level on a good day. Content pipelines, automation, image generation at a level of detail that was research-lab material two years ago. Being honest about the ceiling, because it matters: on SWE-bench Pro, which uses private and previously unseen codebases, top models fall from 70%+ on the older Verified benchmark to the low twenties, and lower on the fully private subset. OpenAI published a piece in 2026 arguing SWE-bench Verified no longer measures frontier coding capability at all. So it's enormous, real, and not finished. Both are true at once. Scalearxiv Whether this becomes AGI I don't know, and METR's data is software, ML and cybersecurity tasks, not general intelligence. For the kind of work I do, the distinction stopped mattering a while ago. What that capability has already done to markets. This is not a forecast. It's the current print. 53% of agency owners now say AI is a credible threat to the agency model, up from 44% the year before, and 68% of brands already run some in-house capability. eMarketer's 2026 report: worldwide ad spending grew 8.6% in 2025 while holding company revenues fell 1.2%. The Big Six went from 44.6% to 29.6% of US ad spend. Forrester's 2026 study with the 4As found nine in ten US agencies using generative AI and half using agentic AI for execution. Revenue Memo + 3 Freelance is harder hit. Upwork writing projects fell 32% year over year in 2025, the steepest drop of any category, eleven of twelve major categories declined, and entry-level project availability fell below 9% from 15%. Ramp's February 2026 "Payrolls to Prompts" study found more than half the businesses that spent on freelance platforms in 2022 had stopped entirely. MediabistroMediabistro I have lawyers, media company owners, agency owners and doctors around me. Some still don't see it. Others have been bleeding for a year because they told themselves clients would keep booking real shoots rather than settle for AI slop. That bet is losing. Ask the uncomfortable question honestly. Why spend five or six figures on a commercial when a competent operator with a few hundred dollars in credits ships something usable? A spec Liquid Death spot made with Veo 3 by a director duo cost roughly $800 in credits and about two weeks, and it spread widely as proof of how far the quality jumped. The IAB expects 39% of ads and online video to be AI-built or AI-enhanced in 2026, up from 30% in 2025. Why fight through models, makeup and set logistics when a new T-shirt design becomes a catalog cover in an afternoon? Shhots AIMarketing Week And the "marketing expert" whose entire output was a Canva graphic every few days and an Instagram post? Be precise about what happened there. That role was already hollow. AI didn't replace it. AI removed the last excuse for keeping it. That's the environment. Here's the part that doesn't fit. The capability curve is vertical. The people with their hands on it are getting less productive, not more. And I want to be precise about the mechanism, because "AI isn't delivering ROI" is the lazy version of this and it's not what I'm seeing. What I'm seeing is people treating ten minutes of prompting as a substitute for eight hours of work. Not as a tool that compresses eight hours into three. As a replacement for the whole day. The output gets pasted in, the ticket gets moved, the status update goes out, and the rest of the day is gone. It's a facade of working. And it rests on an assumption: that nobody is actually going to look. That the output is plausible enough, the volume is high enough, and everyone is busy enough that no one opens it and reads it properly. Someone does open it. It's me. I review this work. And what comes back is thin. Code that runs on the happy path and falls apart on the second edge case. Documents that read fluently and say nothing. Analysis that confidently restates the question. Nothing that survives thirty seconds of real scrutiny. There's a name for the downstream effect now. A June 2026 HBR piece by Oxford's Matthias Holweg and Babson's Thomas Davenport calls it knowledge decay: output that looks finished but contains errors or lacks substance forces colleagues to verify, correct and redo it, and the errors compound across teams until the organization's shared knowledge base degrades. The earlier "workslop" research from BetterUp Labs and Stanford's Social Media Lab found about half of respondents rated colleagues who sent this kind of work as less creative, capable and reliable than before, 42% as less trustworthy. The Next WebToutsurlemarketing That last number is the one I'd underline. The people doing this think they're getting away with it. The data says everyone around them has already downgraded their assessment. They've been repricing themselves for months without noticing. And here is the calculation they're forcing me to make. I don't want to make it. I'm going to be honest about that. But when I review a week of output and it's a facade, I'm sitting there with a real comparison in front of me. On one side: a salary, ongoing, for work I can't ship without rewriting it. On the other side: two hours of my own time building a harness with Codex or Claude Code, plus a couple hundred dollars a month in API credits, producing the same volume at a quality I can actually verify, running on a schedule, without a status meeting. The second option is not theoretical for me. I build these systems. I know what they can and can't hold. And for a growing set of tasks, I know it works, because I've watched it work. That's the trade every technical leader is quietly running right now, and the person on the other side of it usually has no idea it's being run. What I actually wish they'd do instead. The move is so obvious that watching people miss it is what pushed me to write this. You have access to the most capable tools that have ever existed for knowledge work. The correct play is to walk into the room with output nobody expected. To make the case in artifacts rather than in words: don't replace me, look at what comes out the other end when you invest in me. Every euro you put into this seat comes back multiplied and here is the receipt. Instead the message being sent is the exact opposite. It says: I have found the minimum. I am producing exactly enough to appear employed. And I'm doing it with the same tools you could point at this problem directly, for a fraction of what you pay me. You cannot send that signal to someone who builds automation for a living and expect them not to draw the conclusion. The part I don't enjoy. I've worked with some of these people for years. There are real relationships there, birthdays, difficult periods, people who showed up when it counted. I'm not somebody who can end that with a spreadsheet and feel nothing about it, and I don't want to become that. So I'm not writing this as a threat. I'm writing it because I'd rather people hear it from someone who is on their side than find out through a calendar invite. But I'll say the honest part too: they are not making it hard for me to justify. Every review cycle where the work doesn't hold makes the case build itself. And at some point the loyalty I'm extending stops being loyalty and starts being something I have to explain to people I'm accountable to. Where this ends up. The bottom rung is already going. Stanford HAI's 2026 AI Index found employment for software developers aged 22 to 25 fell nearly 20% from 2024 levels, while headcount for developers over 26 kept growing. The entry point into these careers is closing while the people already inside are busy proving they don't need to be there. Tech Times And the upside is documented, not hypothetical. AI-specialized freelancers command 25 to 60% higher rates than generalists in the same field. Designers who adopted early earn 40 to 60% more per hour than before. That gap didn't appear by accident. It went to the people who used the tools to produce more, not less. WE AND THE COLOR I'll grant the obvious objection. If capability keeps compounding the way it has, "good AI operator" has an expiry date too. Probably. I'm not selling permanent safety, nobody can. What I'm saying is narrower. You're standing at the interface between frontier AI and whatever comes next, with your hands on the tools, and most people don't get that seat. Use it to build judgment that can't be handed to a model in a prompt. Be the person whose work is the reason the system ships, not the person whose work the system replaces. Maybe that only buys time. I'd still take the time. Or keep shipping ten-minute answers to eight-hour problems, and let the person reviewing them do the math. Your call. And someone is reading it more carefully than you think.

English
0
0
0
1
Deebz&Sak
Deebz&Sak@eyyeffass·
@PrivOSAI oh this could save teams so much time, the manual check-in grind is real
English
1
0
0
15
PrivOS
PrivOS@PrivOSAI·
📱 Simplify Employee Check-ins with PrivOS Daily check-ins are essential for team alignment, but manual reporting often creates unnecessary overhead. Managers spend time collecting updates while employees switch between multiple tools to share progress. 🤖 With AI-powered analysis, managers can quickly identify underperformed employees, detect faults, and make informed decisions based on real operational data. 🔒 Built on a secure self-hosted workspace, PrivOS helps teams improve collaboration while maintaining control over their information. Transform daily updates into actionable insights with PrivOS. 🚀 #EnterpriseAI #AIWorkspace #TeamCollaboration #Automation #PrivOS
English
2
5
4
69
Deebz&Sak retweetledi
SparkDEX
SparkDEX@SparkDexAI·
🔥 Governance Proposal #2 is LIVE. The SparkDEX DAO is voting on a 5,000,000 SPRK burn. 🔸 0.5% of max supply 🔸 ~25% increase in total SPRK burned 🔸 Lower future inflation Voting ends Friday, June 26 at 3:00 PM UTC. 🗳️ Vote with your xSPRK: sparkdex.ai/governance/pro…
SparkDEX tweet media
English
22
29
69
9.9K
Cointelegraph
Cointelegraph@Cointelegraph·
🚨 UPDATE: Cardano founder Charles Hoskinson says he's "taking a break."
Cointelegraph tweet media
English
657
185
2K
822.1K
Deebz&Sak
Deebz&Sak@eyyeffass·
@Cointelegraph Honestly wonder who runs things while he's out. The foundation?
English
0
0
1
2.2K
Nous Research
Nous Research@NousResearch·
Hermes Agent now has Tool Search, so your agent only loads what it needs
English
159
251
3.4K
1M
Flare ☀️
Flare ☀️@FlareNetworks·
XRPL for issuance. Flare for what comes next. @HugoPhilion on the compute layer: privacy, trading, lending, and collateral for RWAs issued on XRPL. Clip from XRP Las Vegas with @paulbarron - watch the full talk → youtu.be/9V1TOLn12vo
YouTube video
YouTube
English
730
92
364
21K
mickle
mickle@xrpmickle·
It’s like watching my parents talk about AI
English
105
44
588
135.4K
Deebz&Sak
Deebz&Sak@eyyeffass·
Loving the afternoon vibes! ✨ Hope you're all having a great day & looking forward to what Jalisa creates! 🎶
English
0
0
0
18
Deebz&Sak
Deebz&Sak@eyyeffass·
OUR WEDNESDAY SALE STARTS NOW! http://eyyeffass.livejournal.com/18717.html
English
0
0
1
0
Deebz&Sak
Deebz&Sak@eyyeffass·
CLEARANCE SPECIAL FOR TODAY ONLY! While stock last! http://eyyeffass.livejournal.com/18645.html
English
0
0
0
0
Deebz&Sak
Deebz&Sak@eyyeffass·
NDP SPECIAL! CHECK IT OUT HERE: http://eyyeffass.livejournal.com/18217.html
English
0
0
0
0
Deebz&Sak
Deebz&Sak@eyyeffass·
$10 WEEKEND SALE. Everything at $10. Only thins weekend. While stock last: http://eyyeffass.livejournal.com/18116.html
English
0
0
0
0
Deebz&Sak
Deebz&Sak@eyyeffass·
FLEA MY PLACE over at sak's place was awesome. Pictures here: http://www.facebook.com/album.php?aid=230290&id=644201476&l=5c2fcdc124
English
0
0
0
0
Deebz&Sak
Deebz&Sak@eyyeffass·
Check out our latest collection: http://eyyeffass.livejournal.com/17828.html
English
0
0
0
0
Deebz&Sak
Deebz&Sak@eyyeffass·
New collection tomorrow! (:
English
0
0
0
0
Deebz&Sak
Deebz&Sak@eyyeffass·
WEEKEND SALE NOW ON!!! 18-20 JUNE ONLY. ONLY AT http://eyyeffass.livejournal.com
English
0
0
0
0
Deebz&Sak retweetledi
sakinahjam
sakinahjam@sakinahjam·
WEEKEND SALE NOW ON!!! 18-20 JUNE ONLY. HEAD DOWN TO http://eyyeffass.livejournal.com NOW! LIMITED STOCKS. (:
English
0
1
0
0