Shershebnev Alex

691 posts

Shershebnev Alex

Shershebnev Alex

@shershebnev

DevOps, DevRel, DevAskMeAnything @ZencoderAI

Portugal Katılım Temmuz 2009
37 Takip Edilen22 Takipçiler
Sabitlenmiş Tweet
Shershebnev Alex
Shershebnev Alex@shershebnev·
❓Which model is the best? Claude? GPT-5.4? Gemini? I see this question every single day on my X feed. But limiting yourself to one model is like a chef who only ever fries food. You're missing out on the real secret to why some people are actually getting ahead. 👇 Picking one model today is basically choosing to have a smaller tool belt. A good cook knows when to boil and when to fry. Sometimes you do both. You might boil something just to finish it in the pan later. That is how we should be thinking about LLMs. Why pick one when you can have all of them at once? If you run the same prompt through @AnthropicAI Claude, @OpenAI ChatGPT, and @Google Gemini, the results aren't just different. They are often the complete opposite. Claude usually feels more "corporate" to me. Gemini feels more natural for creative stuff like video scripts or blog posts. On multiple occasions I've put them side-by-side on technical questions where they just flat out disagree. ChatGPT tells me to use Approach A and specifically warns me against Approach B. Then I give Claude the same prompt and it picks Approach B, calling Approach A a bad idea. When you limit your models, you're limiting your own access to intelligence. You lose that second opinion. So how do you actually combine models? Here are a few ways. Side-by-side comparison. The simplest one - pass the same prompt to two different models separately, then combine the best parts from each into your own end result. Whether it is a blog post or code, you get to pick the winner for every sentence. Review chain. Take one model and make it produce something good for you. Then ask a second model to review it to find issues and suggest better options. Finally, let the original model or even a third one incorporate those suggestions. If you let multiple models work together, you are essentially replacing your slow review step with a much faster model. You might even get a much smarter reviewer too. That part is debatable, but the speed alone is worth it. Full council. Have one model write the code and three models review it. Or have multiple models each produce a draft and then synthesize the results. Autonomous orchestration. Even better yet, let one agent decide which models and subagents it needs for each subtask. Say you have a main agent based on Opus and you give it a task to build a website. It might spin up one subagent with @grok for image generation to create visual assets, another subagent on Codex to generate the frontend, and another one on Haiku for the backend. You are no longer even required to think about which models to choose or which models should work with or against each other. You just delegate that decision to an agent as well. This isn't for everyone yet. It's expensive. Paying for multiple subscriptions (or hey, here's an idea - get one subscription in Zenflow and use all those models) and token usage for agents talking back and forth adds up fast. But I really believe the future isn't about which model is the winner. It will be about which orchestration is the best. I feel lucky to be seeing this first-hand at @zencoderai . Watching a swarm of agents work together where one writes and three others review it feels totally different. You aren't just working with an AI anymore. It is a collective group working on whatever goal you set.
English
1
1
1
255
Shershebnev Alex
Shershebnev Alex@shershebnev·
First time being a panelist on a panel discussion @DevDaysEurope , it was a lot of fun to talk about AI and security in that format. And 5 star rating is a cherry on top :) Still have two “regular” talks to do tomorrow
Shershebnev Alex tweet media
English
0
0
1
26
Aakash Gupta
Aakash Gupta@aakashgupta·
Anthropic just mass-obsoleted every agent orchestration startup in a single launch. The screenshot tells the full story. That's a production fleet dashboard. 8 agents running. 247 completed tasks. Active status. MCP-connected to HubSpot, pulling deals, generating proposals, reading attachments. This isn't a demo. It's a managed production environment where you define the agent and Anthropic runs the infrastructure. The timing here is surgical. Four days ago, Anthropic blocked OpenClaw and every third-party harness from using subscription credentials. The message was clear: stop building on top of our consumer auth layer. Now here's the replacement. A first-party managed agent platform with fleet monitoring, production-grade MCP integrations, and prototype-to-launch timelines measured in days. Manus spent six months on five harness rewrites. LangChain spent a year on four architectures. Anthropic just shipped the managed version that eliminates the need to build one at all. The real bet: most companies don't want to build agent infrastructure. They want agents that work. Anthropic is pricing this into the platform the same way AWS priced server management into EC2. The 46% of enterprises citing "integration with existing systems" as their primary agent challenge just got a first-party answer from the model provider itself. Every agent startup that raised on "we make Claude reliable in production" just lost their pitch deck.
Claude@claudeai

Introducing Claude Managed Agents: everything you need to build and deploy agents at scale. It pairs an agent harness tuned for performance with production infrastructure, so you can go from prototype to launch in days. Now in public beta on the Claude Platform.

English
178
269
3.2K
1.1M
Shershebnev Alex
Shershebnev Alex@shershebnev·
Hi @sebuzdugan, it is both, actually. The infrastructure is the multi-model orchestration. Plan with Opus, implement with cheaper and faster Gemini. Our benchmarks show 2.7x cheaper per resolved task at equal or better quality vs using just e.g. Opus for everything And then the interface is what makes it accesible to everyone, not just devs who are used to working in the terminal, but also to your marketing team, or sales, or others that could benefit from AI agents And to ship to the whole org you need both
English
0
0
0
85
Sebastian Buzdugan
Sebastian Buzdugan@sebuzdugan·
@zencoderai if it’s “for your entire organization” you’re building an interface, not infrastructure
English
2
0
3
12.5K
zencoderai
zencoderai@zencoderai·
Introducing Zenflow Work. A proactive AI assistant for your entire organization.
English
10
29
347
15.3M
Shershebnev Alex
Shershebnev Alex@shershebnev·
Hi @ashcotXBT, Zenflow is much more than a workflow wrapper. But even at workflow level I believe we are much more flexible and powerful than Claude Cowork. And it comes on several levels - those workflows are flexible on their own, you or agent can customize them as needed. But the real power comes from multi-model orchestration, something you can’t really get in Anthropic’s products simply because they won’t have any other models. With Zenflow you can mix and match e.g. Opus, Codex, Gemini, Grok and whatever else you want. For instance, you might do the planning with Opus, but actual implementation is cheaper and faster with Gemini. I’ve just recently posted an article and a video on that topic, check out the pinned tweet in my profile. Combined with automation and integrations, it allows you to truly get the productivity boost promised by AI :) I suggest you give it a try and see for yourself
English
0
0
0
77
Jimmy Ashcot
Jimmy Ashcot@ashcotXBT·
@zencoderai just another workflow wrapper. what does it do that claude projects cant
English
1
0
1
9.7K
Shershebnev Alex
Shershebnev Alex@shershebnev·
@zencoderai Expanding beyond coding is great, now Zenflow can replace me at even more task 🫣
English
1
0
1
54
zencoderai
zencoderai@zencoderai·
Introducing Zenflow Work! A proactive AI assistant for your entire organization. 👉Define a goal → agents plan, execute, verify 👉Schedule automations that run until the job is done 👉Works with your tools - Jira, Notion, Google Docs, Slack, Gmail, and more 👉Start tasks from Slack, Telegram, Discord or Zenflow directly One interface is all you need Focus on done, not switching tabs
English
9
8
32
1.9K
Shershebnev Alex
Shershebnev Alex@shershebnev·
Stop asking “which AI is best.” That’s the wrong question. The people winning right now are using multiple models at once... And getting way better results. Here’s exactly how 👇 youtu.be/fJ6VwL7oIug
YouTube video
YouTube
English
1
1
1
89
Shershebnev Alex
Shershebnev Alex@shershebnev·
@phuctm97 Check out Zenflow from @zencoderai, I barely use any IDE anymore and manage all agents there. Can also use Claude Code and Codex directly inside the app
English
0
0
0
39
Minh-Phuc Tran
Minh-Phuc Tran@phuctm97·
Just reset my Mac and I don't even bother to install Cursor anymore. Just a year ago, Cursor is my daily driver and a massive success, now it's a forgotten product (at least for me & many friends). It's crazy how fast things change in AI.
English
115
4
328
48.4K
Shershebnev Alex
Shershebnev Alex@shershebnev·
I feel like “.agents” is the most logical location for all things agents. Would be great to see some unification across industry on that for sure. We moved to that directory some time ago @zencoderai. And on that note thank you for trying our product (hopefully you also had a chance to try Zenflow). Let me know if you have any feedback! :)
English
0
0
0
122
Melvyn • Builder
Melvyn • Builder@melvynx·
THIS IS THE WORST THING IN THE AI AGENT AREA
Melvyn • Builder tweet media
English
26
0
53
14.3K
Shershebnev Alex
Shershebnev Alex@shershebnev·
Having fun doing all those different short videos about all the new features of @zencoderai. Anyway, what you need to know about this past week’s updates is that you can have agent work continuously and repeatedly on some task (like following up on a pull request you (or agent) just created) until the task is complete. Imagine not checking the PR all the time and just letting the agent do all the work till the very end
zencoderai@zencoderai

What we released in Zenflow this week, March 23-29? - Create task related automations that agent will follow on until it is done. Created a PR from the task? Ask agent to follow it and address any comments until PR is merged - Instruct agent when to commit instead of committing after every step - You can now create draft PRs (in addition to regular PRs) right from Zenflow - Agents can ask you for clarifications mid-chat, then continue with you answer in "mind" - Get notified when agent review is done - Fixes and improvements across performance and reliability

English
0
0
1
83
Shershebnev Alex
Shershebnev Alex@shershebnev·
@BHolmesDev If you want to combine them both inside the same task, try Zenflow from @zencoderai, you can literally assign different steps of the task to different models seamlessly (or let the agent do that for you)
English
0
0
0
19
Ben Holmes
Ben Holmes@BHolmesDev·
I’ve used Opus 4.6 and GPT 5.4 on a mix of projects since release, and want to break down where I think they uniquely excel. It’s more nuanced than you’d think! Rigor of code - GPT 5.4. It goes the distance validating its work without asking. Opus needs explicit instruction to do this, and even then, it misses more edge cases. Clarity of code - Opus 4.6. Claude is a better communicator, which carries into the code. Variable names are clearer and less mechanical, which improves reviewability. This is very important since code review is the bottleneck for most engineering teams. It also adds the right amount of doc comments. GPT simply never comments or explains its work; it’s like working with an obtuse engineer that wants the solution to speak for itself. Sometimes it does, other times not. Similarly, rigor of plans goes to GPT 5.4, while clarity of plans goes to Opus 4.6. An interesting point though: GPT performs better talking through a strategy without a plan, while Opus needs planning mode to put in any rigor. I find myself forgetting plan mode altogether using GPT 5.4. Quality of research - toss-up. Opus spends longer researching with web search, but GPT spends longer studying the existing codebase. You may think codebase research matters more, but researching how others solve the same problem can be just as important. Maybe more important for greenfield. Quality of conversation - Opus 4.6. It’s just better to talk to, which matters using these things everyday. GPT 5.4 was clearly trained to challenge the user more, which results in a tendency to *always* say you are wrong. I’ve had bizarre interactions where GPT claims something is “not quite right,” the restates exactly what we’ve decided on in the last turn. On a personal level, it’s annoying. On a practical level, it makes iteration on a plan slower. THAT SAID, it takes sufficient pushing for Opus to challenge your thinking in this way. Simply say “I’m impartial” and ask questions to avoid that, as you would a person. Overall winner - Opus to make it work, GPT to make it good. I don’t have a good system of when to switch tools, but on average, I prefer Opus early on and GPT for optimization and discussing architectural decisions. Opus is also better for any design related tasks (but state management in frontend apps is better handled by GPT).
English
137
88
1.5K
203.7K
Shershebnev Alex
Shershebnev Alex@shershebnev·
@rezoundous I’m team “make them work together as a proper multi-agentic system”
English
0
0
0
19
Tyler
Tyler@rezoundous·
Are you team Claude or Codex?
Tyler tweet mediaTyler tweet media
English
628
30
702
311.1K
Shershebnev Alex
Shershebnev Alex@shershebnev·
@deedydas It’s not just FAANG, the second you go outside of your bubble, you’ll see that majority of people are not actually using AI or using it in a very limited capacity
English
0
0
4
2K
Deedy
Deedy@deedydas·
Google Senior Staff Engineer to me: “Yeah, I have no clue what Claude Code / Codex is but I hear it’s all the rage. No, I don’t really care, I just need GOOG to hit $400 and keep this job for 2-3 more years so I can retire!”
English
256
98
5.9K
851.5K
Shershebnev Alex
Shershebnev Alex@shershebnev·
@burkov Better way is have another LLM review it. I just recently wrote a post about different ways to leverage multiple llms for better outcomes. Checkout pinned post on my page
English
0
0
0
13
BURKOV
BURKOV@burkov·
When Claude shows me a three-page-long plan for a fix, I have two options: 1. Read it, confirm that it's all correct, and have it implemented wrong anyway. 2. Not read it, confirm it's all correct, and have it implemented wrong anyway. I prefer to choose option 2, then test, and then tell Claude that it implemented it wrong. How about you?
English
94
7
167
25.2K
Shershebnev Alex
Shershebnev Alex@shershebnev·
@zarazhangrui I feel the same way, but what helps me at least partially is removing window switching and having most of the things inside same app. I control agents and terminals inside Zenflow from @zencoderai. Technically also can move browser there as well, I guess :)
English
0
0
0
444
Zara Zhang
Zara Zhang@zarazhangrui·
My single biggest pain point right now: AI-induced attention deficit I have 5 Claude Code sessions running 10 Terminal tabs open 50 browser tabs open 100 X articles in my bookmarks When there are multiple AIs working for you, you're constantly task switching When you're waiting for AI to output, you're constantly distracted Because what used to be hard is now so easy, you always wanna do more, more, more Because of the speed of the AI news cycle, you feel a constant need to be on this app Add all this up and the result: you live in a constant state of distraction & delirium. Focus & deep work is harder than ever If anyone has figured this out pls teach me how
English
512
129
2.1K
241.9K
Shershebnev Alex
Shershebnev Alex@shershebnev·
@ibuildthecloud On my personal vibe coded project I have e2e test suite, which AI can run everytime it does something, and since I’ve added them (well AI did) I did notice a reduced number of follow ups I have to do to fix things
English
0
0
1
437
Darren Shepherd
Darren Shepherd@ibuildthecloud·
I'm not kidding. Has anyone seen a unit test that AI has written that provided any value whatsoever? This is truly fascinating how useless they are.
English
182
18
973
118.8K
Shershebnev Alex
Shershebnev Alex@shershebnev·
Take a look at this short video where Opus spawns three subagents based on Claude, Codex and Grok to review the code x.com/shershebnev/st…
Shershebnev Alex@shershebnev

🚨@OpenAI recently added subagents to Codex, @AnthropicAI had them for a while in Claude. But what if you could make them work together? Here I'm asking three different CLIs/models to roast the feature implemented by Opus 4.6. Gpt 5.4 through Codex, Opus 4.6 through Claude Code, and @grok through @zencoderai CLI all work under the same orchestrator. You are no longer limited by just one model type (or CLI)

English
0
0
1
125
Shershebnev Alex
Shershebnev Alex@shershebnev·
❓Which model is the best? Claude? GPT-5.4? Gemini? I see this question every single day on my X feed. But limiting yourself to one model is like a chef who only ever fries food. You're missing out on the real secret to why some people are actually getting ahead. 👇 Picking one model today is basically choosing to have a smaller tool belt. A good cook knows when to boil and when to fry. Sometimes you do both. You might boil something just to finish it in the pan later. That is how we should be thinking about LLMs. Why pick one when you can have all of them at once? If you run the same prompt through @AnthropicAI Claude, @OpenAI ChatGPT, and @Google Gemini, the results aren't just different. They are often the complete opposite. Claude usually feels more "corporate" to me. Gemini feels more natural for creative stuff like video scripts or blog posts. On multiple occasions I've put them side-by-side on technical questions where they just flat out disagree. ChatGPT tells me to use Approach A and specifically warns me against Approach B. Then I give Claude the same prompt and it picks Approach B, calling Approach A a bad idea. When you limit your models, you're limiting your own access to intelligence. You lose that second opinion. So how do you actually combine models? Here are a few ways. Side-by-side comparison. The simplest one - pass the same prompt to two different models separately, then combine the best parts from each into your own end result. Whether it is a blog post or code, you get to pick the winner for every sentence. Review chain. Take one model and make it produce something good for you. Then ask a second model to review it to find issues and suggest better options. Finally, let the original model or even a third one incorporate those suggestions. If you let multiple models work together, you are essentially replacing your slow review step with a much faster model. You might even get a much smarter reviewer too. That part is debatable, but the speed alone is worth it. Full council. Have one model write the code and three models review it. Or have multiple models each produce a draft and then synthesize the results. Autonomous orchestration. Even better yet, let one agent decide which models and subagents it needs for each subtask. Say you have a main agent based on Opus and you give it a task to build a website. It might spin up one subagent with @grok for image generation to create visual assets, another subagent on Codex to generate the frontend, and another one on Haiku for the backend. You are no longer even required to think about which models to choose or which models should work with or against each other. You just delegate that decision to an agent as well. This isn't for everyone yet. It's expensive. Paying for multiple subscriptions (or hey, here's an idea - get one subscription in Zenflow and use all those models) and token usage for agents talking back and forth adds up fast. But I really believe the future isn't about which model is the winner. It will be about which orchestration is the best. I feel lucky to be seeing this first-hand at @zencoderai . Watching a swarm of agents work together where one writes and three others review it feels totally different. You aren't just working with an AI anymore. It is a collective group working on whatever goal you set.
English
1
1
1
255
kyzo
kyzo@ky__zo·
everyone: nextjs 16.2.1 is so fast omg ◝(ᵔᵕᵔ)◜ nextjs: •`_´•
kyzo tweet media
English
28
4
230
32.6K