
Alessandro Perilli
22.1K posts

Alessandro Perilli
@perilli
Vice President, Research @ IDC. I lead the AI Strategies research team. We focus on emerging AI tech, global AI market trends, and the state of enterprise AI.




You have 2 AI agents available on your computer/phone. They are identical, but: - Agent 1 has the voice of Paul Bettany as J.A.R.V.I.S. in the Iron Man movies. - Agent 2 has the voice of Scarlett Johansson as Samantha in Her. Do you use them for the same tasks and projects?



I am begging @AnthropicAI please add an automated test suite somewhere that ensures Claude Code on the web never blocks me from cloning or interacting with other public repos from within my existing sessions

I built an app that runs Claude Code, Codex, Cursor, Grok, Antigravity, OpenCode, and Pi side by side in my Mac notch. Agent finishes → notch pops. Needs approval → notch pops. One click back to that exact session. no more switching between apps.

You could spend $235 for the OpenAI programmable keyboard (with knob), or you could spend $18 on Amazon for this programmable keyboard (with knob)

Agentic AI “computer use” is not a gimmick. If you look a little further ahead, you can see it as a critical building block for voice UIs. First, if you haven’t tried the new instant, bidirectional live voice available in some AI systems, please try it now, or this post won’t make sense. The experience is nothing like voice dictation or previous, turn-based voice interactions. It takes a little time to get used to having an actual conversation with your AI, especially because you still try to respect the turns, you still try to remain as silent as possible between turns, and you still try to think hard about what to say before speaking. All that takes away from the spontaneousness of human communication. But none of that is necessary anymore. You can freely ramble in the same way you would with another person. And the more you try it, the more you realise that you are experiencing something truly big, truly transformative. Now, assume we’ll have that for every major AI agent. Second, assume that “computer use” will become ten times faster than it is today. The AI agent will open apps, click elements, and move windows at a speed far superior to that of a person. We have already seen significant progress in less than a year. Finally, connect the dots by assuming that the conversation you are having with your AI will lead to proactive behaviour: You are chatting with your agent and ask for the news of the day, the performance of a stock, the chart your CFO sent you, or an email whose exact wording you want to check. The agent answers by voice, yes, but also proactively launches the relevant apps and displays those digital artefacts for you. The screen populates, in fractions of a second, with all the information you need to see at that moment, then clears again as you move on to something else. If you can imagine that, you can see why progress in “computer use” is so important to watch. And if you are a software vendor willing to bet that we are exceptionally close to this reality, then you probably want to ask yourself: Can my software be controlled in this way, or does its current UI impede that flow?




Introducing GPT-Live, a new generation of voice models for natural human-AI interaction. Rolling out in ChatGPT starting today. You’ll want to turn the sound on for this one.



New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.








