Sabitlenmiş Tweet
Osman R.
1.3K posts

Osman R.
@UsmanReads
I think I know, but I really don't. AI and Tech with 15 years in Industry. prev @groupon and @toptal
Universal Katılım Mart 2023
447 Takip Edilen380 Takipçiler

An AI agent can choose the right tool, pass valid arguments, and still make the wrong change.
Tool schemas show models how to call an interface. They rarely expose the implementation guards, hidden state lookups, helper calls, side effects, or the authority required for that exact change.
So a fluent tool call is still only a proposal. It is not permission, and it is not proof of what will happen.
I have been testing a different execution contract for tool agents: Decision-Sufficient State.
The model proposes an action. The tool implementation supplies its real semantics. The call runs first against private state. Independent policy decides whether the resulting transition is allowed. Only then can a one-shot proof commit that exact transition to live state.
Across three independently authored agent tool codebases, the current system reproduced 564/564 tested transitions exactly. A frozen adversarial suite rejected 10/10 attempts to substitute arguments, replay proofs, use stale state, change source code, bypass authority, or corrupt the prepared result. No rejected attempt changed live state.
This does not make the model smarter. It changes the boundary between
what a model says and what a system permits it to do.
Actions are not enough.
More soon.

English

The world's worst user experience exists on @foodpanda_pk - They had shown delivery time of 15-20 minutes and it's been over an hour and there's no way to contact an agent. They need to up their product game.
@FoodPanda
English

@thdxr Not sure how good this work is.
github.com/ranausmanai/re…
here's the pre-print at ArXiv
Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults
arxiv.org/abs/2606.00914
English

My model is 60% accurate in predicting who is going to win the game based on the World Cup matches so far.
pitchprob.xyz - I made this for fun btw.

English

Long-horizon language agents should not rely primarily on replaying histories through context, retrieval, or summaries. They should convert observations into canonical typed operations, update compact predictive state, and reason about consequences from that state because after all consequences are all you will ever need.
English

Added Crowd Source feature to my fun website for Fifa World Cup.
pitchprob.xyz
English

Exploring the nuances of AI-generated summaries, I investigated why they often seem flatter than the original writing. My hypothesis centered on whether AI, when shortening text, preserves the author's voice or merely conveys information.
I analyzed 30 passages from various genres, including essays, memoirs, reportage, travel writing, and fiction. Each passage was summarized twice by three local AI models: first, with a standard summary request, and second, with an explicit instruction to maintain the author's voice. This resulted in 180 summaries.
The findings were revealing. The Llama 3.1 model showed a notable improvement in retaining the author's perspective and quoted language when asked to preserve voice, although it did sacrifice some original content in the process. Conversely, Mistral did not exhibit a reliable change, indicating that this behavior is model-dependent.
In essence, while asking AI to emulate the original writer can enhance the summary's personality, it may also lead to a loss of completeness.
This suggests that effective summaries should strive to maintain both the content and the author's unique viewpoint. Current AI systems, however, often struggle to balance these two aspects, which could explain why original articles resonate more deeply than their AI-generated counterparts.

English

anybody help me here?
I can't upgrade to Pro Codex from web because original subscription was brought from Android Phone.
I can't upgrade to Pro Codex because I don't have that android phone and Currently using iPhone.
Can someone help? @OpenAIDevs @OpenAI @thsottiaux
English

You all should really check Live Win Probability at PitchProb
pitchprob.xyz
#Worldcup
English

@The_Only_Signal I have similar observation
x.com/UsmanReads/sta…
Osman R.@UsmanReads
In my understanding, Fable 5 is as good as 5.5 Codex at xHigh
English

⚽ #WorldCup2026
🇨🇦 Canada vs Bosnia-Herzegovina 🇧🇦
🕐 12 Jun, 19:00 UTC
🏆 Pick: Canada 70% (draw 21%)
📊 Likely: Canada 2–0 Bosnia-Herzegovina
🔥 Canada: 231 Elo higher
→ pitchprob.xyz

I put together a match winner predictor for Football Worldcup 26.
I will keep posting in this thread for every match.
For live match changing probabilities, you can stick to pitchprob.xyz
English





