Osman R.

1.3K posts

Osman R. banner
Osman R.

Osman R.

@UsmanReads

I think I know, but I really don't. AI and Tech with 15 years in Industry. prev @groupon and @toptal

Universal Katılım Mart 2023
447 Takip Edilen380 Takipçiler
Osman R.
Osman R.@UsmanReads·
An AI agent can choose the right tool, pass valid arguments, and still make the wrong change. Tool schemas show models how to call an interface. They rarely expose the implementation guards, hidden state lookups, helper calls, side effects, or the authority required for that exact change. So a fluent tool call is still only a proposal. It is not permission, and it is not proof of what will happen. I have been testing a different execution contract for tool agents: Decision-Sufficient State. The model proposes an action. The tool implementation supplies its real semantics. The call runs first against private state. Independent policy decides whether the resulting transition is allowed. Only then can a one-shot proof commit that exact transition to live state. Across three independently authored agent tool codebases, the current system reproduced 564/564 tested transitions exactly. A frozen adversarial suite rejected 10/10 attempts to substitute arguments, replay proofs, use stale state, change source code, bypass authority, or corrupt the prepared result. No rejected attempt changed live state. This does not make the model smarter. It changes the boundary between what a model says and what a system permits it to do. Actions are not enough. More soon.
Osman R. tweet media
English
0
0
1
78
Osman R.
Osman R.@UsmanReads·
The world's worst user experience exists on @foodpanda_pk - They had shown delivery time of 15-20 minutes and it's been over an hour and there's no way to contact an agent. They need to up their product game. @FoodPanda
English
0
0
0
49
dax
dax@thdxr·
please i'm begging you show me something you built not another "this is my custom agent setup" post where you pretend you're doing something smarter than vanilla claude code please
English
2.2K
260
8.9K
674.3K
Osman R.
Osman R.@UsmanReads·
My model is 60% accurate in predicting who is going to win the game based on the World Cup matches so far. pitchprob.xyz - I made this for fun btw.
Osman R. tweet media
English
0
0
0
142
Osman R.
Osman R.@UsmanReads·
LLM wont teach you to think.
English
0
0
0
90
Osman R.
Osman R.@UsmanReads·
Long-horizon language agents should not rely primarily on replaying histories through context, retrieval, or summaries. They should convert observations into canonical typed operations, update compact predictive state, and reason about consequences from that state because after all consequences are all you will ever need.
English
0
0
0
79
Furqan Rydhan
Furqan Rydhan@FurqanR·
Spent a lot of time over the last few months deep in different parts of agent infra. Looking to try more platforms and tools, if you’re building something drop it below.
English
17
0
44
3.6K
Osman R.
Osman R.@UsmanReads·
Consequences are all you need.
English
0
0
0
23
Osman R.
Osman R.@UsmanReads·
Added Crowd Source feature to my fun website for Fifa World Cup. pitchprob.xyz
English
0
0
0
35
Thomas Trimoreau
Thomas Trimoreau@TTrimoreau·
Tell me one thing you can do that CLAUDE cannot do yet
English
261
1
127
18K
Osman R.
Osman R.@UsmanReads·
Exploring the nuances of AI-generated summaries, I investigated why they often seem flatter than the original writing. My hypothesis centered on whether AI, when shortening text, preserves the author's voice or merely conveys information. I analyzed 30 passages from various genres, including essays, memoirs, reportage, travel writing, and fiction. Each passage was summarized twice by three local AI models: first, with a standard summary request, and second, with an explicit instruction to maintain the author's voice. This resulted in 180 summaries. The findings were revealing. The Llama 3.1 model showed a notable improvement in retaining the author's perspective and quoted language when asked to preserve voice, although it did sacrifice some original content in the process. Conversely, Mistral did not exhibit a reliable change, indicating that this behavior is model-dependent. In essence, while asking AI to emulate the original writer can enhance the summary's personality, it may also lead to a loss of completeness. This suggests that effective summaries should strive to maintain both the content and the author's unique viewpoint. Current AI systems, however, often struggle to balance these two aspects, which could explain why original articles resonate more deeply than their AI-generated counterparts.
Osman R. tweet media
English
1
0
0
48
clem 🤗
clem 🤗@ClementDelangue·
Lots of people have known for a while that guardrails for frontier model APIs are very easily jailbroken, quite shallow and impossible to fix. They’re mostly a smokescreen and distraction, in my opinion. We need a different paradigm for AI safety!
English
54
26
386
66.3K
Osman R.
Osman R.@UsmanReads·
anybody help me here? I can't upgrade to Pro Codex from web because original subscription was brought from Android Phone. I can't upgrade to Pro Codex because I don't have that android phone and Currently using iPhone. Can someone help? @OpenAIDevs @OpenAI @thsottiaux
English
1
0
1
91
Osman R.
Osman R.@UsmanReads·
I am sorry Anthropic, but I had to request refund after Fable is gone. I will get back to Codex plan.
English
0
0
0
83
Mike Bradley
Mike Bradley@MikeBradleyAI·
Tried Fable 5 and am already back to GPT-5.5. 5.5 audits more thoroughly, is more careful with its work, and doesn’t randomly brick itself on tasks and become useless. Fable 5 is a month late to solving a solved problem, with unnecessarily added restrictions and token burn.
English
21
6
101
6.6K
Osman R.
Osman R.@UsmanReads·
Upcoming WC match. Paraguay has 36% chance to win whereas United State has 33%
Osman R. tweet media
English
0
0
0
37
Osman R.
Osman R.@UsmanReads·
So far, my match predictor correctly predicts Football matches, both 2/2.
English
0
0
0
25
Osman R.
Osman R.@UsmanReads·
#WorldCup2026 🇨🇦 Canada vs Bosnia-Herzegovina 🇧🇦 🕐 12 Jun, 19:00 UTC 🏆 Pick: Canada 70% (draw 21%) 📊 Likely: Canada 2–0 Bosnia-Herzegovina 🔥 Canada: 231 Elo higher → pitchprob.xyz
0
0
0
51
Osman R.
Osman R.@UsmanReads·
I put together a match winner predictor for Football Worldcup 26. I will keep posting in this thread for every match. For live match changing probabilities, you can stick to pitchprob.xyz
English
1
0
0
58