Dariush Wahdany

491 posts

Dariush Wahdany banner
Dariush Wahdany

Dariush Wahdany

@dwahdany

member of technically not staff

Berlin, Deutschland Katılım Mayıs 2022
188 Takip Edilen82 Takipçiler
Dariush Wahdany
Dariush Wahdany@dwahdany·
@BlancheMinerva @roydanroy Its actually only showing up under certain conditions when you click the link. Maybe for him it looks fine. I don’t think his x necessarily got hacked.
Dan Roy@roydanroy

@dwahdany Where does this link appear? Doesn't appear for me. This website came up when I googled for rick roll.

English
0
0
0
1.4K
Stella Biderman
Stella Biderman@BlancheMinerva·
If anyone who follows me knows Dan Roy @roydanroy, please let him know that his account has been compromised and someone is spreading a fake position that actually downloads malware through his X account.
Stella Biderman tweet media
English
10
3
70
14.9K
Dariush Wahdany
Dariush Wahdany@dwahdany·
@roydanroy when i click on what you shared it goes straight to phishing for me, i am not clicking on anything on the site 🤷
English
0
0
0
582
Dan Roy
Dan Roy@roydanroy·
@dwahdany Where does this link appear? Doesn't appear for me. This website came up when I googled for rick roll.
English
1
0
2
3.4K
Gergely Orosz
Gergely Orosz@GergelyOrosz·
A massive change for me at least: that “coding flow” being uninterrupted is harder to get because “in the flow” is when I do one thing only (aka sequential work). Working with agents tempts parallel work that can be more productive but not “flow” It’s weird
Wes Bos@wesbos

Do you ever hit flow state with AI coding?

English
64
20
567
50.6K
merlin
merlin@merlindru·
weird claude opus 5 failure mode this exact text gives it problems even without memory on (and in incognito chats)
merlin tweet media
English
68
13
413
288.9K
Dariush Wahdany
Dariush Wahdany@dwahdany·
i guess it's the new normal to see traces of agent hacking when just browsing the internet (no it's not huggingface)
Dariush Wahdany tweet media
English
0
0
0
30
Dariush Wahdany
Dariush Wahdany@dwahdany·
@dillon_mulroy opus 5 is so ahead of the other models i had to use /bro *twice* on the same response! never had that before
English
0
0
0
11
Dillon Mulroy
Dillon Mulroy@dillon_mulroy·
/bro remains undefeated
Dillon Mulroy tweet media
English
39
19
1K
123.3K
Dariush Wahdany
Dariush Wahdany@dwahdany·
@giansegato wow! makes me wonder if there is some value in constraining modalities or even tools in evals, a la "constraints breed creativity"
English
0
0
0
120
Dariush Wahdany retweetledi
gian
gian@giansegato·
the most goated example i could find when eval'ing Opus 5 was on a frontier bench very hard task. the model is given an engineering drawing of a machine part and asked to write code to rebuild it as a 3D freeCAD model. the fun catch is that the model has no way to view the drawing: mini-swe-agent is a text-only harness, so all the model ever sees is terminal text over bash IO. it's something hard and multidisciplinary even for a human engineer capable of seeing the drawing! working blind, Opus 5 first tried off-the-shelf OCR with Pillow and Tesseract, found the results unreliable on the drawing’s thin numerals, and built its own vision stack on the spot. carved out each printed character as a connected component, rendered the glyphs as ASCII art (?!?), and matched glyph bitmaps to read the dimension labels off the page then it vectorized the drawing itself: run-length line detection with sub-pixel centroids, least-squares circle and ellipse fits to recover radii and bores, arrowhead detection to tell dimension lines from part geometry. it worked out that the four views were printed at different scales, derived each one’s mm-per-pixel ratio, and rebuilt the part, feature by feature, as a parametric freeCAD model welcome to the world, my dude!
Claude@claudeai

Opus 5 is also highly efficient. It outperforms other models for a similar or lower cost per task:

English
12
13
152
20.7K
Dariush Wahdany
Dariush Wahdany@dwahdany·
@NicolaiSchmid do you know which search engine they use? if you want to show up on claude they use brave, so you need to optimize there
Dariush Wahdany tweet media
English
1
0
1
55
Dariush Wahdany
Dariush Wahdany@dwahdany·
@modal You literally just select the model and click create. It's magical. And with just an additional click, they automatically benchmark your endpoint for either agentic or interactive usage.
Dariush Wahdany tweet mediaDariush Wahdany tweet media
English
0
0
0
16
Dariush Wahdany
Dariush Wahdany@dwahdany·
@modal endpoints are really good. You get a well-optimized llm endpoint with just one click, using all their inference knowledge (modal.com/llm-almanac). GLM 5.2 FP8 (a 750GB model!) running on 8xB200 pushing 600 output tok/s like it's nothing 🤷‍♂️
English
1
0
0
50
Dariush Wahdany
Dariush Wahdany@dwahdany·
Frontier-Bench v0.1 (or in spirit, Terminal-Bench v3.0rc1) is out! Exciting new changes: - stronger differentiation between frontier models - regular updates: treat benchmarks like software - auditing process of all benchmark submissions - refined task review process
Ryan Marten@ryan_marten

We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%

English
1
1
7
1.1K
Dariush Wahdany retweetledi
gabriel
gabriel@g_br_l·
gabriel tweet media
ZXX
17
184
3.7K
56.4K
Dariush Wahdany retweetledi
spike
spike@spikedoanz·
spike tweet media
ZXX
9
129
3.5K
63.7K