
Benۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗ☁️
1.7K posts

Benۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗ☁️
@benswerd
A computer enhanced hallucination ۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗ | Vibing @freestyle_dev (YC S24)




hot off the presses - you wanna get agents to code good, but BENCHMARKS ARE NOT ENOUGH - me and @vaibcode went deep on 1) how benchmarks like SWE-bench, terminal bench work 2) What they miss, and why they can't penalize agents for slop code 3) new benchmarks from @cognition, @datacurve, and Abundant that move the needle on harder software problems (and what they still miss) 4) how to adjust your intuition on coding agents accordingly and make them work for YOUR workflow on YOUR team all in a day's🦄 AI that works - enjoy











@benswerd i added it to my website october.dev! it’s so cool!












