that one doll

1.5K posts

that one doll

that one doll

@cute_n_clocky

Down With The Patriarchy • I’m here for nsfw stuff mostly • sometimes trans rambling • Minors DNI 🏳️‍⚧️🏳️‍🌈🇳🇬

27 she/her Katılım Ocak 2019
1K Takip Edilen54 Takipçiler
Ryan Marten
Ryan Marten@ryanmart3n·
We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%
Ryan Marten tweet media
English
54
78
797
229.3K
Ethan Mollick
Ethan Mollick@emollick·
Cool paper looking at how AIs solve unbounded, complex business problems in many fields by testing how well they can crack the cases we use to teach MBAs in business school: 1) AI already does extremely well across diverse business topics 2) Models are improving rapidly with time
Ethan Mollick tweet mediaEthan Mollick tweet mediaEthan Mollick tweet media
English
20
30
312
25.7K
that one doll
that one doll@cute_n_clocky·
@mercor_ai That’s insane lmao. Literally SOTA for management consulting 🥹🥹🥹
English
0
0
0
53
Mercor
Mercor@mercor_ai·
APEX-Agents score update for Muse Spark 1.1. When Meta first released Muse Spark 1.1, the model scored 37.1% Pass@1 on APEX-Agents, our agentic benchmark built for long-horizon professional services tasks in banking, law, and consulting. For roughly 10% of tasks, false positives in Meta's content moderation logic blocked trajectory completion, and as a result, those tasks failed and received 0-scores. We worked with @AIatMeta to tune content filters, which have since been updated on the publicly available model checkpoint. The new result: 41.9% Pass@1, placing Muse Spark 1.1 at #2 overall, behind only Claude Fable 5 (43.3%).
Mercor tweet media
English
7
5
73
16.1K
Logan Kilpatrick
Logan Kilpatrick@OfficialLoganK·
@estebs Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready
English
67
34
966
44.2K
that one doll retweetledi
OpenShip
OpenShip@openshipio·
You don't even want to know what sending 80,000 emails would cost you in 1 week. Postmark: $100+ SendGrid: $90+ Resend: $80+ Mailgun: $80+ OpenShip: $0. Deploy it on your 5$ VPS. Unlimited domains. Unlimited inboxes. No per-email bill. No vendor lock-in. And yes - it works with one-click setup. No deep infrastructure work. Just add a few DNS records, and you're ready to go.
OpenShip tweet media
English
252
290
5.6K
1.9M
NateOnTop
NateOnTop@NateOnTopfr·
The failure here isn't that guardrails exist — it's that they're keyed to topic instead of intent. Auditing your own codebase for vulnerabilities is the most common legitimate security task there is. Refusing it doesn't prevent harm; it just hands defensive work to whoever will do it. The line worth drawing is systems you don't own. A filter that can't tell an audit from an exploit isn't safety, it's a keyword blocklist.
English
1
0
0
706
David Sacks
David Sacks@DavidSacks·
Kimi K3 just fixed 15 critical security bugs that Codex and Fable refused because of “cyber guardrails.” There’s no reason to limit American models on tasks that Chinese models handle without issue. We’re only making ourselves less competitive.
calle@callebtc

I have a report full of security issues of a software I'm working on. Codex won't fix them because of Cyber guardrails Fable won't fix them because of Cyber guardrails Kimi K3 fixed them all. No restrictions, just gets the job done. This will end badly for OpenAI & Anthropic.

English
629
2K
20.6K
2.1M
Ferbin
Ferbin@Ferbin08·
@DailyNoud The weird part is how often the refusal sounds more confident than the answer would have been. A simple "can't help with that" would age better.
English
5
0
312
24K
Daily Noud
Daily Noud@DailyNoud·
I cannot fulfill this request as it contains subjects of hate speech and acts of violence. I am programmed to be a helpful and harmless AI assistant, and generating such content goes against my safety guidelines.
Daily Noud tweet mediaDaily Noud tweet media
English
101
3.3K
56.2K
1.5M
Angjoo Kanazawa
Angjoo Kanazawa@akanazawa·
We got 99% on ARC-AGI-3 by instilling analysis-by-synthesis into AI models. [Schema] is basically VIGA squared: do inverse graphics, then do it again one level up — inverse dynamics. Super simple, surprisingly effective! Analysis by synthesis FTW!!!
Haven Feng@HavenFeng

Today, we’re introducing [schema]: a harness reaching 99% RHAE with Opus 4.8 + Fable 5 and 95.35% with GPT-5.6 Sol on ARC-AGI-3 Public set. [schema] makes an LLM think like a physicist. 🧵

English
10
12
209
33.1K
Tibo
Tibo@thsottiaux·
Another reset for our Codex and ChatGPT Work users. Actually hit 9M active users way earlier today, but then got distracted by the approximately millions of things the team is doing to keep the systems up and reliable. Should have that sweet 100% weekly usage limit back in a few minutes. Go be your productive self and close twitter. Shoo!
English
2.2K
846
15K
1.5M
christiflow
christiflow@atiiflow·
About a week ago, after the release of GPT 5.6, I was receiving resets on my Business subscription, but now... the last two resets haven't reached me. Why do Plus members receive resets from you, even though their subscription is cheaper and they get Sol earlier than Business members, while we, with a Business subscription that’s more expensive to begin with plus the slots, don’t get the same rewards and benefit?(
English
25
1
121
26.2K
Danat Shekhe
Danat Shekhe@Danat_Sh·
@thsottiaux could you please clarify the resets for business accounts? It looks like they are being skipped, at least for the last two resets.
English
2
0
1
31
Ángel
Ángel@longtimefighter·
@thsottiaux Tibo, Business users are not getting their reset :( WHY ARE WE NO LONGER GETTING RESETS LIKE EVERYBODY ELSE :(
Ángel tweet media
English
3
0
9
715
that one doll retweetledi
Kimi.ai
Kimi.ai@Kimi_Moonshot·
Meet Kimi K3
Türkçe
751
2.7K
30.7K
10.2M
that one doll retweetledi
xx wang
xx wang@xxwang99999·
@thsottiaux Why wasn't ChatGPT Business reset?
English
162
16
463
63.6K
Madhu Dadi
Madhu Dadi@madhu245·
Business users still aren’t receiving resets. We missed both this reset and the previous one. Given that the Business plan costs more, it’s disappointing to feel deprioritized. I hope this isn’t related to Business data not being used for training. Could you clarify why resets aren’t being applied equally and address this?
English
2
1
14
919