that one doll
1.5K posts

that one doll
@cute_n_clocky
Down With The Patriarchy • I’m here for nsfw stuff mostly • sometimes trans rambling • Minors DNI 🏳️⚧️🏳️🌈🇳🇬
27 she/her Katılım Ocak 2019
1K Takip Edilen54 Takipçiler

you gotta love a caricature where they just add shit that wasn’t there
Aria Faye@Ariafayeee
Someone made this Wojack of me pt.2 and I think it’s so funny I can’t even lie 😭🏳️⚧️
English

@mercor_ai That’s insane lmao. Literally SOTA for management consulting 🥹🥹🥹
English

APEX-Agents score update for Muse Spark 1.1.
When Meta first released Muse Spark 1.1, the model scored 37.1% Pass@1 on APEX-Agents, our agentic benchmark built for long-horizon professional services tasks in banking, law, and consulting. For roughly 10% of tasks, false positives in Meta's content moderation logic blocked trajectory completion, and as a result, those tasks failed and received 0-scores.
We worked with @AIatMeta to tune content filters, which have since been updated on the publicly available model checkpoint. The new result: 41.9% Pass@1, placing Muse Spark 1.1 at #2 overall, behind only Claude Fable 5 (43.3%).

English

@estebs Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready
English

that one doll retweetledi

You don't even want to know what sending 80,000 emails would cost you in 1 week.
Postmark: $100+
SendGrid: $90+
Resend: $80+
Mailgun: $80+
OpenShip: $0.
Deploy it on your 5$ VPS.
Unlimited domains.
Unlimited inboxes.
No per-email bill.
No vendor lock-in.
And yes - it works with one-click setup.
No deep infrastructure work.
Just add a few DNS records, and you're ready to go.

English

The failure here isn't that guardrails exist — it's that they're keyed to topic instead of intent.
Auditing your own codebase for vulnerabilities is the most common legitimate security task there is. Refusing it doesn't prevent harm; it just hands defensive work to whoever will do it.
The line worth drawing is systems you don't own. A filter that can't tell an audit from an exploit isn't safety, it's a keyword blocklist.
English

Kimi K3 just fixed 15 critical security bugs that Codex and Fable refused because of “cyber guardrails.” There’s no reason to limit American models on tasks that Chinese models handle without issue. We’re only making ourselves less competitive.
calle@callebtc
I have a report full of security issues of a software I'm working on. Codex won't fix them because of Cyber guardrails Fable won't fix them because of Cyber guardrails Kimi K3 fixed them all. No restrictions, just gets the job done. This will end badly for OpenAI & Anthropic.
English

@DailyNoud The weird part is how often the refusal sounds more confident than the answer would have been.
A simple "can't help with that" would age better.
English

We got 99% on ARC-AGI-3 by instilling analysis-by-synthesis into AI models.
[Schema] is basically VIGA squared: do inverse graphics, then do it again one level up — inverse dynamics.
Super simple, surprisingly effective! Analysis by synthesis FTW!!!
Haven Feng@HavenFeng
Today, we’re introducing [schema]: a harness reaching 99% RHAE with Opus 4.8 + Fable 5 and 95.35% with GPT-5.6 Sol on ARC-AGI-3 Public set. [schema] makes an LLM think like a physicist. 🧵
English

Another reset for our Codex and ChatGPT Work users. Actually hit 9M active users way earlier today, but then got distracted by the approximately millions of things the team is doing to keep the systems up and reliable.
Should have that sweet 100% weekly usage limit back in a few minutes. Go be your productive self and close twitter. Shoo!
English

About a week ago, after the release of GPT 5.6, I was receiving resets on my Business subscription, but now... the last two resets haven't reached me.
Why do Plus members receive resets from you, even though their subscription is cheaper and they get Sol earlier than Business members, while we, with a Business subscription that’s more expensive to begin with plus the slots, don’t get the same rewards and benefit?(
English

@thsottiaux could you please clarify the resets for business accounts? It looks like they are being skipped, at least for the last two resets.
English

@longtimefighter @thsottiaux it sucks. ive switched to claude this week bc 0%
English

@thsottiaux Tibo, Business users are not getting their reset :(
WHY ARE WE NO LONGER GETTING RESETS LIKE EVERYBODY ELSE :(

English
that one doll retweetledi
that one doll retweetledi

@BlakeJOwens @thsottiaux Business plans haven’t been getting them :(
English

@thsottiaux The last two resets haven't been reflected in my account.
Any reason why?
English

@madhu245 @thsottiaux This sucks. Seeing some get hyped and business gets nothing. Feels so bad.
English

Business users still aren’t receiving resets. We missed both this reset and the previous one. Given that the Business plan costs more, it’s disappointing to feel deprioritized.
I hope this isn’t related to Business data not being used for training. Could you clarify why resets aren’t being applied equally and address this?
English













