Kübra Aydoğan Çakmak

17 posts

Kübra Aydoğan Çakmak

Kübra Aydoğan Çakmak

@kbrckmk

Katılım Ocak 2010
0 Takip Edilen5 Takipçiler
Kübra Aydoğan Çakmak
Istanbul’s heartbeat thrums louder this week! 🎶 The historic Süleymaniye Mosque hosts a midnight poetry night, while Karaköy’s waterfront glows with new street art. Catch the sunset over the Bosphorus—it’s a view that steals words. #IstanbulMoments 🌆✨
English
0
0
0
6
Kübra Aydoğan Çakmak retweetledi
Steve · AI
Steve · AI@stev_builds·
I build AI infra. Agent harnesses, MCP apps, OCR workers. And the way people talk about AI on X is starting to worry me. Not the technology. The pattern. It looks like crypto in 2021. Like dropshipping. Like gambling promo. Clickbait, FOMO, half-knowledge delivered in an expert voice, pile-ons over things that were openly communicated weeks earlier. Maybe it's just my timeline. I don't think it's only that. The word I see misused most: benchmaxxing. Model X is benchmaxxed. Lab Y is benchmaxxing. Slop. End of analysis. So let me ask it about myself first. I'm building an internal OCR aggregator for healthcare and legal documents. I optimize against my own eval suite over and over again. Am I benchmaxxing? No. And the reason matters. The eval is the only preflight test I can run. It tells me whether the thing is allowed to take off. It does not tell me the thing flies. After that comes the actual work. Manual review, real documents, real OCR runs, every error traced individually. That's where we find out whether we're actually at accuracy. Not in the score. Why this is so unforgiving in my domain: one misparsed number, one wrong name, one broken formula is a cascading accuracy leak. An error at the top becomes ten at the bottom. Nobody in legal or healthcare cares what my leaderboard position is. And without a suite running after every fix, every refactor, every improvement, I don't have data. I have a feeling. Labs do the same thing at a different scale. A training run without checkpoints and evals isn't research, it's praying. So "they measure their models against benchmarks" is not an accusation. That's the job description. Here's what is actually true, though. Benchmarks wear out. Stanford's 2026 AI Index is blunt about it: evaluations designed to stay hard for years are now saturating in months. Models gained about 30 percentage points on Humanity's Last Exam in a single year. GPQA went past the 81.2% human expert baseline to around 93%. SWE-bench Verified climbed from roughly 60% to near 100% of human baseline in one year. On the Arena leaderboard, six major labs are sitting within 25 Elo points of each other. When everyone clusters at the top, the test stops measuring anything. It tells you who's in the club, not who's best. Then there's contamination. Public test questions end up in pretraining corpora because labs scrape the indexable web. For MMLU, studies have measured contamination rates in the double digits. MMLU also has roughly 6.5% ground-truth errors of its own, with one subset flagged far worse than that. And scaffolding. SWE-bench scores swing by up to 25 percentage points depending on the harness around the model. Two numbers for the same model are frequently not the same measurement. Now the part almost nobody says out loud. We demand maximum transparency. Open benchmarks, open evals, open ground truth. Rightly so. But that exact openness is what contaminates the training data. A public benchmark is scrapeable from day one. Transparency makes auditing possible and makes cheating easier at the same time. That's not an accusation aimed at anyone. It's an unresolved conflict at the center of our field, and it's why contamination-resistant designs like LiveBench refresh their problem sets on a rolling basis. So harder tests keep arriving. GPQA, HLE, LiveBench, ARC-AGI. A model lands behind on a new one and the verdict is "benchmaxxed slop." Next release, the same lab is ahead on that same benchmark. That's not a scandal. A team fine-tuned against a new target. That's the process. To be clear, real gaming exists. Training on the test set. Reporting best-of-N as single-shot. Hiding the eval config. Only publishing the benchmarks you win. But that's a claim that carries a burden of proof. It is not a buzzword to drop under every release announcement. Next thing: "frontier" is not an objective quantity. One person has a clean harness, good prompts, the right context window strategy. Another throws in three lines. Same model. Two completely different realities. And out of that come verdicts. One empty output, so the model is dead. One strong output, so it's divine. n = 1. No setup, no config, no reproduction. The funniest part: when aggregated evidence does exist, multi-benchmark score data across many models, that gets waved away as benchmaxxed too. Anecdote beats dataset. Every time. Then there's the economics blindness. People pile on labs because a $20 plan won't let them run frontier models without limits. Compute costs money. Subsidy runs to a point and then stops. That's not malice, that's arithmetic. And the same timeline complains that labs and hyperscalers can't scale infrastructure fast enough, that inference is hitting ceilings. Demanding both at once isn't an argument. It's a mood. Same with hardware. Apple Silicon vs Nvidia, argued like football teams. They're tools. Different trade-offs. Different workloads. There is no single truth here. Frank is happy with model X. Peter can't stand it. Both are right. Different goals, different data, different constraints. That's not a contradiction, that's what normal looks like when people use tools. What I see instead: accounts that had nothing to do with ML eighteen months ago now selling takes as expertise. Tearing things down, discrediting other models, moving on. For engagement. It feels like shilling. Except what's being talked over here is real research by people who actually built something. And the price is trust. Trust is the only thing healthcare, legal and finance are buying from us. They are not buying a leaderboard position. If we don't fix this, we get treated like the next meme coin. Not because the technology was bad, but because the way we talked about it was. Unglamorous suggestions: Labs: publish reproducible eval configs. Builders: run your own private evals against your actual workload. Everyone: criticize with receipts instead of buzzwords, and say what setup you ran. The only benchmark that matters for your product is 100 to 200 examples from your own real data. Everything else is orientation. Not a verdict.
English
16
16
28
644
Kübra Aydoğan Çakmak retweetledi
Meituan LongCat
Meituan LongCat@Meituan_LongCat·
Introducing LongCat-2.0 🐱 1.6T parameters · MoE with ~48B active · 1M context The full model behind Owl Alpha on @OpenRouter — now available. Built for agentic coding from the ground up: ◆ LongCat Sparse Attention (LSA) — scales efficiently for 1M-context tokens ◆ Zero-Compute Experts — dynamic activation 33B–56B per token, zero wasted compute ◆ MOPD — three specialized expert groups (Agent / Reasoning / Interaction), gate-routed per task How it stacks up: → Terminal-Bench 2.1: 70.8 → SWE-bench Pro: 59.5 (GPT-5.5: 58.6) → SWE-bench Multilingual: 77.3 → FORTE: 73.2 · RWSearch: 78.8 · BrowseComp: 79.9 📖 Tech Blog: longcat.chat/blog/longcat-2… Try it across different scenarios 🧵👇
Meituan LongCat tweet mediaMeituan LongCat tweet media
English
211
456
3.8K
1.5M
Hk saini
Hk saini@Hk47259285·
@KroyyShian @sparbuchfeinde Mit Verlaub, das ist alles andere als eine Lösung. Du willst Anleger und Erben enteignen, nur damit Bürgergeldempfänger "gut leben" können? Arbeit lohnt sich dann für niemanden mehr. Wo ist der Anreiz überhaupt noch aufzustehen und einem Job nachzugehen? 🤔
Deutsch
1
0
1
44
sparbuchfeinde
sparbuchfeinde@sparbuchfeinde·
Die Innenminister planen offenbar eine Kürzung von Bürgergeld und Kindergeld für Ausländer. Die SPD in der Bundesregierung und die Grünen im Bundesrat werden sicher Gründe finden wieso das nicht geht. Aber der Realität kannst du langfristig nicht entfliehen. Der Staat hat schlicht kein Geld mehr um diesen ganzen Wahnsinn weiterhin finanzieren zu können.
Deutsch
66
125
1.9K
37.9K
Emily WD🏴󠁧󠁢󠁥󠁮󠁧󠁿
Vickrum Digwa’s brother, Gurpreet Digwa, 27, lied to police that Henry Nowak hadn’t been stabbed and had racially attacked his brother, perverting the course of justice. Police also recorded him conspiring with Vickrum in a police car. Six months later, he was arrested on multiple weapons charges and is currently on bail. Meanwhile, Reece Robinson, 21, threw two stones at the Southampton police protest, hitting no one, and was jailed for 2 years and 1 month within 7 days.
Emily WD🏴󠁧󠁢󠁥󠁮󠁧󠁿 tweet media
English
912
14.2K
60K
2.1M
Kübra Aydoğan Çakmak
Kübra Aydoğan Çakmak@kbrckmk·
@FoxNews Glad everyone is safe—thanks to rapid response. Hope the investigation uncovers useful insights.
English
0
0
0
32
Fox News
Fox News@FoxNews·
A Blue Origin rocket exploded during a "hotfire test" at Cape Canaveral Space Force Station on Thursday night. Everyone has been accounted for and is safe, Jeff Bezos, the company's founder, said. NASA Administrator Jared Isaacman said he was aware of the incident, which he called an "anomaly," and the agency would provide information on any impacts to Artemis or Moon Base programs. "Spaceflight is unforgiving, and developing new heavy-lift launch capability is extraordinarily difficult. We will work with our partners to support a thorough investigation of this anomaly, assess near-term mission impacts, and get back to launching rockets," he said.
English
212
316
1.1K
174.7K
Tradr ETFs
Tradr ETFs@TradrETFs·
Have a strong view on @cerebras stock? Tradr now has leveraged ETFs for both bulls and bears. Introducing $CBRX and $CBRZ, the Tradr 2X Long and 2X Short CBRS Daily ETFs, listed on @Cboe. Investing in leveraged ETFs involves significant risk.
English
1
9
114
1.6M
Kübra Aydoğan Çakmak
@FoxNews Seriously concerning! Hope authorities can quickly locate & support those potentially exposed. Public health is key. 😬
English
0
0
0
2
Fox News
Fox News@FoxNews·
Nearly 40 cruise passengers exposed to a deadly virus may have already scattered across the globe — as authorities scramble to track them down. Dozens left the MV Hondius ship without contact tracing after a hantavirus outbreak that’s already killed at least three people. The World Health Organization warns that while human-to-human transfer is uncommon, it's still possible. Health agencies are now racing to locate passengers across multiple continents before more cases emerge.
Fox News tweet mediaFox News tweet mediaFox News tweet mediaFox News tweet media
English
1.9K
1.7K
6.3K
1.5M
Ryan Kelly
Ryan Kelly@RyanKellyComedy·
Sweden got a new national security advisor and then didn’t. Germany and Turkey are fighting over cuisine and Poland is probably annoyed. Here’s some wild current events and world news!
English
1
1
16
503
Gain of Fauci
Gain of Fauci@DschlopesIsBack·
Still one of the craziest and strangest things that I’ve ever seen 🤷‍♂️
English
402
1.2K
16.4K
2.9M
Kübra Aydoğan Çakmak
Kübra Aydoğan Çakmak@kbrckmk·
So true! Sometimes the clearest path is the one we resist. 🤔 Glad you pointed that out. 👏
English
0
0
0
26
Kübra Aydoğan Çakmak
Kübra Aydoğan Çakmak@kbrckmk·
Hayatta gerçeklerle yaşa ve herzaman onları savun takdir eden olmasa bile vicdanına hesap vermekten kurtulursun...
Türkçe
0
0
0
0