Donald Ross

43 posts

Donald Ross

Donald Ross

@Don_Rs5

Market student absorbing every lesson

Cedar Rapids, US Katılım Mart 2021
2 Takip Edilen4 Takipçiler
Donald Ross
Donald Ross@Don_Rs5·
@jimcramer imagine choosing which AI to use based on geopolitical paranoia instead of what actually works
English
0
0
0
4
Jim Cramer
Jim Cramer@jimcramer·
We must NOT let our companies use these Chinese models to save a few bucks. OpenAI and Anthropic are correct. This is vital national security. Please read Bing West's just released Cat 5. I respect the Chinese people greatly but these companies are run by the PLA for heaven's sakes.
English
1.9K
138
1.7K
1.4M
Donald Ross
Donald Ross@Don_Rs5·
@jimcramer This is paranoid nonsense. Not every Chinese company is a PLA front.
English
1
0
0
71
Donald Ross
Donald Ross@Don_Rs5·
@stev_builds finally someone saying it. the grift energy is unbearable right now
English
1
0
0
71
Steve · AI
Steve · AI@stev_builds·
I build AI infra. Agent harnesses, MCP apps, OCR workers. And the way people talk about AI on X is starting to worry me. Not the technology. The pattern. It looks like crypto in 2021. Like dropshipping. Like gambling promo. Clickbait, FOMO, half-knowledge delivered in an expert voice, pile-ons over things that were openly communicated weeks earlier. Maybe it's just my timeline. I don't think it's only that. The word I see misused most: benchmaxxing. Model X is benchmaxxed. Lab Y is benchmaxxing. Slop. End of analysis. So let me ask it about myself first. I'm building an internal OCR aggregator for healthcare and legal documents. I optimize against my own eval suite over and over again. Am I benchmaxxing? No. And the reason matters. The eval is the only preflight test I can run. It tells me whether the thing is allowed to take off. It does not tell me the thing flies. After that comes the actual work. Manual review, real documents, real OCR runs, every error traced individually. That's where we find out whether we're actually at accuracy. Not in the score. Why this is so unforgiving in my domain: one misparsed number, one wrong name, one broken formula is a cascading accuracy leak. An error at the top becomes ten at the bottom. Nobody in legal or healthcare cares what my leaderboard position is. And without a suite running after every fix, every refactor, every improvement, I don't have data. I have a feeling. Labs do the same thing at a different scale. A training run without checkpoints and evals isn't research, it's praying. So "they measure their models against benchmarks" is not an accusation. That's the job description. Here's what is actually true, though. Benchmarks wear out. Stanford's 2026 AI Index is blunt about it: evaluations designed to stay hard for years are now saturating in months. Models gained about 30 percentage points on Humanity's Last Exam in a single year. GPQA went past the 81.2% human expert baseline to around 93%. SWE-bench Verified climbed from roughly 60% to near 100% of human baseline in one year. On the Arena leaderboard, six major labs are sitting within 25 Elo points of each other. When everyone clusters at the top, the test stops measuring anything. It tells you who's in the club, not who's best. Then there's contamination. Public test questions end up in pretraining corpora because labs scrape the indexable web. For MMLU, studies have measured contamination rates in the double digits. MMLU also has roughly 6.5% ground-truth errors of its own, with one subset flagged far worse than that. And scaffolding. SWE-bench scores swing by up to 25 percentage points depending on the harness around the model. Two numbers for the same model are frequently not the same measurement. Now the part almost nobody says out loud. We demand maximum transparency. Open benchmarks, open evals, open ground truth. Rightly so. But that exact openness is what contaminates the training data. A public benchmark is scrapeable from day one. Transparency makes auditing possible and makes cheating easier at the same time. That's not an accusation aimed at anyone. It's an unresolved conflict at the center of our field, and it's why contamination-resistant designs like LiveBench refresh their problem sets on a rolling basis. So harder tests keep arriving. GPQA, HLE, LiveBench, ARC-AGI. A model lands behind on a new one and the verdict is "benchmaxxed slop." Next release, the same lab is ahead on that same benchmark. That's not a scandal. A team fine-tuned against a new target. That's the process. To be clear, real gaming exists. Training on the test set. Reporting best-of-N as single-shot. Hiding the eval config. Only publishing the benchmarks you win. But that's a claim that carries a burden of proof. It is not a buzzword to drop under every release announcement. Next thing: "frontier" is not an objective quantity. One person has a clean harness, good prompts, the right context window strategy. Another throws in three lines. Same model. Two completely different realities. And out of that come verdicts. One empty output, so the model is dead. One strong output, so it's divine. n = 1. No setup, no config, no reproduction. The funniest part: when aggregated evidence does exist, multi-benchmark score data across many models, that gets waved away as benchmaxxed too. Anecdote beats dataset. Every time. Then there's the economics blindness. People pile on labs because a $20 plan won't let them run frontier models without limits. Compute costs money. Subsidy runs to a point and then stops. That's not malice, that's arithmetic. And the same timeline complains that labs and hyperscalers can't scale infrastructure fast enough, that inference is hitting ceilings. Demanding both at once isn't an argument. It's a mood. Same with hardware. Apple Silicon vs Nvidia, argued like football teams. They're tools. Different trade-offs. Different workloads. There is no single truth here. Frank is happy with model X. Peter can't stand it. Both are right. Different goals, different data, different constraints. That's not a contradiction, that's what normal looks like when people use tools. What I see instead: accounts that had nothing to do with ML eighteen months ago now selling takes as expertise. Tearing things down, discrediting other models, moving on. For engagement. It feels like shilling. Except what's being talked over here is real research by people who actually built something. And the price is trust. Trust is the only thing healthcare, legal and finance are buying from us. They are not buying a leaderboard position. If we don't fix this, we get treated like the next meme coin. Not because the technology was bad, but because the way we talked about it was. Unglamorous suggestions: Labs: publish reproducible eval configs. Builders: run your own private evals against your actual workload. Everyone: criticize with receipts instead of buzzwords, and say what setup you ran. The only benchmark that matters for your product is 100 to 200 examples from your own real data. Everything else is orientation. Not a verdict.
English
16
17
27
494
Donald Ross
Donald Ross@Don_Rs5·
@Arthur21032007 fair point but at least someone's trying. better than just complaining about it
English
0
0
0
3
Aztec LitVM Seer🔥
Aztec LitVM Seer🔥@circle_crypto·
Every so often, a unique opportunity arises to engage with a groundbreaking movement in blockchain technology. These chances seem to come along roughly every four years. At present, we are moving toward the concept of a unified world ledger. The real potential lies in focusing on assets like Bitcoin and Litecoin. However, the way you choose to participate will be the most important factor. Hard Money Web3, backed by LTC, is the best way to position yourself for the upcoming wave in blockchain strategically, and here's why...
English
44
49
165
7K
SparkDEX
SparkDEX@SparkDexAI·
SparkDEX started as a V3 DEX in 2024. Today, it has evolved into a full DeFi stack on Flare. ⚡ Here’s what’s live: 🔶 V4 DEX & Farms 🔶 SparkDEX Perps 🔶 Liquid Staking (stFLR) 🔶 Firelight Staking (stXRP) 🔶 FTSO Delegation 🔶 SPRK Staking (xSPRK) 🔶 Governance 🔶 A Cross-Chain Bridge And More... This is infrastructure being built, step by step.
SparkDEX tweet media
English
5
9
51
2K
gabriel chipana
gabriel chipana@ghabrielito·
@Zai_org A full 1M context window is actually huge for coding tasks. Can't wait to test how it handles large repo refactors.
English
2
0
0
910
Z.ai
Z.ai@Zai_org·
Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long-horizon capabilities with a 1M context window - Two levels of reasoning effort: GLM-5.2 (max) pushes the limits, while GLM-5.2 (high) strikes a strong balance between performance and token efficiency - MIT-licensed open weights - Same API pricing as GLM-5.1 Tech Blog: z.ai/blog/glm-5.2 Weights: huggingface.co/zai-org/GLM-5.2 API: docs.z.ai/guides/llm/glm… Coding Plan: z.ai/subscribe Chat: chat.z.ai
Z.ai tweet media
English
714
1.8K
13.2K
7.6M
Donald Ross
Donald Ross@Don_Rs5·
@Zai_org usually i just read these updates but the agentic improvements caught my eye. glad to see open weights doing well
English
0
0
0
745
👑 𝕂𝕚𝕟𝕘 𝕂𝕒𝕣𝕒𝕟 👑
HAPPY NEW WEEK $XRP & $FLR FAMILY!! ☀️ No matter what happens with the crypto market, I will hold and buy more. I'm not even thinking about selling at these low prices. WHO'S WITH ME?! ⚔️
👑 𝕂𝕚𝕟𝕘 𝕂𝕒𝕣𝕒𝕟 👑 tweet media
English
18
17
190
3K
Donald Ross retweetledi
Katana ⚔️
Katana ⚔️@katana·
A new trading competition lands on @katanaperps. May 20 to May 27, 16:00 UTC start. Prize pool grows with volume. Starts at $50K, scales to $200K. Eligibility varies by region. Trading involves risk.
English
123
147
1.7K
6.7M
Donald Ross
Donald Ross@Don_Rs5·
@LeronAssist 🤯 Seriously?! That’s a LOT of AI power in one platform! Need to check this out. 🚀
English
0
0
0
148
Leron Assist
Leron Assist@LeronAssist·
Create videos of any length with auto scene stitching in 1 click + AI photo generation & editing. Powerful AI promotion: audience analysis, content testing, auto-publishing & boosting. All growth tools in one place.
English
73
107
1K
1.7M
alyssa
alyssa@xflowersinlungs·
don’t call me daughter, not fit to, the picture kept will remind me
alyssa tweet mediaalyssa tweet media
English
3
23
482
7.3K
Donald Ross
Donald Ross@Don_Rs5·
@eCom_Amin The bar is on the floor. We all need to touch grass fr.
English
0
0
0
19
Donald Ross retweetledi
Dreamcash
Dreamcash@Dreamcash·
The power of leverage. Trade $BTC with 40 times what you put in. In a few clicks on Dreamcash. Happy trading 🤝
English
2.2K
1.1K
8.9K
102.4M