🕸️

1.9K posts

🕸️

🕸️

@psycvcs

compsci student

Katılım Nisan 2020
123 Takip Edilen10 Takipçiler
🕸️
🕸️@psycvcs·
@thsottiaux i hate 5 hours limit so much. please keep 1 week usage limit
English
0
0
0
9
Tibo
Tibo@thsottiaux·
Hello people of Sol! I've reset usage limits for all ChatGPT Work and Codex users. Together with that, a quick update on GPT-5.6 Sol usage limits. Over the past few weeks, many of you have told us that Sol was using your Codex limits faster than expected. To be clear, we have not reduced usage on any subscription plans. We’ve been digging into what was happening and have landed several improvements. As a result, we expect your usage to last around 18% longer during typical use of Sol. Some of you should already see significantly larger improvements from today. Tomorrow, we’ll also restore the five-hour limit that we temporarily paused while investigating. Here’s what we found: - GPT-5.6 Sol is much more willing to work for longer, make additional tool calls, and coordinate complex workflows across tools and subagents. That makes it better at solving hard problems, but some tasks were using far more than we intended. - Sol also works harder at the same reasoning effort than previous models. High on Sol can use more tokens than High did on GPT-5.5. - Programmatic tool calling, also referred to as code mode, gives Sol much more flexibility to run tool calls in parallel or continue working while waiting. But it also led to more responses per turn, more cached input tokens, and higher usage than expected. - This was particularly noticeable when Sol was waiting for tool calls to finish or running many web searches. We’ve improved how we handle both cases and are continuing to make code mode more efficient. - The impact was also very uneven. The median user actually found Sol quite token efficient, while some power users working on harder tasks saw their usage drain much faster. We were very focused on average and median usage before launch and missed some cases where the long tail could use significantly more usage. Sol is a significant step forward in what Codex can do, but capability and efficiency do not always improve at the same pace, and some issues only become clear once people are using the model at real-world scale. We should have recognized this sooner and been more upfront about it. You keep pushing the frontier and we’ll keep improving efficiency and sharing updates as we go.
English
3.4K
1.2K
21.1K
2.9M
🕸️
🕸️@psycvcs·
oracle always free kenapa out of capacity mulu dah
English
0
0
0
10
🕸️
🕸️@psycvcs·
opus 5 nih jelek banget ya. outputnya suka ga konsisten kalau dikirim prompt berulang dan kemampuan yang dikeluarin ga sama di tiap prompt. outputnya berasa di fallback ke opus 4.8
Indonesia
0
0
0
17
🕸️
🕸️@psycvcs·
claude opus 5 punya knowledge yang lebih luas dibanding gpt sol di AI research. tapi gpt sol punya kemampuan koreksi yang lebih bagus
Indonesia
1
0
0
33
Tibo
Tibo@thsottiaux·
@theo Poor guy cleans up his bio and goes to sleep. Wakes up with the world in shambles
English
33
6
808
39.4K
Logan Kilpatrick
Logan Kilpatrick@OfficialLoganK·
Automating AI research is going to look a lot more like data cleaning than it is going to look like inventing the transformer
English
206
166
3.7K
283.8K
🕸️ retweetledi
Happiest Person
Happiest Person@Mrs_Person1·
@bisulpecahhhh Setuju banget. . Dan kabar baiknya gugatan di MK infonya sudah terkabul yaa. Semoga bukan omon² aja,😁 S: detikNews
Happiest Person tweet media
Indonesia
26
159
7.8K
240.9K
🕸️ retweetledi
Chhavi Yadav
Chhavi Yadav@chhaviyadav_·
we have certainly moved on from 8B models.. reviewers are asking for evals on bigger & more advanced models in the open-source space, eg. someone asked us for evals on GPT-OSS (21B MoE), which also needs to be fine-tuned ofc.. what open models are you using in your papers? @NeurIPSConf
English
14
4
113
42.8K
bluedev
bluedev@blueemi99·
holy shit, t3 code looks so insane right now great job t3 team!
bluedev tweet media
English
26
3
615
145.5K
🕸️
🕸️@psycvcs·
@eliebakouch is claude still reliable for AI research? isn't sol better?
English
0
0
0
6
🕸️
🕸️@psycvcs·
@levifikri ga juga. ada beberapa pertanyaan AI masih tolol. mau itu chatgpt, claude, grok, gemini, kimi, qwen, meta ga bisa jawab sama sekali. algoritma google search masih yg paling bagus
Indonesia
0
0
0
24
Levi | still learning
Levi | still learning@levifikri·
Agentic coding sudah bagus dan lebih produktif, sehingga buka VSCode dan inspecting file satu2 terlihat ancient relic sekarang Voice input sudah akurat dan lebih cepat, sehingga ngetik pake keyboard karakter per karakter terlihat ancient relic sekarang Asking AI sudah lebih cepat dapat jawaban, sehingga buka Google dan klik link satu per satu terlihat ancient relic sekarang Padahal itu semua activities yang normal kita lakukan beberapa tahun lalu. Things move that fast Apa lagi ya?
Indonesia
9
11
171
11.7K
🕸️ retweetledi
nico
nico@nicochristie·
Opus 5 is a shocking step up from Fable 5 and 5.6 Sol for spreadsheets Almost couldn’t believe our evals at first A step function increase in intelligence, while as efficient as Sol Feels like another Opus 4.5 moment for knowledge work
nico tweet media
English
25
67
1.4K
109.9K
Beto
Beto@betomoedano·
Opus 5 is now the cheapest frontier model
Beto tweet media
English
118
26
1.5K
357.3K
🕸️ retweetledi
Boris Cherny
Boris Cherny@bcherny·
Opus 5 is a great model for coding, data analysis, design, biology, knowledge work. More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully. And when layering defenses -- strong model alignment, combined with prompt injection probes, combined with Auto Mode in Claude Code -- the success rate for prompt injection attacks drops to ~0. This is new and exciting! More about this soon. #page=73" target="_blank" rel="nofollow noopener">www-cdn.anthropic.com/c5fbac3f0b1280…
Boris Cherny tweet media
Claude@claudeai

On several coding and knowledge work evaluations, Opus 5 is the new state-of-the-art:

English
345
475
6.2K
654.4K
🕸️ retweetledi
General Reasoning
General Reasoning@GenReasoning·
🔍 Introducing BackSearch. LLMs are increasingly asked to predict the future, but a good backtest requires a snapshot of the internet at a point in time. BackSearch allows LLMs to search the web as it was on a particular date. It’s great for: 🔮 Forecasting and prediction markets. 📈 Quantitative finance. 🌍 RL environments that simulate the world. 🥶 Freezing websearch for benchmark reproducibility We’re releasing a narrow slice of our index to begin with, focused on the news domain for 2026. Based on feedback we’ll open up more of our index in subsequent releases. 👇 Try BackSearch out with the link below.
English
73
140
2.1K
340K
🕸️ retweetledi
Artificial Analysis
Artificial Analysis@ArtificialAnlys·
Claude Opus 5 is narrowly the most intelligent model on the Artificial Analysis Intelligence Index, offering comparable intelligence to Fable 5 at 26% lower Cost per Task We supported @AnthropicAI to evaluate Claude Opus 5 ahead of release: it sets the highest GDPval-AA v2 and AA-Briefcase scores so far. Opus 5 (max) scores 61 on the Artificial Analysis Intelligence Index, effectively tied with Claude Fable 5 (max, 60), and ahead of GPT-5.6 Sol (max, 59), Kimi K3 (57), and Claude Opus 4.8 (max, 56) Key takeaways: ➤ New leader in agentic knowledge work: Claude Opus 5 (max) scores 1861 Elo on GDPval-AA v2, >100 points ahead of Claude Fable 5 and GPT-5.6 Sol (max). On AA-Briefcase, our proprietary agentic knowledge work benchmark, it scores 1720 Elo, +146 ahead of Fable 5. These benchmarks test the ability of models to produce accurate and well-presented professional outputs using our open source reference agent harness, Stirrup ➤ Joint first place on the Coding Agent Index: Claude Opus 5 (xhigh) with Claude Code leads the Artificial Analysis Coding Index, including the highest score on SWE-Atlas-QnA ➤ Frontier intelligence with reduced cost: Claude Opus 5 (max) costs $2.03 on average per Intelligence Index task, below Claude Fable 5 (with fallback) at $2.75, but still above Claude Opus 4.8 (max) at $1.80 and Claude Sonnet 5 (max) at $1.53. However, at high and xhigh reasoning efforts Opus 5 can outperform both Opus 4.8 and Claude Sonnet 5 at a lower cost per task ➤ Frontier agentic terminal use: 89% on Terminal-Bench v2.1 at max effort, roughly in line with the leader, GPT-5.6 Sol (xhigh) ➤ Outperformance on scientific reasoning: Along with leading agentic performance, Claude Opus 5 scores 53% on Humanity’s Last Exam in line with Fable 5; on CritPt, a frontier physics evaluation developed by Argonne and UIUC researchers, it also matches Fable 5 but sits behind GPT-5.6 Sol, GPT-5.5 Pro, and GPT-5.6 Terra ➤ Factual knowledge still lags Fable 5: As expected from the models’ size classes, Opus 5 still has lower factual knowledge on AA-Omniscience than Fable 5. It improves +7 points on AA-Omniscience Accuracy over Opus 4.8, but answers more often when uncertain - its hallucination rate rises +14 points to 50% ➤ Improving efficiency, but only on the Intelligence vs. Cost per Task Pareto frontier at high Intelligence levels: Opus 5 outperforms Fable 5 at lower cost, but at lower effort levels it sits just behind the GPT-5.6 family on the Intelligence vs. Cost per Task frontier Other model details: ➤ Context window: 1 million tokens (equivalent to Opus 4.8) ➤ Pricing: As with recent Opus launches, tokens cost $5/$25 per million tokens of input/output; cache pricing remains at a 25% premium for cache writes ($6.25 per million tokens) with 5-minute time to live, and 90% discount for cache hits ($0.50 per million tokens) ➤ Five effort settings (low, medium, high, xhigh, max), and support for server-side fallback as with Fable 5. Intelligence Index evaluations were run with Opus 4.8 fallback enabled
Artificial Analysis tweet media
English
65
210
2.1K
217.4K
🕸️
🕸️@psycvcs·
please timeline AI gw balik lagi dong. udah muak gw liat gambar tap and hold
Indonesia
0
0
0
13