Gavin Fuller

2K posts

Gavin Fuller

Gavin Fuller

@GavinFullerTX

Azure Cloud & AI Consultant. My opinions are my own.

Katılım Aralık 2022
1.5K Takip Edilen97 Takipçiler
Gavin Fuller retweetledi
Artificial Analysis
Artificial Analysis@ArtificialAnlys·
DeepSeek V4 Flash 0731 is now open weights! @deepseek_ai has just released the weights for its new flash tier model, DeepSeek V4 Flash 0731. With a score of 50 on the Artificial Analysis Intelligence Index, it lands among the top 3 open weights models on the leaderboard. The weights are released under the MIT license, allowing unrestricted commercial use and modification. DeepSeek V4 Flash 0731 shares identical architecture and pricing with the earlier DeepSeek V4 Flash. At a size of 284B total parameters (13B active), released in mixed FP4/FP8 precision at ~167GB total file size, it lands on our Pareto frontier for Intelligence Index vs. Total Parameters. Among open weights models, DeepSeek V4 Flash 0731 delivers a significant leap in intelligence for its size class. DeepSeek V4 Flash 0731 is also available now through DeepSeek's first-party API. Check out Artificial Analysis to compare DeepSeek V4 Flash 0731 with other leading open weights and proprietary models: artificialanalysis.ai/models
Artificial Analysis tweet media
English
27
63
714
34.7K
Gavin Fuller retweetledi
Anthropic
Anthropic@AnthropicAI·
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. anthropic.com/news/investiga…
English
1.8K
2.2K
12.9K
17.3M
Gavin Fuller retweetledi
Peter Dimov
Peter Dimov@psdimov·
@MTSlive We’re like 2 more rounds of them trying to one up each other until they start confessing to international war crimes
English
0
3
296
7.7K
Gavin Fuller retweetledi
Gavin Fuller retweetledi
Gavin Fuller retweetledi
Victor E. Nunez
Victor E. Nunez@nunezvice·
let’s put this 80% price drop for 5.6🌙 in perspective three weeks ago, gpt-5.5 xhigh was our frontier. it completed 67% of tasks on DeepSWE. today, luna max matches that score at about $0.12 per task instead of $7.23. same score. roughly 60x cheaper. three weeks later
Simon Willison@simonw

OK, GPT-5.6 Luna is a bit of a beast. Given the 80% price drop today I decided to try it in Datasette Agent, and it's furiously quick and generates all the SQL, HTML and JavaScript (for Datasette Apps) I could possibly want

English
54
93
1.8K
216.9K
Gavin Fuller retweetledi
DeepSeek
DeepSeek@deepseek_ai·
⚠️ Note 🔷 DeepSeek-V4-Flash-0731 keeps the exact same model architecture and size as the preview version. 🔷 Today's upgrade applies ONLY to the DeepSeek-V4-Flash API. The DeepSeek-V4-Pro API and App/Web models remain unchanged for now. The official release of DeepSeek-V4-Pro is coming ASAP! Stay tuned.
English
65
127
3.2K
342.6K
Gavin Fuller retweetledi
DeepSeek
DeepSeek@deepseek_ai·
🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! 👇 🔷 The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex! Check out the configuration details in our official API docs: api-docs.deepseek.com/quick_start/ag…
DeepSeek tweet media
English
1.4K
3.1K
25.7K
6.9M
Gavin Fuller retweetledi
Cline
Cline@cline·
DeepSeek silently updated their changelog with a new V4-Flash upgrade 1 hour ago. Their new Terminal-Bench score is 82.7, a massive +25.8 point leap from its initial April preview score of 56.9. Currently only available via their API, open weights release will follow shortly.
Cline tweet media
English
175
361
4.8K
1.1M
Chubby♨️
Chubby♨️@kimmonismus·
This is just insane. For me the flash release is another DeepSeek moment. Intelligence so cheap and so good that it’s literally too cheap to meter. Many don’t get how insane this release is.
Chubby♨️ tweet media
Cline@cline

DeepSeek silently updated their changelog with a new V4-Flash upgrade 1 hour ago. Their new Terminal-Bench score is 82.7, a massive +25.8 point leap from its initial April preview score of 56.9. Currently only available via their API, open weights release will follow shortly.

English
195
416
6.4K
464.9K
Gavin Fuller
Gavin Fuller@GavinFullerTX·
@thsottiaux I’ve found multiple concurrent threads seem to choke pretty hard even on high spec machines, to the point where even my terminal visibly stalls. Same issue does not replicate on omp and a few other harnesses at the same scale with the same models
English
0
0
0
22
Tibo
Tibo@thsottiaux·
What should we improve on Codex to improve the everyday experience? Nothing too small
English
7.4K
107
5.8K
1.1M
Gavin Fuller
Gavin Fuller@GavinFullerTX·
Just realized I burned $100 in Codex credits in 1 hour and 8 minutes last night because I forgot to use a banked rate limit reset before going to bed
GIF
English
0
0
2
217
Gavin Fuller
Gavin Fuller@GavinFullerTX·
Realizing OpenAI has completely hooked me on GPT-5.6 Sol Ultra on fast mode with @thsottiaux’s (almost) daily rate limit resets. Don’t know how I can go back at this point, it just keeps nailing everything I care about.
English
0
0
1
40
Gavin Fuller retweetledi
Cognition
Cognition@cognition·
We've updated FrontierCode 1.1 to reflect new discounts for GPT-5.6 Terra and GPT-5.6 Luna. With these new costs, the GPT-5.6 series sits on the pareto curve of price/performance efficiency.
English
45
64
1.4K
389.4K
Gavin Fuller
Gavin Fuller@GavinFullerTX·
@zephyr_z9 Luna on max reasoning is roughly on par with Sonnet 5 at a fraction of the price
English
1
0
13
4K
Zephyr
Zephyr@zephyr_z9·
OpenAI cut Luna prices from $1/$6 to $0.2/$1.20 A 5x price cut OpenAI begins the price war, with the goal of crushing the Chinese players Not going after Anthropic's margins yet
English
105
125
3.7K
363.3K
Gavin Fuller retweetledi
Tibo
Tibo@thsottiaux·
Changes both affect API prices and usage in paid subscriptions. Something for everyone.
English
140
36
1.7K
213.6K