Dr. Alexander Sheppert

731 posts

Dr. Alexander Sheppert banner
Dr. Alexander Sheppert

Dr. Alexander Sheppert

@alexsheppert

D.O., Ph.D., M.B.A. - Internal Medicine Resident - Artificial Intelligence PhD/Engineer - Founder and CTO at Matic. Thoughts on science, medicine, and AI.

Vancouver, WA Katılım Ağustos 2021
146 Takip Edilen243 Takipçiler
Dr. Alexander Sheppert
Dr. Alexander Sheppert@alexsheppert·
@theo When fable decides it doesn’t want to answer a biology question in a workflow, opus 5 feels like a definite upgrade as a fallback
English
0
0
0
139
Theo - t3.gg
Theo - t3.gg@theo·
How are we feeling about Opus 5 so far?
English
550
10
2.2K
346.1K
Dr. Alexander Sheppert
Dr. Alexander Sheppert@alexsheppert·
Day 2 of nights. Trying to transcend the need for a consistent circadian rhythm.
English
0
0
2
23
Samswara
Samswara@samswoora·
The true blackpill is how quickly humans adapt to godlike technological advancement and take it for granted
Caelin@caelin_sutch

English
70
222
9.7K
279.6K
Dr. Alexander Sheppert
Dr. Alexander Sheppert@alexsheppert·
@thsottiaux The weekly limit is amazing. It allows me to decide I have an important enough task to burn half the week overnight if I want to.
English
0
0
0
21
Tibo
Tibo@thsottiaux·
We've had no 5h limit in Codex plus and pro for a few days. Do you think it is better or are you finding it difficult to manage the usage included in the weekly limit effectively? If we were to make this different, what should it look like in an ideal world?
English
7K
231
13.1K
1.2M
Tibo
Tibo@thsottiaux·
Now that we merged ChatGPT and Codex, what should we merge next? What's the double or nothing move.
English
2.8K
88
5K
810.3K
Dr. Alexander Sheppert retweetledi
phillip
phillip@philliplede·
I used to like Nietzsche until I found out, from his habitual use of the em dash, that he was using ChatGPT to write his laconisms.
phillip tweet media
English
10
13
179
91.5K
Dr. Alexander Sheppert
Dr. Alexander Sheppert@alexsheppert·
@CanesDavid @HeyEpic The useful classification is too specific to have on a dropdown anyway. It’s not like my telephone notes are very long anyway, half the note is why I called the patient!
English
1
0
0
172
𝙳𝚊𝚟𝚒𝚍 𝙲𝚊𝚗𝚎𝚜
Not being able to sign a telephone note because "no reason for call supplied" or some such hard stop nonsense from @HeyEpic tells you everything you need to know about the current state of healthcare. Ok fine... "OTHER" - or "RESULTS" - happy now? Can I sign it? Want to know the reason? Read the note. 😂🤣
English
4
1
39
5.2K
Dr. Alexander Sheppert
Dr. Alexander Sheppert@alexsheppert·
@elonmusk This has been my experience. Grok doesn’t seem to be lazy or take the same shortcuts other models do. I tell it “work until you complete x, y, z” and it actually will!
English
1
0
1
386
Elon Musk
Elon Musk@elonmusk·
Grok 4.5 reaches #1 position on Long-Horizon Terminal-Bench
tetsuo@tetsuoai

The Long-Horizon Terminal-Bench paper landed around May and concluded that the results showed headroom for improvement. The best of the 15 models they tested finished seven of the 46 tasks, and the mean across all models was about two. That ceiling is what fifth place looks like on the current board. Grok 4.5 is now at 13, and Fable 5 is at 12. A single task costs around 9.9M tokens, 231 episodes, and 85 minutes of wall clock time. That means agents are holding a plan across all of it and finishing, and that capability nearly doubled in two months. SpaceXAI is on top, and they marketed the 4.2x output token efficiency, which undersells it. Two dollars in, six out, per million. On a benchmark where one task burns ten million tokens, the bill is dominated by input replay, and they say Grok 4.5 solves tasks in under half the number of steps, so there is less accumulated context to resend on every call. The efficiency compounds on the input side, which is the side that costs money. Fable 5 is one task behind. Their own launch chart has them losing DeepSWE 1.1 to Fable by 17 points, and Grok 4.20 sits on this same board at 0.080 with zero completions, so whatever happened in 4.5 is not a family trait. My read is that the 4.5 jump came out of training alongside Cursor, which is a stream of real agentic edit trajectories nobody else has at that volume, and nothing in the counterevidence argues against it compounding into the next checkpoint.

English
1.9K
2.2K
8.6K
3.2M
Dr. Alexander Sheppert
Dr. Alexander Sheppert@alexsheppert·
I think grok 4.5 is the first model that understands humor to some degree.
English
0
0
0
39
Claude
Claude@claudeai·
We're extending Claude Fable 5 access on all paid plans, as well as keeping Claude Code’s weekly rate limits 50% higher, through July 19.
English
7K
7.3K
74K
27.9M
Polymarket
Polymarket@Polymarket·
JUST IN: JPMorgan reveals its AI investing agents beat a traditional 60/40 portfolio in 20-year backtests.
English
244
243
5K
638.1K
Dr. Alexander Sheppert retweetledi
gabe
gabe@allgarbled·
I explained what a SQL join is to a sales guy and he said it seemed “awesome” and “really useful”
English
194
421
15.8K
678.1K
Bojan Tunguz
Bojan Tunguz@tunguz·
Staying up late so I can take advantage of the next 5 hour usage window.
English
13
2
113
7.2K
ben
ben@contraben·
It's never been a better time to be creative. (ever) The two best frontier models (ever) have been released within 30 days of each other. Here’s what we learned from running 4 frontier models head to head. >Sol has taste and Fable takes direction. 10 identical landing page briefs, judged blind by working creatives. GPT 5.6 Sol won 82% on loose briefs. But when handed a real design spec it finished last. 🧵
OpenAI@OpenAI

Sol, Terra, and Luna, our GPT‑5.6 family of models, are starting to roll out now in ChatGPT, Codex, and the API.

English
17
13
131
35.3K
Ferrariic, MD
Ferrariic, MD@ciirarref·
All attendings/residents should actually read the ct scan instead of trusting what the radiologist says at face value. The radiologist doesn't always know what you're looking for and don't have the context of the physical exam.
English
34
11
141
62.3K
Dr. Alexander Sheppert
Dr. Alexander Sheppert@alexsheppert·
FYI for AI users, my current take: - Fable remains the model with the best taste - GPT 5.6 is fantastic for the money, better than fable for the money - grok build 4.5 is AMAZING. The most cost-effective system in my testing right now. For most important coding where cost isn't as important: - set up codex gpt 5.6 as a callable agent for fable to orchestrate For most other tasks: - grok build 4.5 is impressive and economical - gpt 5.6 sol is great standalone in codex also
English
0
0
0
96
Dr. Alexander Sheppert
Dr. Alexander Sheppert@alexsheppert·
@ciirarref One week out of intern year: “I hope to be the kind of senior resident who inspires the next generation to break the cycle of abuse”
English
0
0
3
120
Ferrariic, MD
Ferrariic, MD@ciirarref·
I thank god everyday that I'm not an intern anymore
English
1
0
31
2.6K
Elon Musk
Elon Musk@elonmusk·
Grok 4.5 is not yet using our internally developed C/C++ inference software that exact maps to the GB300 hardware. Doubling or more of the current speed is probably achievable.
Farzad 🇺🇸 🇮🇷@farzyness

I think AI has just hit a gigantic threshold, and Grok 4.5 is the PERFECT example as to why that is. One of the hardest parts of working with AI is iterating on a project or task that you're working on. As the models have gotten smarter (and more expensive), it's taking longer and longer to get an answer or action back. This creates a ton of stall time per query or action, which is actually quite bad for creativity and staying in a state of flow. You have SO many extended starts and stops. Which inevitably leads to your brain going somewhere else. And then when the AI comes back, you have to redirect your brain to that original task, spool your brain back up to what you were working on at that moment, and then adjust as needed. There's a ton of mental friction involved. This ESPECIALLY sucks when the AI takes a REALLY long time to get something back for you, but it's not quite what you were looking for or asked for. And what sucks EVEN MORE is that these "mistakes" are getting MORE expensive!!! So wait time is going up. AND it costs more per run. HOWEVER - even after using Grok 4.5 for about an hour - what's become obvious is that it's SO MUCH MORE ENJOYABLE AND BETTER to use a model that is FAST... and capable ENOUGH. Capable ENOUGH is the real unlock here. Imagine having Fable 5 performance but at the speed of Gemini 3.5 flash. Or Haiku. That's where we're inevitably going. I think Grok 4.5 (and models like it) have really solved for one of the biggest unlocks in AI - a model that will get you a GOOD ENOUGH answer VERY FAST, at which point iteration can happen VERY QUICKLY. This - counter intuitively - keeps the user in a state of flow and creativity for MUCH longer because you are constantly ENGAGED with your project... instead of letting the AI loose for a long time. And as long as humans are involved, I think 'not quite right' will be a FOREVER problem with AI - because AIs, by default, CANNOT have human taste. Because they are NOT human. But they can be UNBELIEVABLE tools. And unbelievable tools are the ones that are VERY GOOD and VERY FAST. I think that's the true unlock with Grok 4.5 and models like it. Difficult to describe until you experience it. I think this is a VERY big deal for @SpaceXAI and @elonmusk.

English
912
1.5K
13.1K
2.4M