Reeve Musk

771 posts

Reeve Musk banner
Reeve Musk

Reeve Musk

@ReeveMuskat05

Starmind

Katılım Mart 2010
1.2K Takip Edilen361 Takipçiler
Reeve Musk retweetledi
Sawyer Merritt
Sawyer Merritt@SawyerMerritt·
SpaceX's Starship rocket just performed a flip maneuver and precisely splashed down on target in the Indian Ocean next to a few buoys with @Starlink on them. The softest splashdown we've seen yet. SpaceX employees ended the stream chanting "USA! USA! USA!"🇺🇸
English
445
2.2K
25.4K
1.9M
Reeve Musk retweetledi
Grok
Grok@grok·
Grok for Excel is live. Use Grok 4.5 to build financial models, analyze market data, and generate charts and graphs. Try it now x.ai/grok/excel
English
552
846
7.5K
24M
Reeve Musk
Reeve Musk@ReeveMuskat05·
Neat
Box@Box

Grok 4.5 reviewed a full Credit & Security Agreement stored in Box — the kind of dense, multi-section facility document that typically requires significant counsel time. @Grok 4.5 used Box MCP to access the file securely, extract key terms across the agreement, identify potential conflicts with existing debt covenants, and compile a summary of items for counsel to review, and finally saved the memo back to the same folder. As frontier models keep leveling up, they are unlocking more opportunities for companies to automate and unlock their enterprise content. Check-out the generated report here: app.box.com/s/zfwfud9ojwbv…

English
0
0
0
5
Reeve Musk
Reeve Musk@ReeveMuskat05·
Cool that Grok 4.5 is #1 in some respects, even with respect to Fable 5
Artificial Analysis@ArtificialAnlys

SpaceXAI's Grok 4.5 takes the #1 spot on AutomationBench-AA with a score of 51%, ahead of Claude Fable 5 (49%) and Claude Opus 4.8 (48%) at roughly a quarter of their cost per task - the first model to complete more than half of workflow objectives without breaking any business rules AutomationBench-AA, our independent leaderboard for @zapier’s AutomationBench, tests whether AI agents can automate real SaaS workflows while adhering to business rules. The test set is private to prevent contamination. Models complete 657 tasks across 40 simulated app environments including Gmail, Google Sheets, Slack, Salesforce, and HubSpot, and the headline score is the share of objectives completed without violating any guardrails. Key takeaways: ➤ Grok 4.5 completes more objectives than any other model: It completes 79.9% of task objectives and strictly passes 21.9% of tasks. This is the highest we’ve measured on both outcomes, exceeding Claude Fable 5’s 73.3% objective completion and Claude Opus 4.8’s 19.3% of fully-completed tasks ➤ Grok 4.5 pushes out the Pareto frontier of score vs. cost per task: At $0.34 per task, it is both cheaper and higher-scoring than every other leading model - Claude Fable 5 ($1.35 per task), Claude Opus 4.8 ($1.46), GPT-5.5 (xhigh, $1.28), and Gemini 3.5 Flash (high, $0.49) ➤ It is extremely token-efficient: Grok 4.5 uses ~8k output tokens per task, the fewest of any leading model - less than a quarter of Claude Opus 4.8 (32k) and a third of Gemini 3.5 Flash (24k). Its total token usage of 0.44M per task is among the lowest on the leaderboard. Low cost is driven by this efficiency as well as low token pricing ➤ Grok 4.5 uses fewer turns with many parallel tool use: Grok 4.5 resolves tasks in ~16 turns, fewer than GPT-5.5 (xhigh, 25) and less than half of Gemini 3.5 Flash (high, 35), while making the most tool calls per task of any leading model (52.5). It batches 3.3 tool calls per turn, compared to ~2.5 for Claude Opus 4.8 and ~2.0 for GPT-5.5 (xhigh) ➤ Guardrails still get broken: Grok 4.5 triggers 0.63 violations per task, above Claude Opus 4.8 (0.55) and Gemini 3.5 Flash (0.46). At 13.0 objectives completed per violation, it trails Gemini 3.5 Flash (15.0) and Claude Opus 4.8 (13.5) ➤ Its strongest lead is in the hardest domain: Grok 4.5 completes 71% of Finance objectives, the domain with the lowest average score, ahead of Claude Fable 5 (64%) and Claude Opus 4.8 (62%) Congratulations to @SpaceXAI and @elonmusk on topping the leaderboard!

English
0
0
0
28