George Cameron

281 posts

George Cameron banner
George Cameron

George Cameron

@grmcameron

Co-Founder @ArtificialAnlys | Message me to play 🎾 in SF

San Francisco Katılım Ocak 2022
375 Takip Edilen931 Takipçiler
Kyle Wong
Kyle Wong@ewveggies·
Yesterday was my 23rd birthday. As a present, I got laid off from Amazon AGI! I’m actively looking for new opportunities, particularly related to computer use agents and post-training. Since October 2024, I've been working on GUI agents non-stop: computer use at @SimularAI, mobile use @Apple, and browser use @AmazonScience. With recent benchmarks like OSWorld 2.0, ALE, etc., it's clear CUA is far from solved, and would love to continue pushing. Please reach out if you have any openings!
Polymarket@Polymarket

JUST IN: Amazon cuts jobs from its artificial general intelligence team.

English
75
17
960
191K
George Cameron
George Cameron@grmcameron·
For those impacted by Amazon AGI team layoffs - Artificial Analysis is hiring Members of Technical Staff to build evals that benchmark the next frontiers of AI. Please reach out
English
1
3
86
17.3K
George Cameron retweetledi
Artificial Analysis
Artificial Analysis@ArtificialAnlys·
Kimi K3 launched yesterday, and within hours our billboards across San Francisco had already been updated to reflect the new frontier The Artificial Analysis Intelligence Index moves as fast as AI does - whether you're checking our website or waiting for the 38 Geary.
Artificial Analysis tweet media
English
21
87
1.3K
70.1K
George Cameron
George Cameron@grmcameron·
Interesting fact people might not know: Moonshot AI (Kimi models) founder Yang Zhilin, who also goes by the nickname 'Kimi', studied under Jie Tang who is a co-founder of Z AI (GLM models) at Tsinghua university (before CMU)
George Cameron tweet media
English
23
71
1.4K
249.2K
George Cameron retweetledi
Artificial Analysis
Artificial Analysis@ArtificialAnlys·
Kimi K3 scores 57 on the Artificial Analysis Intelligence Index. Its intelligence is comparable to Opus 4.8 and GPT-5.5 but remains behind Fable 5 and GPT-5.6 Sol. Moonshot AI has expressed plans to release the 2.8T parameter model's weights, which would make it the leading open weights model Key results: ➤ Strong agentic task performance: @Kimi_Moonshot's Kimi K3 reaches an Elo rating of 1668 on GDPval v2. This is a marked improvement over K2.6’s 1190, surpassing GLM-5.2 (1514), GPT-5.5 (1494), and Claude Opus 4.8 (1600). However, it still lags behind Claude Fable 5 (1760). Kimi K3 also scores an impressive 53% and takes the #1 position on AutomationBench-AA, our implementation of Zapier’s Agentic SaaS workflow evaluation. ➤ Second-highest performance on AA-Briefcase (agentic knowledge work): On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5. It is well-rounded: its rubric scoring and analytical quality almost reach Claude Fable 5’s scores, while GPT-5.6 Sol continues to outperform other leading models on presentation quality. ➤ Set to lead open weights models once weights are released: Moonshot AI has not yet released the weights but expressed plans to do so. Once available, Kimi K3 would clearly lead other open weights models including GLM-5.2 (51) and DeepSeek v4 Pro (44). However, at 2.8T parameters, it is significantly larger than its open weights peers (eg. GLM-5.2 at 753B params and DeepSeek V4 Pro at 1.6T), as well as the Kimi K2 to K2.6 models (1T params). ➤ Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers: Moonshot AI’s pricing for K3 is significantly higher than their K2 pricing (K3’s output token price is $15/1M tokens while K2.6 was $4). This positions the model as cheaper on a cost per task basis than Opus 4.8, similar to GPT-5.6 Sol ($1.04) and more expensive than open weights peers, GLM-5.2 ($0.32) and DeepSeek V4 Pro ($0.04) ➤ Improved token efficiency alongside higher intelligence: Kimi K3’s token usage on the Artificial Analysis Intelligence Index decreased significantly, using 21% fewer output tokens than K2.6. The new model used approximately 132M output tokens to complete all nine evaluations, compared to approximately 166M for K2.6, while achieving higher scores. ➤ Native multimodal capabilities: Kimi K3, like K2.6, is released with native image and text multimodal input. If weights are released, this will position Kimi K3 as one of the leading open weights models with multimodal input capabilities Other model details: Context window: 1M Size: 2.8T total parameters Pricing: The first-party API is priced at $3.00/$15.00 per 1M input/output tokens, with cached input discounted 90% to $0.30 per 1M tokens. Modality: Native multimodal input supports text and images, and the model remains text-only for output. Accessibility: Accessible at launch through Moonshot’s first party API. Model weights are not yet released but Moonshot AI has expressed plans to do so.
Artificial Analysis tweet media
English
180
714
6.2K
1.7M
George Cameron retweetledi
Artificial Analysis
Artificial Analysis@ArtificialAnlys·
Our co-founders Micah (@_micah_h) and George (@grmcameron) spoke at AI Engineer World’s Fair about the cost of intelligence for agentic tasks. While accessing a fixed level of intelligence is becoming cheaper, the cost of completing complex agentic tasks is increasing. As models take on more complex agentic real-world tasks, they have larger context requirements and are relying on additional turns and tool calls which lead to higher cost per task. To reflect these workloads and capabilities, we developed AA-Briefcase, our new agentic knowledge work benchmark, designed to measure models on long-horizon knowledge work tasks. Models are evaluated on tasks across multi-week projects and several thousand input files. As discussed in the talk, cost per task is becoming an increasingly key metric and is a combination of the four following factors: ➤ Token price: The price of accessing a given level of intelligence continues to decline on a per-token basis. However, as models become capable of completing longer horizon work, intelligence and token usage have grown faster, resulting in a higher overall cost per task. ➤ Turns: Agentic work requires multiple ‘turns.’ All turns are included in the context of subsequent turns, converting prior output tokens into input tokens and causing costs to accumulate throughout the trajectory, even with input caching discounts. Some models require significantly more turns than others, without necessarily higher performance. ➤ Token efficiency: Models vary significantly in how many tokens they use to complete the same task. A more verbose model may cost more per task even when its token prices are lower. Looking at the broader picture of ‘cost per task’ helps to contextualize token efficiency and token cost. ➤ Prompt caching & hit rate: Long agentic tasks repeatedly pass the same instructions, research, tool outputs, and prior work back into the model. Prompt caching reduces the price of these reused input tokens, making the difference between standard input pricing and cache-hit pricing increasingly important as the context increases. Watch the full keynote: youtube.com/live/4sX_He5c4…
YouTube video
YouTube
English
3
2
42
17.2K
George Cameron
George Cameron@grmcameron·
@swyx Cinematic and inspiring. Lots to relate to and lessons to learn from
English
0
0
1
450
clem 🤗
clem 🤗@ClementDelangue·
We just crossed $100M annual run-rate. I know many AI companies are capturing much more $$$ these days, but still proud of the milestone! Maximizing short-term revenue has never been our priority. In fact, we're proud to manage to store and serve hundreds of petabytes of models and datasets while keeping HF free and open-source for 97% of our users. As a platform, we’re happy to hopefully create orders of magnitude more value for the community than what we capture. To me, that’s the very definition of a platform. And it has helped us build one of the most loved platform in tech, with network effects, a defensible position and a sustainable business which is quite unique in AI. Many many thanks to all the community members for building with us, we wouldn't be anywhere without you! Can’t wait for what’s next, especially as more companies start to see the value of open and local AI! Next milestone $1B?
English
234
118
2.4K
194.7K
George Cameron
George Cameron@grmcameron·
@oleksoleksoleks Which benchmarks did you run? We do pass@1 across repeats and don't re-run on the same checkpoint unless actual errors encountered. one thing to check is if you're setting reasoning effort correctly on glm-5.2? It's their first model with configurable reasoning effort
English
0
0
5
55
George Cameron
George Cameron@grmcameron·
Our biggest release ever at Artificial Analysis. Exciting and very interesting results - it has been hard keeping quiet
Artificial Analysis@ArtificialAnlys

Announcing AA-Briefcase, the benchmark for the next era of agentic knowledge work AA-Briefcase is our new benchmark for testing models on long-horizon knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week projects, each with many linked tasks and thousands of input source files. We evaluated Claude Fable 5 from @AnthropicAI before it became unavailable, and it currently leads with an Elo score of 1587, followed by Claude Opus 4.8 (max, 1356), Opus 4.7, and the recently-released GLM 5.2 (max, 1266) from @Zai_org. Claude Fable 5 cost $31 on average to run each AA-Briefcase task, followed by Claude Opus 4.8 at $10.40, GPT-5.5 (xhigh) at $3.68 and GLM-5.2 (max) at $2.40. AA-Briefcase comprises four private scenarios, each representing a multi-week knowledge work project set in a realistic organizational context. A public fifth scenario has been released via @huggingface as a representation of scenario structure, submission, and grading (AA-Briefcase Lite). This does not count toward official AA-Briefcase results, and is demonstrative only. Key elements of AA-Briefcase: ➤ Realistic long-horizon projects: AA-Briefcase moves beyond single, disconnected prompts by evaluating models across a coherent long-horizon project. Tasks build week by week, draw on shared institutional context, and require deliverables such as financial models, board presentations, and design mock-ups ➤ Large volumes of fragmented context: AA-Briefcase requires models to reason across thousands of inputs, including company documents, meeting transcripts, large-scale data exports, 25,000+ Slack messages and 3,500+ emails. These sources are fragmented, messy, and often contain realistic contradiction, testing whether models can navigate the ambiguity of real-world knowledge work ➤ Composite rubric and pairwise grading: AA-Briefcase combines binary rubric checks for ground-truth correctness with pairwise grading on analytical quality and presentation quality. Unlike many evaluations that focus on a single metric, AA-Briefcase tests agentic capabilities more comprehensively, exposing cases where models produce outputs that look polished but are incorrect or lack analytical rigor ➤ Built by industry experts: AA-Briefcase scenarios mirror real-world knowledge work, with tasks developed over months by experts across data science, product management and corporate strategy from companies including Google, McKinsey & Company and BCG. Task challenges are drawn from professional experience, making AA-Briefcase more reflective of the ambiguity, messy context and competing priorities that define real-world knowledge work Key results: ➤ Claude Fable 5 leads AA-Briefcase at 1587 Elo: This is followed by Claude Opus 4.8 (1356) with the next-best non-Anthropic model, GLM-5.2 (max), ~90 points back at 1266. Note that Claude Fable 5 did not use the Opus 4.8 fallback for any task in AA-Briefcase ➤ Cost per task varies by ~800x across models tested: Claude Fable 5 leads the benchmark but costs more than $31 per task on average, compared to ~$0.04 for DeepSeek V4 Flash (max). The strongest price/performance options are open weights models such as GLM-5.2 (max) and DeepSeek V4 Pro (max), with GLM-5.2 (max) scoring only ~90 Elo below Claude Opus 4.8 (max) for less than 25% of the cost ➤ Real-world complexity remains difficult for models: The top performer, Claude Fable 5, satisfies all rubric criteria on just 3% of AA-Briefcase tasks. On 31 of 91 tasks, no model scores above 50% on the rubric criteria ➤ Task difficulty scales with the number of required input files: For each rubric check, we identify the set of source files needed to pass. Across all models, pass rates fall as this file count increases, though top-tier models degrade less than weaker models More details below in thread ⬇️

English
0
0
14
1.6K
George Cameron retweetledi
Artificial Analysis
Artificial Analysis@ArtificialAnlys·
Today we're releasing the first results for AA-AgentPerf, our new agentic inference benchmark: initially covering DeepSeek V4 Pro across NVIDIA Blackwell, Hopper, and AMD. AA-AgentPerf is the first benchmark built for agentic inference. We use real, long-context agentic coding trajectory data as the workload, and inference with real production optimizations such as KV cache reuse and speculative decoding, leading to the most realistic evaluation of inference performance available today. AA-AgentPerf’s lead metric is Agents per Megawatt. In a power-constrained world, this answers the most relevant question for AI infrastructure providers - “how many real agents can I deploy per unit of power available?”. First results for DeepSeek V4 Pro (at the easiest defined service level of 20 tokens/s and 10s TTFT): ➤ GB300 (rack-scale, disaggregated): 61,354 Agents/MW ➤ B300 (single node, disaggregated): 21,053 Agents/MW ➤ MI355X: 3,551 Agents/MW ➤ H200: 2,594 Agents/MW Further AA-AgentPerf details: ➤ Real agent workloads, beyond synthetic queries: AA-AgentPerf replays real coding agent trajectories where our agents used up to 200 turns and worked with sequence lengths >100K tokens - the workloads that matter in 2026 ➤ Production optimizations allowed: KV cache reuse, speculative decoding, and prefill/decode disaggregation are all permitted, with accuracy verification to control for quality loss - we want results to reflect what real deployments actually look like ➤ Lead metric is Agents per Megawatt: simultaneous agents supported at production performance targets (e.g. 20 tokens/s per user, ≤10s TTFT) per megawatt consumed. Agents per TCO and $/hr will be supported soon Key findings: ➤ Rack-scale disaggregated inference (GB300) is ~3× more power-efficient than single-node Blackwell (B300), and similarly ahead in raw agents per GPU ➤ Blackwell represents a large generational step over Hopper in both power efficiency and raw compute per GPU ➤ In this test, NVIDIA's Blackwell systems currently lead AMD MI355X by a clear margin. Important context: our MI355X configs are approximately two weeks older than our Blackwell configs and couldn’t stably use speculative decoding. MI355X power draw under heavy load is also well below TDP, indicating there is much room to improve on DeepSeek V4 Pro, which we will measure and publish in the coming weeks ➤ Config and inference framework version matter enormously - we've seen meaningful improvements daily since the DeepSeek V4 Pro release and look forward to tracking performance over time AA-AgentPerf is a live benchmark and we publish results on a rolling basis as submissions come in. Some of the new features coming in v1.1: more models (gpt-oss-120b), more hardware (GB200, B200, H100, MI300X), better AMD configurations, $/hr and cost-per-task normalization, Agents per TCO, and performance tracking over time.
Artificial Analysis tweet media
English
17
26
281
2.1M
Nathan Lambert
Nathan Lambert@natolambert·
My time at Ai2 / @allen_ai has come to an end. Ai2 is a wonderful place. The last 2.5+ years building Olmo, Tulu, and other projects will be one of the peaks of my entire career. I'm extremely thankful for my teammates and the open community who made this work possible. For me, it's time to try something different. I will still be working in the open model & open science spaces (more news on that soon). In the meantime I'll be spending a few months learning, chatting with a broader network, getting married (!!) and most importantly recharging from pouring my soul into this place. I've attached the note I shared with the team and some fun photos from our time together. I'll keep cheering for Ai2 and am excited to see what you build next.
Nathan Lambert tweet mediaNathan Lambert tweet mediaNathan Lambert tweet mediaNathan Lambert tweet media
English
142
41
1.8K
152.5K
Sarah Chieng
Sarah Chieng@MilksandMatcha·
Changed my last name and fled the country 💍
Sarah Chieng tweet media
English
146
8
1.3K
155.6K
Lin Qiao
Lin Qiao@lqiao·
We just hit a major milestone — @FireworksAI_HQ passed $800M annualized run rate and reached 4x revenue growth, apart from Cursor, in Q1. We invite curious and courageous minds to join us and define new frontiers of specialized intelligence!
English
64
35
857
199.6K
George Cameron
George Cameron@grmcameron·
Cursor Composer models previously never made sense to use considering other models out there. Now there is a reason to use it, almost as good as Opus 4.7 and GPT-5.5 but >10X cheaper
Artificial Analysis@ArtificialAnlys

Cursor's new Composer 2.5 takes third on the Artificial Analysis Coding Agent Index and is ~10-60x lower cost than the higher-effort Opus 4.7 and GPT-5.5 variants above it. This release puts Composer among the leading coding agent models, something that wasn’t clear for past releases @cursor_ai has released Composer 2.5, the latest model in its Composer line. Composer 2.5 scored 62 on our Coding Agent Index, a 14 point gain over Composer 2 (48). This puts it in third place of our tested agents, behind only Claude Opus 4.7 (max) in Claude Code (66) and GPT-5.5 (xhigh reasoning) in Codex (65). These cost $4.10 and $4.82 per task respectively, ~10x the cost of Composer 2.5 Fast ($0.44) and ~60x the cost of Composer 2.5 standard ($0.07). Key results for Composer 2.5 in Cursor CLI: ➤ Cost-quality Pareto frontier: At $0.07 (standard) and $0.44 (Fast) per task, Composer 2.5 is cheaper than every other agent scoring above 60 on the Index. Medium-effort peers cost $1.24–$2.21 per task; higher-effort variants land 3-4 points above at $4.10–$4.82 ➤ Per-benchmark gains vs Composer 2: +35 points on SWE-Bench-Pro-Hard-AA (12% → 47%), +2 points on Terminal-Bench v2 (64% → 66%), and +3 points on SWE-Atlas-QnA (69% → 72%). At 47%, Composer 2.5's score on SWE-Bench-Pro-Hard-AA is comparable to Claude Opus 4.7 (max) in Claude Code ➤ Among the fastest coding agents: Composer 2.5 Fast runs at an average wall time of 6.7 minutes per task, the third-fastest agent on the Artificial Analysis Coding Agent Index, behind only Claude Opus 4.7 (medium) in Claude Code (5.8m) and GPT-5.5 (medium) in Cursor CLI (6.2m) ➤ Fast mode enables better responsiveness at 6x pricing: Fast runs 30% faster than standard Composer 2.5, but is ~6x the cost per task ($0.44 vs $0.07). Token pricing is 6x higher for Fast: $3.00/$15.00 vs $0.50/$2.50 per million input/output tokens Model details: ➤ Base model: Continued training on @Kimi_Moonshot's open weights Kimi K2.5 as with Composer 2, with Cursor reporting ~85% of total compute from its own additional training and reinforcement learning ➤ Pricing: $0.50/$2.50 per million input/output tokens for the standard variant; $3.00/$15.00 for the Fast variant (the default in Cursor) ➤ Available exclusively in Cursor: both Cursor IDE and Cursor CLI, an externally accessible API is not available Congratulations @cursor_ai and @mntruell on the impressive release!

English
1
0
2
433
Paddy Srinivasan
Paddy Srinivasan@paddix·
We just set the bar for inference performance. Fastest DeepSeek, MiniMax & Qwen on Blackwell Ultra. #1 output speed. Sub-1s latency on @digitalocean's AI-Native Cloud Platform as validated by @ArtificialAnlys 230 tok/s on DeepSeek V3.2 which is 3.9x faster than AWS Bedrock and Sub-1s TTFT. This is what we unveiled at Deploy earlier this week: 👉 An AI-Native Cloud built for the Inference Era 👉 A unified Inference Engine — speed, cost, simplicity Red the article below to understand how we engineered this with techniques inlcuding Tensor Parallelism, Kernel Fusion, Programmatic Dependent Launch, Speculative Decoding and Multi-Token Prediction (MTP) including working closely with the creators of vLLM at @inferact digitalocean.com/blog/how-we-bu…
English
7
14
80
6.8K
George Cameron
George Cameron@grmcameron·
@__codewithkai__ @ArtificialAnlys This is just a legend component for where publishing claimed results. As none of the bars in the chart have striped pattern none are lab claimed - all these are independently benchmarked by us!
English
0
0
0
23
George Cameron retweetledi
Artificial Analysis
Artificial Analysis@ArtificialAnlys·
Anthropic launched Claude Opus 4.7 today, the new #1 in our GDPval-AA benchmark for performance on agentic real-world work tasks Opus 4.7 scored 1753 on GDPval-AA at launch with its ‘max’ effort setting, surpassing GPT-5.4 xhigh. This is a significant upgrade, placing Opus back on top of Sonnet on the GDPval-AA leaderboard. Compared to OpenAI’s GPT-5.4, it has an implied win rate of ~60% when compared head-to-head on the GDPval task set. We supported @AnthropicAI with testing this model ahead of release and appreciate them referencing our evaluations in their announcement post and system card for both GDPval-AA and AA-Omniscience. We’re actively conducting the rest of the Artificial Analysis Intelligence Index evaluations and will share complete results soon!
Artificial Analysis tweet media
English
28
35
541
36.4K