Rick A.F.

1.1K posts

Rick A.F. banner
Rick A.F.

Rick A.F.

@FoxRick01

Building Collie - the first AI Harness for Non-Coders. Former VC/Investor/Consultant and White Collar NPC. -- DM now for Early Access --

Vienna, Shanghai, London Katılım Eylül 2017
543 Takip Edilen162 Takipçiler
Sabitlenmiş Tweet
Rick A.F.
Rick A.F.@FoxRick01·
THREAD TIME: definitive guide on how to improve your AI output & systems x3 EASILY, any model works:
Rick A.F. tweet media
English
1
0
5
316
Brooklyn!
Brooklyn!@imbabybrooklyn·
Do we like node workflows?
English
46
3
323
12.6K
Rick A.F.
Rick A.F.@FoxRick01·
@jenzhuscott Rough but true. Model switching is so easy, literally 0 lock in as long as harness is fine
English
0
0
0
33
Jen Zhu
Jen Zhu@jenzhuscott·
DeepSeek Kill Zone. Models that are inferior and significantly more expensive should cut spending and be in survival mode ASAP. Models that are superior but also more expensive or slightly inferior but with comparable pricing can survive a bit longer.
Jen Zhu tweet media
English
10
8
73
3.7K
Eno
Eno@enodrift·
Which company ships the next 'holy shit' AI moment? Not OpenAI . Anthropic . Google Who?
English
1
0
1
108
Rick A.F.
Rick A.F.@FoxRick01·
@HarryStebbings 7x and then 4x in year 2 and 3. Lets go! I mean, I would love to go even faster.
English
0
0
0
214
Harry Stebbings
Harry Stebbings@HarryStebbings·
I do not care at all about how fast a startup scales $0-$3M ARR. High quality customers, that love the products But I care immensely that you can go from $3M to $20M and 20M to $80M in the subsequent two years. Don’t rush the foundations for scale but once met, run faster than ever.
English
26
8
166
16.3K
Rick A.F.
Rick A.F.@FoxRick01·
@teortaxesTex I spent 2b tokens for $10 last month on Deepseek V4 Pro, the value is unreal!
English
0
0
0
42
Rick A.F.
Rick A.F.@FoxRick01·
@levie Absolutely, the harness has become the bottleneck and tbh its mind-boggling that it creates such a big difference - like what, the frontier intelligence can only figure out stuff with ingrained system prompts? Impressive though.
English
0
0
0
50
Aaron Levie
Aaron Levie@levie·
No idea if these specific numbers generalize across tasks, but directionally it’s clear that the harness is going to become the most important variable -right next to model capability- in the AI stack. The ability for harnesses to break down work in the most efficient way and route to the right model at the right time is going to be a huge variable for maximizing accuracy and reducing costs. We’re actually still incredibly early in this journey. The harness didn’t matter that much when tasks only took hundreds of thousands or millions of tokens. But as we have tasks that take tens of millions and hundreds of millions of tokens, this becomes a major variable. Huge opportunity ahead.
Composio@composio

Hermes and Pi Agent led on the average cost per task, while Claude Code cost about 3.7x as much as Pi: - $0.39 Hermes Agent - $0.40 Pi Agent - $0.47 Codex - $0.51 OpenCode - $0.54 Kimi Code - $1.47 Claude Code The median cost tells the same story: $0.29 in Pi Agent and Hermes, $0.35 in OpenCode, $0.38 in Kimi Code, $0.39 in Codex and $0.72 in Claude Code, so the cost gap holds for a typical task and is not driven by a few expensive runs. We calculated these costs using Kimi K3’s list prices: $3/1M input tokens, $0.30/1M cached input tokens, and $15/1M output tokens.

English
21
23
189
66.2K
Rick A.F.
Rick A.F.@FoxRick01·
@imbabybrooklyn Love Hermes too and thats I am building a harness like Hermes but for non Coders entirely.
English
0
0
0
82
Brooklyn!
Brooklyn!@imbabybrooklyn·
One of the most meaningful milestones for Hermes Desktop: I’ve successfully replaced Cursor for myself. I was effectively a one-person, $20K–$40K/month Cursor enterprise customer for far too long. Now, Hermes’s parallelism and autonomy are genuinely better and faster for how I work. I think that’s because Hermes Desktop is built around a simple premise: coding is no longer about completing one or two tasks at a time. It empowers you to run multiple streams of work without getting in your way, losing the plot, or taking away control. What makes this especially exciting is how much room there still is to improve: context optimization, smarter built-in coding nudges, and so much more. There’s still a long way to go, but reaching the point where I genuinely prefer the product I built feels like a major milestone worth sharing. @NousResearch @Teknium
English
43
23
514
196.5K
Rick A.F.
Rick A.F.@FoxRick01·
@GWazumi @Shogun_AI_ And the harness, some frontier models can be improved by a factor 1-5x just with a better harness.
English
0
0
0
2
Gota. W
Gota. W@GWazumi·
The next AI infrastructure wave isn't about bigger models. It's about memory, context, and world models: systems that sense, remember, and adapt to how you actually work. That's exactly what we're building with @Shogun_AI_.
English
2
2
4
55
Rick A.F.
Rick A.F.@FoxRick01·
@superguenter Insane value of getting local LLM to work now. One 4-8k machine and forever free frontier
English
0
0
1
10
Daniel Han
Daniel Han@danielhanchen·
Unsloth UD quants for DeepSeek V4 Flash 0731 are out! UD-Q8_K_XL 162GB is fully lossless (MXFP4 + BF16) UD-Q4_K_XL 155GB 96% top-1% UD-Q3_K_XL 129GB 87% UD-IQ3_S 117GB 83% UD-Q2_K_XL 97GB 79% UD-IQ1_S 83GB 73%
NVIDIA RTX Spark@NVIDIARTXSpark

Congratulations @deepseek_ai on the launch of DeepSeek V4 Flash 0731! Get started today by running the model locally via the @UnslothAI GGUF. 👇

Čeština
6
11
71
7.2K
Rick A.F.
Rick A.F.@FoxRick01·
@MiaAI_lab 100% every YC founder and new startup will buy like 50k of DGX spark and run them forever. 1% improvement everyday, and even more and here we go.
English
0
0
0
30
Mia
Mia@MiaAI_lab·
If DeepSeek v4 Flash GA is really that good, and you can run it locally on 2 DGX Sparks at 70+ tok/s, I fully expect prices to go up. Right now you can still get the Asus Ascent GX10 for under $4000.
Mia tweet media
English
117
72
1.9K
165.1K
Rick A.F.
Rick A.F.@FoxRick01·
@MiaAI_lab I am building an upgrade on top of PI, Nanobot and Opencode. lets see if I can improve harness performance! Lets go heycollie.com
English
0
0
0
87
Rick A.F.
Rick A.F.@FoxRick01·
@AISystemGuy very interesting, thank you for the overview. Just incredible how it can serve at this cost.
English
0
0
0
16
🤖DR 🇺🇸
🤖DR 🇺🇸@AISystemGuy·
Architecture note · DeepSeek-V4-Flash What’s actually new in V4? V4 is not V3.2 with more parameters. The residual stream, expert router, KV cache, attention output, and shipped numeric formats have all changed. These are notes from porting V4 to MLX, where each of those components had to be implemented from scratch. model_type: deepseek_v4 is currently unsupported by both Transformers and MLX-LM. MLX-LM supports DeepSeek, V2, V3, and V3.2, but V3.2 provides little more than a starting skeleton for MLA and MoE gating. V4 has 304B total parameters across 43 layers, 256 routed experts with six active plus one shared expert, and a one-million-token context window. None of those numbers is the interesting part. Hyper-Connections hc_mult: 4 hc_sinkhorn_iters: 20 The residual stream carries four copies. This is the deepest architectural change. A conventional transformer block reads one hidden state and adds one back. Each V4 block instead carries four parallel copies of the hidden state. Every sublayer: Reduces the four streams to one using learned per-stream weights. Runs attention or the MoE over that state. Expands the result back to four streams. Mixes it with the previous four through a learned combination matrix. That matrix goes through 20 Sinkhorn iterations, alternately normalizing its rows and columns until it is nearly doubly stochastic. This prevents the four residual streams from quietly collapsing into copies of one another. The reduce, expand, and combine weights all come from a single projection of the flattened state, split three ways. Every downstream tensor shape inherits this structure. The hidden state is: [batch, sequence, 4, 4096] The mixing projection is [24, 16384] per sublayer, where: 24 = (2 + 4) × 4 Hash routing num_hash_layers: 3 The first three MoE layers do not score experts at all. A conventional MoE scores every expert against the hidden state and selects the top-k. V4 does that from layer three onward. In the first three layers, it looks up experts in a fixed table indexed by token ID: [129280, 6] int32 Each vocabulary entry maps directly to six experts. Routing in these layers is entirely context-independent: the same token always reaches the same six experts, regardless of what surrounds it. KV compression compress_ratios: [4, 128, …] Compression alternates by layer—and is learned rather than fixed. Attention layers alternate between compression ratios of 4 and 128. Compression uses gated pooling over consecutive tokens: a projection scores each position, applies a softmax across the window, and produces a weighted sum. Ratio-4 layers use overlapping windows. Each position also receives a per-slot embedding so the pooling operation knows where within the window a token appeared. Compressed entries use their own RoPE base—160,000 instead of the standard 10,000—because they represent spans rather than individual positions. Layers without compression use pure 128-token sliding-window attention and disable YaRN entirely. Sparse indexing index_topk: 512 A second attention mechanism decides what the primary attention may read. Alongside the sliding window, each ratio-4 layer runs a 64-head indexer with its own compressor. It scores the compressed KV entries, selects the top 512, and passes those positions to the main attention layer. The main attention then reads: • The local sliding window • Exactly those 512 selected compressed entries The indexer is quantization-aware by construction. Before simulating FP4 on its queries and keys, it applies a Hadamard rotation to distribute information more evenly across channels. Attention output o_groups: 8 o_lora_rank: 1024 V4 adds low rank on the way out—and runs RoPE backward. V3 factored the query projection through a low-rank bottleneck. V4 applies the same idea to the output projection, dividing it into eight groups, each with its own rank-1024 factorization. The detail most likely to be implemented incorrectly: attention applies the inverse rotary embedding to its output before the output projection. Positional rotation is undone on the way out, not merely applied on the way in. Each head also has a learned attention sink: a per-head logit that competes with the real keys, allowing the head to attend to nothing in particular. Numerics quant: fp8 e4m3 expert_dtype: fp4 The checkpoint ships with two quantization formats at once. The published checkpoint is not BF16. It is 167 GB for 304B parameters—about 4.4 bits per weight—because different parts of the model use different formats: • Attention and shared experts: FP8 E4M3, E8M0 scales, 128×128 blocks • Routed experts: packed FP4 E2M1, E8M0 scales, 32-weight blocks The scales are exponent-only. A stored byte b represents 2^(b−127), making every scale an exact power of two and its application lossless. The routed experts are the more interesting case. FP4 E2M1 provides eight magnitudes, packs two values into each byte, and shares one scale across every 32 weights along the input dimension. Smaller changes that still bite scoring_func swiglu_limit: 10.0 The router uses sqrt(softplus(x)) rather than softmax or sigmoid. The bias used to shift expert selection intentionally does not affect the routing weights: routing scores are captured before the bias is added. SwiGLU clamps asymmetrically: • The up projection is bounded at both −10 and +10 • The gate projection is bounded only from above MLA uses a single KV head with a head dimension of 512. What can be omitted Two subsystems live behind a separate forward_spec path and are unnecessary for ordinary generation: • The multi-token-prediction head • DSpark DSpark is a speculative decoder that operates on five-token blocks in the final three layers. It combines a Markov head over token IDs with a confidence head that decides how much of each proposed block to retain. Leaving both out is equivalent to omitting the MTP head when porting another MoE model. V3.2 → V4 summary • Residual stream: single → four streams plus Sinkhorn mixing • Expert routing: scored → hashed for three layers, then scored • KV cache: latent MLA → MLA plus learned compression • Sparse attention: indexer → indexer over compressed KV • Output projection: dense → eight-group low-rank • Router score: sigmoid → square-root softplus • Shipped weights: FP8 → FP8 plus packed FP4 • Low-rank queries: retained from V3.2 The architecture details come from the published DeepSeek’s reference implementation distributed with the weights. The quantization formats were also checked directly against the checkpoint tensors. Both decoding paths reproduce the reference bit-for-bit.
DeepSeek@deepseek_ai

🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! 👇 🔷 The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex! Check out the configuration details in our official API docs: api-docs.deepseek.com/quick_start/ag…

English
1
1
14
1K
Rick A.F.
Rick A.F.@FoxRick01·
@chamath Very nice overview, this is also why I am building the first AI Harness for Non-Coders, e.g. all of white collar without Coding. There are no solutions today and most revert to useless chatbots. Using a companys alpha truly with my harness. Lets go
English
0
0
0
967
Chamath Palihapitiya
Here is my AI investing guide. Sitting here August 2026, my current best thoughts are as follows: 1. LPS (Land Power Shell) is still the most obvious and fastest path to cash on cash returns. Lots of value can be assembled and traded quickly at this layer. And as data centers get more pushback, energized land can explode in value. Very bullish here. I’ve stepped into this layer very aggressively. My partner @anitavlallian and I have acquired almost 6GW coming online in a ramp from today thru 2029 of grid power and behind the meter. 2. Silicon - I helped get @GroqInc off the ground in 2015 and we licensed it to @nvidia for $20B Dec2025. I won’t invest or incubate anything in this layer now. The perf demands of the chips are too high, manufacturing precision is too complex and supply chain influence to get adjacent components like memory isn’t possible for a startup anymore. Lots of capital will be wasted here chasing Groq and Cerebras’ success. Note that both startups made sense a decade ago when these constraints were much more modest. 3. Clouds - Clouds are very very lucrative but very hard to build and very expensive and technically complicated to maintain. And as alignment becomes a more important issue, I expect the clouds will be asked to build robust KYC and attest to it. This makes the risk:reward ratio skewed. I don’t want to be responsible when the USG says a cloud allowed a bad actor to do something bad because of poor KYC. 4. Models are complicated. The big open question is how much of the revenue being generated by them today is because of tokenmaxxing and poor model behavior. If it’s a lot, then the annualized revenues will diminish meaningfully even as token consumption inflects upwards. This is the big economic question at this layer. 5. Harnesses are where the action is and why I started @8090solutions two years ago. In a nutshell, the harness helps enterprises owns their proprietary context (what Alex Karp calls their ‘alpha’). This is an enterprise’s data, workflows, evals, and business rules. A harness that gives this to an enterprise is what creates very low model-agnostic switching costs, which further reinforces my views of #4 above. 6. Applications will be another long term winner along with harnesses. This is where the differentiation between “off the shelf” and “custom time and materials” melts away. Every company, with the right harness, can now imbue their alpha into the software that runs their company. I expect this to mean that “off the shelf” is largely replaced with custom software creating a huge opportunity to write these solutions for companies. Build once and sell repeatedly is a laggard GTM motion for a SaaS world that isn’t needed here. Think custom by design, alpha embedded, proprietary by nature. Fin. Good luck to all the players!
Chamath Palihapitiya tweet media
English
130
170
1.9K
159.2K
Kyle
Kyle@Kyle_Structure·
Drop your account below 👇 I’ll follow people who are: Builders. Creators. Learners. Problem solvers.
English
3
0
5
296
Harsha from Slashy
Harsha from Slashy@GaddipatiHarsha·
I'm almost gonna get to 2k followers. X so easy
English
7
0
9
623
Rick A.F.
Rick A.F.@FoxRick01·
@AnkanXplorer Because it’s daring and not easy. Being in public is tough even on X. Gaining followers is hard too
English
1
0
1
15
Ankan Chowdhury
Ankan Chowdhury@AnkanXplorer·
why do so many genuinely talented people stay completely invisible ?
English
14
2
13
487