Gaurav T

58 posts

Gaurav T

Gaurav T

@jaguarperch

Building enterprise AI agents @Microsoft. Author of “Behind the Interface”. Built https://t.co/im26t1XmMj, https://t.co/HaoRlwfwdw

Katılım Temmuz 2024
229 Takip Edilen6 Takipçiler
Gaurav T
Gaurav T@jaguarperch·
Agentic AI products are going to switch to Grok 4.5. It’s at least 2x faster and an order of magnitude cheaper (cheaper pricing + efficient trajectories) at the same quality as flagship models from A/ and OAI. If your company isn’t making the move, you are going to lose @SpaceXAI
English
1
0
0
14
Gaurav T
Gaurav T@jaguarperch·
Why @AnthropicAI uses @fin_ai for their own website’s chatbot? And chatbot doesn’t even use Anthropic models? Fin’s website says they use their own proprietary models. I guess SaaS not dead yet
Gaurav T tweet media
English
0
0
0
15
Paras Chopra
Paras Chopra@paraschopra·
What software have you built for yourself (and yourself *only*) using Claude Code / Codex / etc.?
English
442
24
838
146K
Dwarkesh Patel
Dwarkesh Patel@dwarkesh_sp·
Who should I interview on my podcast? Open to more AI, but also to random history/econ/etc professors that I might not have heard of before.
English
1.2K
37
1.9K
408.9K
Gaurav T
Gaurav T@jaguarperch·
@dwarkesh_sp @reinerpope The “Behind the Interface” book explains all that in a much easier way - basically how to build a computer from silicon - one layer at a time a.co/d/0i7FOqVh
English
0
0
1
131
Dwarkesh Patel
Dwarkesh Patel@dwarkesh_sp·
New blackboard lecture w @reinerpope How do chips actually work – starting with basic logic gates, and working up to why GPUs, TPUs, FPGAs, and the human brain each look the way they do. 0:00:00 – Building a multiply-accumulate from logic gates 0:16:20 – Muxes and the cost of data movement 0:25:59 – How systolic arrays work 0:39:00 – Clock cycles and pipeline registers 0:51:40 – FPGAs vs ASICs 1:03:14 – Cache vs scratchpad 1:07:16 – Why CPU cores are much bigger than GPU cores 1:11:49 – Brains vs chips 1:15:22 – A GPU is just a bunch of tiny TPUs Look up Dwarkesh Podcast on YouTube/Spotify/etc to watch. Enjoy!
English
92
714
5.6K
934.3K
Gaurav T
Gaurav T@jaguarperch·
@kimmonismus lol the reason isn’t true at all. Rather GitHub Copilot CLI is a great harness and has multi model support.
English
0
0
0
61
Chubby♨️
Chubby♨️@kimmonismus·
Microsoft put $13 billion into OpenAI and built the cloud infrastructure Anthropic runs on. This week it canceled its internal Claude Code licenses because the token bill was too high. Even for MSFT Claude is too expensive.
Hedgie@HedgieMarkets

🦔Microsoft canceled its internal Claude Code licenses this week after token-based billing made the cost untenable, even for a company with effectively infinite cloud resources. Uber's CTO sent an internal memo warning the company burned through its entire 2026 AI budget in just four months. American AI software prices have jumped 20% to 37%, and GitHub (owned by Microsoft) is dropping flat-rate plans for usage-based billing across its products. My Take The AI subsidy era is ending in real time. The same company that put $13 billion into OpenAI and built the Azure infrastructure powering most of Anthropic's compute just looked at the bill from a competitor's coding tool and decided it was not worth paying. That is not a productivity failure on Anthropic's end. Token-based pricing is forcing every enterprise customer to confront the actual cost of running these models at scale, and the number turns out to be far higher than the flat-rate experiments suggested. This ties directly to my Gemini Flash post yesterday. Anthropic, OpenAI, and Google all raised effective prices in the last six months. Enterprises that built workflows assuming AI costs would keep falling are now watching annual budgets evaporate in months. Two outcomes look likely from here. Either enterprises scale back AI usage to fit budgets, which slows the revenue ramp the labs need to justify their valuations ahead of IPOs, or the labs cut prices and absorb the losses, which makes the unit economics worse at exactly the wrong moment. Both paths land in the same place, the numbers stop working, and somebody has to take the writedown. Hedgie🤗

English
67
28
524
84.5K
Arjun Divecha
Arjun Divecha@arjundivecha·
@rohanpaul_ai @adrgrondin What is the use case for running these models on your phone? Your Gemini ChatGPT or Claude app will do a better job If it was so good, why doesn’t Google build Gemma into the Gemini app?
English
5
0
2
793
Rohan Paul
Rohan Paul@rohanpaul_ai·
Incredible possibilities for on-device small models. Here @adrgrondin is running Google’s Gemma 4 E2B on iPhone 17 Pro. ~40tk/s with MLX optimized for Apple Silicon SOTA coding & math on mobile with 128K context. Fully offline with thinking mode.
English
20
35
358
108K
Gaurav T
Gaurav T@jaguarperch·
@scootfleabag @adrgrondin MLX loads entire model in memory. Usage further peaks during a new message on already long conversation. When it crosses 4-5Gb iOS crashes app. Mobile needs to be built different. Instead try it on @get_ten_ai app which optimised for much lower memory fp (with mmap) - same speed.
English
0
0
1
34
Scotty
Scotty@scootfleabag·
@adrgrondin First time an app has caused my phone to fully restart 😅 it works great unless I want to ask more the one question
English
2
0
1
602
Adrien Grondin
Adrien Grondin@adrgrondin·
Google’s Gemma 4 E2B running on-device on iPhone 17 Pro Gemma 4 is built from the same research as Gemini 3, has image understanding capabilities and can reason if needed Running at ~40tk/s with MLX optimized for Apple Silicon
English
246
403
5.9K
1M
Dariton
Dariton@Dariton4000·
@LocallyAIApp Are there any plans on adding web search? That would actually make the app better
English
2
0
4
1.1K
Locally AI - Local AI Chat
Locally AI - Local AI Chat@LocallyAIApp·
Try Gemma 4 now — Google’s most intelligent open models to date. Built on the same research as Gemini 3, with support for reasoning and image understanding. Available on supported iPhones and iPads.
Locally AI - Local AI Chat tweet media
English
47
41
734
80.8K
clem 🤗
clem 🤗@ClementDelangue·
Google just re-entered the game 🔥🔥 They want to take the crown 👑 back from Chinese open source AI. And... Gemma 4 is FINALLY Apache 2.0 aka real-open-source-licensed. From what I've seen it's going to be a pretty significant model. But give it a try yourself today: brew install llama.cpp --HEAD if you have at least 24GB of RAM or VRAM, run the (very good) 26B MOE: llama-server -hf ggml-org/gemma-4-26B-A4B-it-GGUF:Q4_K_M if you have 16GB of RAM or VRAM, run the dense E4B: llama-server -hf ggml-org/gemma-4-E4B-it-GGUF:Q8_0
clem 🤗 tweet media
English
37
62
579
38.1K
Gaurav T
Gaurav T@jaguarperch·
@neural_avb Open source community is 🔥 day 0 support always on MLX / llama.cpp 🥹
English
0
0
0
860
AVB
AVB@neural_avb·
This guy is BEYOND CRACKED. Gemma 4 already on MLX, bro has uploaded all models with quantization. 125 models uploaded in last few hours 🤯 New mlx-vlm repo also supports turbo-quant, and rf-detr too (among other things) If you are a mac dev, you better be jumping at this. Bookmark him, turn his notifications on, sponsor his work.
AVB tweet media
Prince Canuma@Prince_Canuma

mlx-vlm v0.4.3 is here 🚀 Day-0 support: 🔥 Gemma 4 (vision, audio, MoE) by @GoogleDeepMind 🦅 Falcon-OCR + Falcon Perception by @TIIuae 🪨 Granite Vision 4.0 by @IBMResearch New models: 🎯 SAM 3.1 with Object Multiplex by @facebook 🔍 RF-DETR detection & segmentation by @roboflow Infra: ⚡ TurboQuant (KV cache compression) 🖥️ CUDA support for vision models (Sam and RF-DETR) Get started today: > uv pip install -U mlx-vlm Leave us a star ⭐️ github.com/Blaizzy/mlx-vlm

English
66
278
4.2K
654.4K
clem 🤗
clem 🤗@ClementDelangue·
Local is the future of AI because it’s free, safer & faster. It’s also the best way to mitigate current and future compute shortages and distribute control & power! Local AI is having its moment and we’re here for it! The blogpost with all the weights and best ways to use is here: hf.co/blog/gemma4
clem 🤗 tweet media
English
29
29
262
29.8K
clem 🤗
clem 🤗@ClementDelangue·
So happy to see Google release Gemma 4 today in apache 2.0 that gives you frontier capabilities locally. You can use it right away in all your favorite open agent platforms like openclaw, opencode, pi, Hermes by asking it to change your model to local gemma 4 with llama-server. Local AI is having its moment and we’re here for it!
clem 🤗 tweet media
English
74
48
508
26.5K
Gaurav T
Gaurav T@jaguarperch·
@neural_avb @get_ten_ai B is qwen. At least at 2B param level vision, qwen is doing better in my experience. For example here Gemma fails to say there is a bridge
English
0
0
2
107
Gaurav T
Gaurav T@jaguarperch·
@neural_avb Guess which one is Gemma 4 2B and which is qwen 3.5 2B. You can try both on the @get_ten_ai app.
Gaurav T tweet mediaGaurav T tweet media
English
1
0
1
252
AVB
AVB@neural_avb·
Damn, Gemma-4 will have 2B and 4B models that are mutimodal (text, vision, audio) and agentic-pilled. I wanna know how good they are versus the Qwen3.5 2B and 4B models. I love what's happening with these small open models right now. More options is a good thing.
Google DeepMind@GoogleDeepMind

Meet Gemma 4: our new family of open models you can run on your own hardware. Built for advanced reasoning and agentic workflows, we’re releasing them under an Apache 2.0 license. Here’s what’s new 🧵

English
14
7
176
22.3K
Georgi Gerganov
Georgi Gerganov@ggerganov·
Pro tip - hook your PC and Phone with Tailscale and enjoy fast and private inference on the go. Here is Gemma 4, hosted on Mac Studio, streaming to my iPhone. No 3rd party apps. Same WebUI experience everywhere.
English
57
99
1.8K
371.1K
Gaurav T
Gaurav T@jaguarperch·
6/6 It’s fully open-source and free. GitHub 👇 github.com/gauravtendolka… ⭐ Star it if you’d use this 🔁 Repost if you know someone paying for money management apps 💬 Would love contributions
English
1
0
1
39
Gaurav T
Gaurav T@jaguarperch·
1/6 I just asked my bank accounts: “If I cut restaurant expenses to once a month and use that for a travel fund, how many months until I can afford a 3 day Disneyland trip?” It did the math. Across every account! You can do this too - 100% privately. 🧵
English
5
0
0
119
Gaurav T
Gaurav T@jaguarperch·
5/6 That last part is the whole point. YNAB, Copilot, Monarch - your transactions run through their servers. You’re trusting a privacy policy. With PFA, the AI is local. Your bank data never leaves your computer. Not because of a promise. Because it physically can’t.
English
0
0
0
21
Gaurav T
Gaurav T@jaguarperch·
4/6 Setup takes minutes and you stay in full control: → Connect any bank via SimpleFin (read-only access) → A local LLM (Qwen 3.5) runs entirely on YOUR machine → Ask anything in plain English → Get real answers across all your accounts No subscription. No one sees your data.
English
0
0
0
44