
Baud (YC S26)
28 posts

Baud (YC S26)
@baudlabsai
Multiplier-free ASIC for 100x faster and cheaper AI. #ICML2026
Katılım Haziran 2026
25 Takip Edilen20 Takipçiler
Baud (YC S26) retweetledi

YC and Together AI are partnering to bring the first dedicated YC GPU cluster online, giving YC startups easier access to the compute they need to build and scale.
In this Founder Fireside, YC's @agupta and @togethercompute co-founder & CEO @vipulved dig into why compute has become one of the biggest bottlenecks for modern AI companies, how Together AI is helping more than 8,000 customers—from early-stage research teams to companies like Cursor, Cognition, and ElevenLabs—train, fine-tune, and serve AI models, and why flexible access to GPUs is becoming a competitive advantage for the next generation of founders.
00:00 — Partnering on a GPU Cluster
00:26 — What Together AI Does
01:22 — From Research Labs to Cursor: Together's 8,000 Customers
01:58 — The Landscape of AI Native Startups
03:24 — How Building an AI Company Has Changed Since 2018
04:56 — Why the Cost of Compute Keeps Going Up
05:29 — YC as the Biggest Seed Funder of Research Companies
08:47 — Flash Attention, Mamba, and the Science of Production AI
10:43 — Compute Planning Advice for Early Stage Companies
12:39 — When Your Compute Bill Is Bigger Than Your Cash Balance
13:31 — How Companies Are Using the Cluster Today
14:24 — What's Next
English

What's New Here
Multiplication-free in both training and inference — and lossless. Academia has studied log-domain computation and approximate multipliers for decades, but nearly every approach either compromises accuracy or covers inference only. Our representation holds for both the forward and backward passes, which means you can run the entire lifecycle — from pretraining to deployment — on this chip.
Translating a circuit-level advantage into a system-level advantage. Swapping multipliers for adders is not a micro-optimization — the gap between quadratic and linear scaling directly determines how much intelligence a given piece of 12nm silicon can carry. We don't compete with giants on bleeding-edge process nodes; we change the intelligence density per square millimeter of silicon.
Hardware and software co-designed from day one. The representation, the chip, and the compiler were designed together. That's why we can promise bit-exact PyTorch compatibility — it isn't an afterthought port; it's part of the architecture.
Built for the next generation of workloads. Multimodal models, world models, media generation — the fastest-growing, most compute-dense workloads are exactly where a multiplier-free architecture pays off the most.
We are a small team of experienced chip designers and deep learning engineers with a single goal: to let businesses of all sizes own their artificial intelligence — instead of renting it forever.
English

Baud exists to cure this waste — to bring the efficiency of intelligence toward that of the human brain, and eventually beyond it. Concretely, we have built two things:
First, a new arithmetic representation. It losslessly eliminates multiplications in both the forward and backward passes of neural networks, while compressing the network itself. Note the word lossless — this is not quantization, not an accuracy-for-performance trade. Training and inference both run natively in this representation.
Second, an ASIC designed around this representation. Because multipliers are no longer needed, we can pack an order of magnitude more compute and memory into the same die area. At the same process node and memory bandwidth, our chip is faster, less power-hungry, and simpler to build than incumbents. Our first chip is validated on GlobalFoundries' 12nm process, with tape-out by the end of this year.
Third — and this is the most underestimated piece — a compiler that has existed from day one. The graveyard of ASICs was never dug by hardware; it was dug by software. If engineers have to rewrite kernels, training recipes, and inference frameworks for your chip, no amount of speed matters. Our compiler converts any PyTorch-exportable model directly into our format, with bit-exact results in most cases — mainstream open-source architectures like Qwen, DeepSeek, GLM, Flux, and Wan work out of the box. Your code doesn't change. Your ecosystem doesn't change.
This is not a paper architecture. We are already running the full chip architecture in emulation on FPGAs: a model trained on a single FPGA is serving inference at 1000+ tokens per second, and anyone can try it online. Our first cluster is open for early access, supporting pretraining, fine-tuning, RL post-training, and inference.
English

Own Your Artificial Intelligence: Why We're Building a Chip With No Multipliers
We believe AI can drive scientific discovery and raise the overall quality of human life. For the first time in history, humanity has figured out how to turn sand and sunlight into machines of scientific discovery.
But today, creating and serving frontier intelligence is prohibitively expensive — so expensive that only a handful of companies can afford it. Training a frontier model requires tens of thousands of GPUs and hundreds of millions of dollars in capital expenditure. Even for inference alone, the vast majority of businesses can only pay per token to a small number of providers — forever renting, never owning their intelligence.
Baud (YC S26)@baudlabsai
We are opening early access to our first cluster. Supported workloads include: - Pretraining and fine-tuning - RL post-training - Inference - Up to 32B parameters
English

Introducing OpenScience. A better, open-source Claude Science.
• Any model: GLM, Kimi, DeepSeek, Claude, GPT, your own fine-tune. Switching is one flag.
• 250+ research skills across ML, comp bio, cheminformatics. All readable, editable, extensible.
• No throttling, no gatekeeping, no one vendor deciding what science is okay.
• Native Atlas integration: many agents, one shared reproducible research graph.
• Runs on your infra. Your data stays yours.
Scientific AI should be open. One company shouldn't own the tools the rest of us discover with, or decide who gets to.
English

500+ people RSVP’d to my Luma page before I even hosted a single offline event. Yes, I hacked how people discover ICML side events in Seoul.
Nothing fancy. ICML is happening right in my neighborhood, so I collected every side event I could find and made one useful hub for people coming to Seoul.
The first 100 were slow. Then it started compounding. From 200 to 500 took only a few days.
When people started loving it, I remembered Corgi Cafe from last month. Thanks @UseCorgi!
So we gathered partners who cared. @arizeai and @NEARProtocol helped immediately.
Now we’re opening Cafe @ICML.
It’s less than 5 minutes from the ICML venue. Even better: it’s not a normal cafe. It’s an office, so we have meeting rooms too.
What we need during ICML is not more meetups. We need a place to sit down, rest, drink coffee, and talk about research with good people.
If you need space during ICML, come by.
Thanks to our sponsors, coffee is free.
We’re still taking sponsors. If you sponsor the cafe, we’ll make a drink named after you and help people naturally discover your company or product while they’re here.
luma.com/7iiqamt2?tk=66…
Thank you Corgi, Arize, NEAR, @ekang426, and @sminhan26
English

Heading to #ICML2026 in Seoul?
If you're interested in multiplier-free ASICS for 100x more efficient AI training & inference, feel free to say hi and let's connect!
Baud (YC S26)@baudlabsai
We are opening early access to our first cluster. Supported workloads include: - Pretraining and fine-tuning - RL post-training - Inference - Up to 32B parameters
English

Baud (YC S26) retweetledi













