Scott Condron

3.2K posts

Scott Condron banner
Scott Condron

Scott Condron

@_ScottCondron

Helping build AI dev tools at @weights_biases. I post about AI, data visualisation and the stuff I’m working on at wandb.

Dublin, Ireland Katılım Nisan 2018
2.1K Takip Edilen5.7K Takipçiler
Sabitlenmiş Tweet
Scott Condron
Scott Condron@_ScottCondron·
Here's an animation of a @PyTorch DataLoader. It turns your dataset into a shuffled, batched tensors iterator. (This is my first animation using @manim_community, the community fork of @3blue1brown's manim) Here's a little summary of the different parts for those curious: 1/5
English
34
487
2.6K
0
Scott Condron retweetledi
CoreWeave
CoreWeave@CoreWeave·
Quantitative trading is entering the foundation model era. Flow Traders has selected CoreWeave as its primary AI cloud platform provider for its AI and deep learning division.. As AI becomes central to quantitative trading, firms need consistent performance at scale, not cloud designed for general-purpose computing. We're excited to support @FlowTraders as they build the next generation of AI-driven trading. Learn more here: utm.io/urQl5
CoreWeave tweet media
English
8
14
161
12.9K
Scott Condron
Scott Condron@_ScottCondron·
@Vtrivedy10 - easier infra for serving & finetuning custom models We've made decent strides here. You can upload a LORA adapter to a bucket and you can use and pay for these tokens as normal by just swapping the model slug. No hosting needed really
English
1
0
1
107
Viv
Viv@Vtrivedy10·
big opportunity to create open source tooling + recipes to bring the same data curation work that’s done for the labs inside every org part of owning your intelligence is making your data useful for your agents there’s amazing energy around open source rn -open models crossed a frontier level intelligence threshold - easier infra for serving & finetuning custom models there’s a lot still to do at the data layer! data companies feel like black boxes to the average builder -> there’s some massive data creation/collection process that the labs buy up open tooling tooling to interview domain experts and turn their knowledge into evals and skills open tools to iteratively turn trace data into environments for training and hill climbing lots to do in helping everyone own the pipeline for data -> better ai for themselves
English
10
2
65
4.1K
Scott Condron retweetledi
Alex Ratner
Alex Ratner@ajratner·
All enterprises need some kind of specialized AI to be non-commodity in the AI era - and for that, they need a data flywheel. However: data flywheels are *not* built by just passively collecting user/agent interaction traces and then tuning on these. That is like telling a student to study only using their ungraded practice exams. Most of the alpha and effort in AI today is getting *ground truth* - i.e. the grading keys/grades/answers to the practice exams - that enable tuning/RL to work. This is all about *high quality data* in the format of rubrics (grading keys), human evals (grades), and/or "gold" traces (correct answers). Without this feedback signal, a data flywheel is fundamentally incomplete - and tuning on it will only reinforce inaccurate agent behavior. Specialized AI that works is built with data flywheels that are continuously *developed* with the right annotation and data!
English
5
14
49
3.1K
Scott Condron
Scott Condron@_ScottCondron·
I’ve got this working with Marimo-pair (giving coding agents a way to see the marimo runtime), it’s a great way to combine custom domain specific interfaces with coding agents. Will show a demo next week
Hamel Husain@HamelHusain

Underrated way to start evals with just a coding agent 1. Cluster your traces 2. Build your own annotation app 3. Have AI monitor your annotations in real time to adapt your sampling + propose new annotations to accept /reject @sh_reya shows how here: youtube.com/watch?v=tqUDjc…

English
0
0
0
653
Scott Condron
Scott Condron@_ScottCondron·
Please let this catch on so we can all enjoy a standardised trajectories format and make fine-tuning, memory and running judges on trajectories more efficient and consistent thank you
Charles Packer@charlespacker

At @Letta_AI we spent a lot of time parsing harness-agnostic agent trajectories for model training and dreaming inference. There were no great existing libraries for parsing trajectories from various sources into a canonical token-efficient format for agent consumption (ie, by the dreamer), so we released what we've been building internally on GitHub and as an easy-to-use npm package. ~5x less tokens than raw jsonl, ready for agent consumption. Trajectory is in production inside of the Letta Code harness as the preprocessor for dreaming cycles. Enjoy, and open PRs if you run into issues!

English
2
1
3
442
Scott Condron
Scott Condron@_ScottCondron·
Maybe we should embrace the SaaSpocalypse and run a hackathon using CoreWeave inference to vibe code a clone of wandb and weave? Who’s in!?
English
0
0
3
336
Scott Condron retweetledi
CoreWeave
CoreWeave@CoreWeave·
Artificial Analysis benchmarked the speed of every provider serving @MiniMax_AI M3. We're the tall bar 😅 357 output tokens per second, 1.8x the next provider, plus the fastest time to first answer token at 6.6s. Blended price matches the lowest on the board at $0.22/M.
CoreWeave tweet media
English
13
22
207
46.3K
Scott Condron
Scott Condron@_ScottCondron·
@urieli17 let's put that in 2 - they might know what to optimize (happy users etc.) but unsure how to tie an eval to that - a bit of a mix of 2 & 3
English
0
0
1
16
Scott Condron
Scott Condron@_ScottCondron·
Why do some AI products not prioritize evals?
English
1
0
0
333
Scott Condron
Scott Condron@_ScottCondron·
Reasons: - prod / beta tester failures are cheap and they have enough signal from them to know what to improve - they’re not interested in expensive regression tests for the same reason as above - they don’t have PMF or their system isn’t stable enough to know what to optimize - it’s too much up front investment and they’re not sure where / how to start Any I missed?
English
0
0
0
93
Nicholas Charriere
Nicholas Charriere@nichochar·
when you talk to builders evals: - labs people: “we do it all day every day it’s the most important thing” - application layer agent builders: “im embarrassed to say we barely do them at all” there is a reason for this
English
1
0
0
347
Scott Condron
Scott Condron@_ScottCondron·
@johnowhitaker @hortonhearsafoo Say for an internal document where I transcribe it, or even just if I dump in the general point and have it do some rephrasing. Where do you draw the line?
English
0
0
1
23
William Horton
William Horton@hortonhearsafoo·
I don’t get stuff like this. The patterns are what makes AI writing easy to spot, but they’re not what _make_ it slop.
Peter Yang@petergyang

I’m sick of reading AI slop, so today I’m open-sourcing my /no-ai-slop skill that removes 20+ slop patterns from any piece of writing. 📌 Get the free skill here: github.com/petergyang/no-… If you find it useful, please consider starring the repo so more people can find it. Why I built the skill: I use AI to edit my writing because it helps me fix spelling, grammar, and clarity. But even the best models keep producing the same slop that this skill removes: → Binary contrasts: “It’s not X. It’s Y.” → Throat-clearing openers: “Here’s what nobody tells you.” → Fake-profound endings: “The future isn’t coming. It’s already here.” Use this skill responsibly. I always write a first draft manually before iterating with AI on edits and I make sure to do another manual pass at the end as well. That’s in contrast to using AI to automate pumping out slop end-to-end. 📌 Read my full post for more on how I try to use AI responsibly to edit without giving into the dark side: creatoreconomy.so/p/use-my-no-ai…

English
2
0
6
459
Scott Condron retweetledi
Weights & Biases
🗣️ Just ask ARIA! Why the loss spiked. Which run won. Whether a teammate already tried it. What to launch next. An AI research agent built into W&B, your project context already loaded. Public preview is live and in your workspace now.
English
1
5
14
1.9K
Scott Condron
Scott Condron@_ScottCondron·
@thdxr Spicy! I dunno, I find myself happy enough to sit in some products and talk to their agents - especially if they can take action in the product and control things visually. That said, I do often like having the MCP for most apps
English
0
0
0
259
dax
dax@thdxr·
the agent you're building into your product is most likely not going to work just drop it, what's going to happen is there will be a few agent products that individuals and teams choose to use and they will want everything there. might mean they don't even want your product
English
120
27
1.2K
111.3K
Scott Condron
Scott Condron@_ScottCondron·
@HamelHusain Evals have gone through so much in the last couple of years but you haven’t wavered!
English
0
0
0
514
Hamel Husain
Hamel Husain@HamelHusain·
It’s happening ✨
Hamel Husain tweet mediaHamel Husain tweet media
English
14
9
132
12.7K
Scott Condron retweetledi
Morgan McGuire
Morgan McGuire@morgymcg·
keep track of your /goal research studies
Morgan McGuire tweet media
English
0
1
2
765
Scott Condron
Scott Condron@_ScottCondron·
I miss doing dev rel. I’ve been so focused on the product side of things at W&B and growing Weave, which has been a great experience, but here’s hoping I can get back to making more educational content and building fun demos. I feel like I’ve seen so much cool stuff from customers, had ideas and haven’t found time to share it Things I’m excited about: - auto-research (the OpenResearch project looks slick, want to play more and make W&B better with external research agents) - automated insights on traces - lots of CoreWeave capacity we could put to use giving teams quick insights into production agents to make them more reliable - Marimo, especially using molab for shareable experimental UIs for evals - marimo-pair - cracked way to give an agent a way to control the runtime of your apps, kind of like computer use but can sit in between processes and change variables - easier fine-tuning / distillation using some of the tools from OpenPipe
Scott Condron@_ScottCondron

OpenResearch from @askalphaxiv runs experiments on your own code and compute Each experiment gets an isolated worktree, @wandb backed runs, and a graph showing how the research branches and progresses /reproduce-paper <paper URL or title> on <compute>

English
2
0
2
331
Scott Condron retweetledi
Suhail
Suhail@Suhail·
Now I can go to bed. Good luck little AI research agent. May your wandb metrics avoid collapse.
Suhail tweet media
English
7
2
48
14.7K