joseph
110 posts

joseph
@gitcommit90
nano bio.md director of @publicbytesorg

box is the cheapest full VM sandbox for agents each box gets 4vCPU, 8gb ram, >50gb storage, free disk snapshots, desktop, ssh, https hosting, devtools, docker run >1000 concurrent boxes, self-serve at only $0.00001/s/box imagine all the agents you could run! try box









i didn't expect the american open weight comeback to actually deliver the thing builders kept asking for. for a year now "open weights" has meant chinese labs, full stop, while the american side went quiet, and every time the west did drop something "open" it was late, or a toy, or it needed a rack of h100s to breathe. builders kept asking for one boring thing the whole time: an open model actually good enough to matter, that i can run myself. on hardware i own. poolside's laguna s 2.1 is the first american drop that just answers it, and i didn't take the launch chart's word for it, i put the whole thing on one dgx spark and measured every number. it's a 118b mixture of experts, 8.5b active per token, open weights under a real license, and agentic coding is the entire point of it. i served it on vllm with the nvfp4 build and the dflash drafter, and the part nobody publishes is what it actually does on your own metal. it fits. 67 gigs of weights on ram, and because only 12 of its 48 layers run full attention, the entire million token context is about 26 gigs of kv, so the full window lands near 100 of 128 gigs with room to spare. a frontier coding model, its whole context, on one box that sits next to a monitor. speed is honest anon. single stream it's modest, about 19 tokens a second, this box was never a sprinter. but the dflash drafter climbs it to 25-30 on normal chat and 45 sustained on code, the workload it's built for, and under load it does 140 tokens a second across 16 streams, because the spark was always a serving box, not a single-stream drag race. it even holds at depth, barely fading from 19 to 14 out at 230k of context while free memory never moves. that's the comeback that actually matters. not a press release, not a leaderboard screenshot, an open model that shows up with what people kept asking for and holds up when you measure it but still america didn't win open source back, china still sets the pace. but this is the first one that landed the ask.









Today we're releasing Laguna S 2.1, our most capable model to date. It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes. Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark. Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface poolside.ai/blog/introduci…











