Ryan Marten
645 posts

Ryan Marten
@ryan_marten
Building @harborframework and @terminalbench

Over the past year I've been building out Preseen. We’re focused on building the best AI forecasting agent for decision makers. Every decision relies on a forecast. Too often decisions are poorly made because the underlying forecast is bad. Humans are constrained. But if you can have a system that can forecast anything as well as the best humans, then human decision making will be improved. AIs are reaching the point where they are better than some of the best human forecasters. This is going to change decision making. Preseen will help drive that change.





We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%


We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%

We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%





