

gNucleus AI
125 posts

@gNucleusAI
Accelerate Engineering AI Transformation From engineering data labeling, proprietary AI model training, to managed and private cloud deployment!







Frontier-Bench was made possible by our sponsors. Thank you to our compute sponsors @modal_labs, @AnthropicAI, @OpenAI, @Google and our data partners @scale_AI, @SnorkelAI Open Benchmarks, @turingcom, @gNucleusAI, Boolean AI, and @joinHandshake. Frontier-Bench is hosted by @harborframework and @LaudeInstitute.

The featured task is : frontier-bench/freecad-platform-drawing It is one of three complex FreeCAD evaluations contributed by gNucleus AI to Frontier-Bench v0.1: • freecad-platform-drawing — an image-to-CAD reconstruction challenge • freecad-impeller — a complex text-to-CAD modeling task • freecad-spring-clip — another rigorous text-to-CAD evaluation @claudeai @alexgshaw @terminalbench @harborframework



We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%



Introducing Claude Opus 5. It's a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price.

🚀 Huge milestone for gNucleus and the entire generative engineering space! As a contributor to Frontier-Bench, we are excited to have our CAD evaluation tasks was featured in Anthropic’s announcement of Claude Opus 5. As our CTO Mei Chen puts it: Creating meaningful evaluations for AI in engineering is difficult. The tasks must reflect the geometric reasoning, tool use, iteration, and real-world friction involved in professional CAD workflows. Congrats to Ryan Marten Alex Shaw and Anthropic team! @claudeai Check out our post for the full backstory! 👇 linkedin.com/feed/update/ur… #IndustrialAI #CAD #LLM #AI #gNucleus AI





Introducing CADGenBench: measure how well AI systems produce engineering-grade 3D parts! While current models can generate 3D parts, they are far from precise enough to build functional parts. We built a benchmark to systematically measure their capabilities on two tasks: 1. Generation from an engineering drawing of a part 2. Editing: given an existing STEP file and a requested change The benchmark is tool-agnostic. It makes no assumptions about how you build the model. You can vary the LLM, and you can vary the environment. Use build123d, Onshape, Autodesk, or a model without an LLM entirely. We open sourced the scoring engine and a reference baseline on top of build123d. A collaboration between Hugging Face and @mecadoinc! Submission space: huggingface.co/spaces/Hugging… Code repository: github.com/huggingface/ca…
