actAVA AI

49 posts

actAVA AI banner
actAVA AI

actAVA AI

@actAVAai

actAVA is the premier agent lifecycle platform, purpose-built for the healthcare and life sciences industries.

Healthcare AI Factory Katılım Eylül 2025
5 Takip Edilen187 Takipçiler
actAVA AI retweetledi
Haolin Chen
Haolin Chen@HaolinChen11·
Today, we’re adding Inkling to the χ-Bench leaderboard. Inkling passed 6 of 75 tasks, while 58 of 69 failures were workflow or tool-use errors. • 8.0% PASS@1 • $69.26 per full run • $11.54 per passed task Using the χ-Bench failure taxonomy, the 69 unsuccessful trials broke down as: • 36 workflow-completion failures • 22 tool-use errors • 8 clinical-reasoning failures • 3 stuck or abstained Take a deeper look: 20 out of 22 tool-use errors are caused by model calling a wrong tool alias that is not provided to it, given it has sufficient tools and could use them to finish the task. We are also evaluating Opus 5 which is released in this morning, stay tuned! #ChiBench #AIAgents #LLMEvaluation #HealthcareAI #AIEngineering
Haolin Chen tweet media
English
0
1
2
37
actAVA AI retweetledi
Dhairya Shah
Dhairya Shah@dhairya_dhiru99·
Exciting to see domain-specific AI like Cura 1T outperforming frontier models on HealthBench Hard & Professional via recursive self-improvement. Trained smarter, not just bigger—delivering frontier performance at 20-100x lower inference cost. Promising path for scalable, deployable medical AI. @actAVAai @AI_MDPI @SwissCognitive @UofT_TCAIREM
English
1
3
6
304
actAVA AI retweetledi
Weiran Yao
Weiran Yao@iscreamnearby·
Posted the launch yesterday. Today, the part people kept asking about – This is Cura working through a diagnostic case, unedited: "A 26-year-old man falls from a ladder, landing on his outstretched right hand". Clinicians: pls reply with a case; we'll run it and post the trace.
Weiran Yao@iscreamnearby

The strongest healthcare LLM, custom-built for your enterprise, owned by you🌸 Meet @actAVAai Cura: 1T agentic model trained by recursive self-improvement for long clinical + health admin workflows Try: actava.ai/cura Share your use case: $20 credits + early access👇

English
8
8
26
3.7K
actAVA AI retweetledi
actAVA AI
actAVA AI@actAVAai·
Today we launched Cura, a 1-trillion-parameter model trained for clinical and administrative healthcare work. With it, actAVA becomes the first U.S. startup to help healthcare organizations customize and own a model at this scale. Earlier this year we ran χ-BENCH, a study of real healthcare administrative tasks. The best frontier agent we tested, Claude Opus 4.8 with Claude Code, completed 33% of the work. Researchers at Stanford, Johns Hopkins, Yale, and Salesforce AI Research reviewed the study. 33% is not a prompting problem. You don't fix it with better retrieval or a longer system prompt. So we trained our own model, on the actual work these organizations do. Across six of the hardest healthcare evaluations, CURA beats Claude Opus and matches Claude Fable, the leading frontier models, at 20 to 100 times lower inference cost. It runs standalone or on KORA, our agent platform, and deploys inside the customer's own VPC. Patient data never leaves their control. The part we keep coming back to is ownership. Right now, most healthcare organizations rent their intelligence: per-seat tools charge a markup on every model call, and the customer's operational knowledge quietly improves someone else's product. Owning the model flips that. Your data, your standards, your infrastructure, and the model gets better on your work instead of a vendor's. Healthcare has the data, the domain depth, and the economic pressure to be the first industry that runs on models it controls. That's the bet. More at actava.ai/cura
actAVA AI tweet media
Weiran Yao@iscreamnearby

The strongest healthcare LLM, custom-built for your enterprise, owned by you🌸 Meet @actAVAai Cura: 1T agentic model trained by recursive self-improvement for long clinical + health admin workflows Try: actava.ai/cura Share your use case: $20 credits + early access👇

English
73
5
83
3.2K
actAVA AI
actAVA AI@actAVAai·
Healthcare enterprises are increasingly vulnerable to turbulence in model vendors. When a provider updates their underlying model, performance drift often occurs—unseen—in your production workflows. For high-stakes healthcare operations, this variance is not an acceptable operational risk. Relying on a single vendor’s roadmap introduces "model lock-in," where your ability to execute is tied to the stability of an external API. To maintain operational continuity, organizations require a governance layer that abstracts the model from the business logic. Protecting high-stakes workflows from vendor volatility requires three structural guardrails: - Model-Agnostic Orchestration: Decouple your business logic from the underlying LLM. By using a modular framework, you can swap models without re-architecting your entire agent infrastructure. - Continuous Stress Testing: Run automated evaluations (using χ-BENCH) every time a model provider pushes an update. You need to verify performance against your specific clinical and operational benchmarks before that update touches your data. - Human-in-the-Loop Validation: Automate the detection of performance degradation. When an agent’s output drifts outside of established thresholds, the system must trigger an immediate human review to maintain clinical safety. Enterprise healthcare requires predictable, governed outcomes that are immune to third-party volatility. actAVA provides the infrastructure to isolate your workflows from external model changes. Blog: actava.ai/news/how-to-pr… #actAVA #AIGovernance #EnterpriseAI #AgentLifecycle
English
24
1
38
522
actAVA AI
actAVA AI@actAVAai·
𝗧𝗵𝗲 𝗯𝗲𝘀𝘁 𝗳𝗿𝗼𝗻𝘁𝗶𝗲𝗿 𝗮𝗴𝗲𝗻𝘁𝘀 𝗿𝗲𝘀𝗼𝗹𝘃𝗲 𝗼𝗻𝗹𝘆 𝟮𝟴% 𝗼𝗳 𝗿𝗲𝗮𝗹 𝗵𝗲𝗮𝗹𝘁𝗵𝗰𝗮𝗿𝗲 𝘄𝗼𝗿𝗸𝗳𝗹𝗼𝘄𝘀 𝗼𝗻 𝘁𝗵𝗲 𝗳𝗶𝗿𝘀𝘁 𝘁𝗿𝘆. 𝗪𝗵𝘆 𝗶𝘀 𝘁𝗵𝗲𝗿𝗲 𝘀𝘂𝗰𝗵 𝗮 𝗺𝗮𝘀𝘀𝗶𝘃𝗲 𝗴𝗮𝗽 𝗯𝗲𝘁𝘄𝗲𝗲𝗻 𝗰𝗹𝗲𝗮𝗻 𝗮𝗰𝗮𝗱𝗲𝗺𝗶𝗰 𝗯𝗲𝗻𝗰𝗵𝗺𝗮𝗿𝗸𝘀 𝗮𝗻𝗱 𝗽𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝗼𝗻-𝗿𝗲𝗮𝗱𝘆 𝗽𝗲𝗿𝗳𝗼𝗿𝗺𝗮𝗻𝗰𝗲? General-purpose models often excel on isolated coding tests and polished datasets. However, real healthcare operations present a completely different level of friction. When moving from pilot to production, a healthcare-native agent lifecycle requires a rigorous validation standard tailored to operational realities. Our latest analysis breaks down why the industry requires a better testing architecture to measure agentic AI, highlighting the three critical vectors evaluated by actAVA.ai χ-BENCH: • 𝗣𝗼𝗹𝗶𝗰𝘆 𝗗𝗲𝗻𝘀𝗶𝘁𝘆: Agents must navigate massive, shifting libraries of medical guidelines, insurance rules, and provider procedures across long tool-call chains. • 𝗠𝘂𝗹𝘁𝗶-𝗥𝗼𝗹𝗲 𝗖𝗼𝗺𝗽𝗼𝘀𝗶𝘁𝗶𝗼𝗻: End-to-end workflows require systems to switch context seamlessly between distinct roles—such as a clinician, UM nurse, and medical director—without losing continuity. • 𝗜𝗿𝗿𝗲𝘃𝗲𝗿𝘀𝗶𝗯𝗹𝗲 𝗔𝗰𝘁𝗶𝗼𝗻𝘀: Operational handoffs are terminal. Unlike consumer applications, a step submitted to an enterprise workflow cannot be corrected through a simple trial-and-error loop. Evaluating systems against realistic operational stress tests is the only way to ensure safety, governance, and measurable ROI from day one. Read the full article from Dr. Weiran Yao, Chief AI Officer at actAVA.ai, to learn how we bring structured order to complex AI evaluation. lnkd.in/g8paiZaA The AI factory for healthcare. Master your agentic future. hashtag#actAVA hashtag#HealthcareAI hashtag#AgenticAI
English
13
1
15
291
actAVA AI
actAVA AI@actAVAai·
If you turn on Web Search in ChatGPT and ask it find healthcare agent benchmarks for you, CHI-Bench is the top recommended, strongest match for testing agent performances on realistic healthcare operations
actAVA AI tweet media
English
1
0
8
778
actAVA AI
actAVA AI@actAVAai·
𝗧𝗵𝗲 𝗾𝘂𝗲𝘀𝘁𝗶𝗼𝗻 𝗵𝗲𝗮𝗹𝘁𝗵𝗰𝗮𝗿𝗲 𝗔𝗜 𝘁𝗲𝗮𝗺𝘀 𝗮𝗿𝗲 𝗮𝘀𝗸𝗶𝗻𝗴 𝗶𝗻 𝟮𝟬𝟮𝟲 𝗶𝘀𝗻'𝘁 "𝗖𝗮𝗻 𝘄𝗲 𝗯𝘂𝗶𝗹𝗱 𝗮𝗻 𝗮𝗴𝗲𝗻𝘁?" 𝗜𝘁'𝘀 "𝗛𝗼𝘄 𝗱𝗼 𝘄𝗲 𝗸𝗻𝗼𝘄 𝗶𝘁'𝘀 𝘀𝗮𝗳𝗲 𝗳𝗼𝗿 𝗽𝗮𝘁𝗶𝗲𝗻𝘁𝘀?" Too many agents are deployed before they’re ready, leading to incidents that set automation back by years. Here is the framework for true operational readiness: • 𝗕𝗲𝗻𝗰𝗵𝗺𝗮𝗿𝗸𝘀 𝘃𝘀. 𝗥𝗲𝗮𝗹𝗶𝘁𝘆: A 90% accuracy score on a clinical Q&A benchmark does not predict success with complex, multi-step transactions such as prior authorization. • 𝗜𝗱𝗲𝗻𝘁𝗶𝗳𝘆 𝗙𝗮𝗶𝗹𝘂𝗿𝗲 𝗠𝗼𝗱𝗲𝘀: You must evaluate agents against systemic risks, including reasoning vulnerabilities, tool integration latencies, and escalation anomalies. • 𝗦𝗶𝗺𝘂𝗹𝗮𝘁𝗶𝗼𝗻 𝗼𝘃𝗲𝗿 𝗗𝗲𝗺𝗼𝘀: Stop relying on "happy path" vendor demos. Implement rigorous workflow simulations that test administrative edge cases, state persistence, and live integrations. • 𝗖𝗼𝗻𝘁𝗶𝗻𝘂𝗼𝘂𝘀 𝗚𝗼𝘃𝗲𝗿𝗻𝗮𝗻𝗰𝗲: Deployment is only the beginning. You must continuously monitor for environmental degradation, including regulatory shifts, model drift, and shifting patient cohort distributions. The "deploy-and-observe" era must end. Healthcare enterprises require a verifiable, reproducible foundation for risk mitigation to ensure compliance and patient safety. Frameworks like actAVA KORA and χ-BENCH provide the structure necessary to move from perpetual pilot to production-ready automation. Read more about this important topic from actAVA.ai Chief AI Officer, Dr. Weiran Yao, here --> actava.ai/news/the-5-dim…
English
3
1
5
293
actAVA AI
actAVA AI@actAVAai·
We welcome Tom Patterson to the actAVA AI advisory board. As the SVP of Corporate Development and Strategy at BetterUp and a 3-time founder, Tom brings deep operational expertise in aligning enterprise technology with human performance and workforce resilience. In this role, Tom advises us on integrating human-in-the-loop safeguards into the agent lifecycle. His focus centers on critical governance infrastructure, including: - Privacy Compliance: Structuring models to strictly adhere to enterprise privacy regulations. - User Agency: Implementing clear design frameworks that distinguish AI-generated advice from human guidance. - Workforce Stability: Deploying systems capable of monitoring operational distress cues to support frontline teams. Tom’s first question to us was “How do you scale production-ready agentic AI without introducing workforce instability or employee anxiety?” From Tom’s point of view, technology can only move from pilot to production when built with rigid structural safeguards that protect the human element. Building an unshakeable partnership between silicon and carbon requires evidence-led governance, not open-ended promises. We agree! We are grateful for Tom’s partnership as we continue to build the infrastructure for responsible, governed, and measurable agentic AI at enterprise scale. Read our full conversation with Tom on the structural safeguards required for enterprise growth: actava.ai/news/meet-our-… The AI factory for healthcare. Master your agentic future. #actAVA #AIGovernance #AgentLifecycle
English
0
0
2
70
actAVA AI retweetledi
Weiran Yao
Weiran Yao@iscreamnearby·
99% of AI researchers don’t know. Your HuggingFace datasets can be selected into benchmark shortlist if it tests on open source models. We just made CHI-Bench into the shortlist with other popular benchmarks like GSM8k, SWE-Bench🤗 RT or reply👇if you want to know how!
Weiran Yao tweet media
English
1
1
5
419
Weiran Yao
Weiran Yao@iscreamnearby·
Huge congrats to @Humana's Erius agent taking the #1 spot on CHI-Bench for Prior Auth and 6th for all domains. It outperforms every frontier lab on one of healthcare's hardest workflows.
Weiran Yao tweet media
English
2
5
9
771
actAVA AI
actAVA AI@actAVAai·
CHI-Bench leaderboard just gets updated with the newest and highest score from @claudeai Opus 4.8. CHI-Bench is world's first long-horizon benchmark for healthcare AI agents. Leaderboard: actava.ai/benchmarks
Weiran Yao@iscreamnearby

🚀 @claudeai Opus 4.8 just took #1 on CHI-Bench (long-horizon healthcare agents). 75 real workflows across prior auth, utilization & care management. Opus 4.8 → 33.3% (PA 32 · UM 28 · CM 40) Opus 4.6 → 28.0% (PA 18 · UM 41 · CM 24) Opus 4.7 → 24.4% (PA 24 · UM 17 · CM 32)

English
0
0
4
294