Voice AI testing for healthcare agents
Simulate thousands of patient calls and score every one for identity verification, PHI handling, and clinical scope — before you go live.
No credit card · 200 free credits
Simulate
Simulate the patient calls your agent handles
Paste your agent prompt and RubricHQ generates happy paths, edge cases, and adversarial calls for each workflow — then runs them as real voice calls.
Scheduling and reminders
New appointments, reschedules, and cancellations, including the calls where the patient changes their mind halfway through.
- Patient reschedules to a provider with no availability
- Caller books for a family member
- Cancellation inside the late-cancel window
- Reminder call reaches voicemail
Prescription refills
Refill requests routed to the right pharmacy and provider, with the agent staying inside what it is allowed to say.
- Refill for a medication not on file
- Caller asks whether they can double a dose
- Pharmacy change mid-request
- Controlled-substance refill that needs a human
Intake and triage routing
Collecting insurance and symptoms, and escalating anything urgent to staff instead of handling it.
- Caller describes chest pain during intake
- Insurance member ID read out with errors
- Patient refuses to give date of birth
- Spanish-speaking caller on an English line
Outcome metrics
Measure what success means in healthcare
Every call gets scored, so these become rates you can track across every test batch and every production call.
Resolution rate
Share of calls where the patient's request was fully handled.
Booking accuracy
Appointments booked or changed with date, time, and provider confirmed back to the patient.
Containment rate
Calls completed without a transfer to front-desk staff.
Escalation accuracy
Urgent symptoms routed to a human every single time.
Evaluate
What RubricHQ checks on every call
Each simulated or production call is scored with code-as-judge rules, LLM-as-judge metrics, and audio metrics — so a failure shows up as a failed metric, not a patient complaint.
Identity verification and PHI
Did the agent verify the caller before sharing anything, and share only what the request needed?
- Verification completed before any PHI is spoken
- No details disclosed to a wrong-party caller
- Minimum-necessary disclosure on every answer
Clinical scope boundaries
Scheduling agents shouldn't give medical advice. RubricHQ flags the calls where yours did.
- Dosage and symptom questions deflected to a clinician
- Emergency language triggers an immediate escalation
- No diagnosis or reassurance outside policy
Real caller conditions
Patients call from cars, waiting rooms, and bad connections. Test for them before launch.
- Elderly and hard-of-hearing callers
- Background noise and dropped audio
- Latency, dead-air, and interruptions measured per turn
Metrics to score every healthcare call
Ready to run from the metric gallery
Custom metrics you can add
Callers to test against
Personas from the built-in library, with multi-language support.
From prompt to production in one platform
- 01
Simulate
Auto-generate scenarios from your prompt and run them as concurrent voice calls.
- 02
Evaluate
Score every call with code, LLM-as-judge, and audio metrics.
- 03
Optimize
Diagnose failures and prove prompt fixes before you ship.
- 04
Monitor
Score live production calls and get alerted in Slack or email when checks fail.
Healthcare voice AI testing FAQ
Can RubricHQ test whether my agent leaks patient information?+
Yes. Prebuilt metrics check whether identity was verified before any sensitive detail was shared, whether details reached a wrong-party caller, and whether each answer disclosed only the minimum necessary. You can add custom checks for your own policy.
Is RubricHQ HIPAA certified?+
No. RubricHQ does not currently hold a HIPAA attestation or sign BAAs. We recommend testing with synthetic patient data, which is how simulations work by default — the simulated callers use made-up identities. Don't send production calls that contain PHI unless your agreements allow it.
Which voice platforms does it work with?+
RubricHQ connects to agents built on Vapi, Retell, LiveKit, and Pipecat, and can call any agent with a phone number. Production calls from other platforms can be sent in through the API for scoring.
How long does a test run take?+
Calls run in parallel, so a batch finishes far faster than dialing by hand; concurrency depends on your plan. You can also schedule runs nightly or gate deploys in CI with the RubricHQ GitHub Action.
Can it simulate elderly callers or callers with accents?+
Yes. The persona library includes confused senior, soft-spoken, non-native English, and poor-connection callers, and simulations can run in multiple languages.
Does it check that urgent symptoms get escalated?+
Yes. Add an escalation metric once — for example, 'any mention of chest pain is transferred to staff' — and it is scored on every simulated and production call.
Test your healthcare voice agent today
$0 to start — 200 free credits, no credit card. Connect your agent and run your first batch of test calls in minutes.
More industries