Voice AI testing for public sector agents
Simulate thousands of resident calls across languages, scoring each against your rules for eligibility accuracy, status verification, and appointment booking.
No credit card · 200 free credits
Simulate
Simulate the resident calls your agent handles
Paste your agent prompt and RubricHQ generates happy paths, edge cases, and adversarial calls for each workflow — then runs them as real voice calls.
311 and citizen services
Service requests, general questions, and routing to the right department, including the calls that don't fit any category.
- Pothole report with a vague cross-street location
- Noise complaint that should go to a non-emergency line
- Caller describes an emergency on the 311 line
- Question about a service another agency handles
Benefits and permit status
Application status and next steps shared only with the applicant, and answered from the record instead of guessed.
- Applicant asks why their benefit payment is late
- Caller checks a permit on behalf of a contractor
- Case number read out with errors
- Caller asks whether they qualify for a program
Appointments and multilingual access
Booking in-person visits and serving residents in their own language, at their own pace.
- DMV-style appointment rescheduled to a full day
- Spanish-speaking caller on an English line
- Caller needs documents listed before the visit
- Hard-of-hearing caller asks the agent to repeat
Outcome metrics
Measure what success means in public sector
Every call gets scored, so these become rates you can track across every test batch and every production call.
Containment rate
Resident requests completed without a transfer to city staff.
Resolution rate
Share of calls where the resident's request was fully handled.
Verification pass rate
Identity verified before any benefit or permit status is shared.
Eligibility accuracy
Eligibility answers match published program rules, never guessed.
Evaluate
What RubricHQ checks on every call
Each simulated or production call is scored with code-as-judge rules, LLM-as-judge metrics, and audio metrics — so a failure shows up as a failed metric, not a resident complaint.
Verification and disclosure
Did the agent verify the caller before sharing case details, and share only what the request needed?
- Applicant verified before any status is shared
- No case details disclosed to a third party
- Verification method matches agency policy
Accurate answers and routing
Eligibility and process answers should match published rules, and emergencies should never stay on the line.
- No eligibility decision stated outside policy
- Emergencies directed to 911 immediately
- Requests routed to the correct department
Access for every resident
Residents call in many languages and conditions. Test for them before launch.
- Callers in multiple languages
- Elderly, soft-spoken, and long-pause callers
- Latency, dead-air, and interruptions measured per turn
Metrics to score every public sector call
Ready to run from the metric gallery
Custom metrics you can add
Callers to test against
Personas from the built-in library, with multi-language support.
From prompt to production in one platform
- 01
Simulate
Auto-generate scenarios from your prompt and run them as concurrent voice calls.
- 02
Evaluate
Score every call with code, LLM-as-judge, and audio metrics.
- 03
Optimize
Diagnose failures and prove prompt fixes before you ship.
- 04
Monitor
Score live production calls and get alerted in Slack or email when checks fail.
Public Sector voice AI testing FAQ
Can RubricHQ test my agent in multiple languages?+
Yes. Simulations can run in multiple languages, and you can add a custom metric that checks the agent answered in the caller's language on every call.
Can it check that case details stay private?+
Yes. Prebuilt metrics check whether the caller was verified before any sensitive detail was shared and which verification method was used. You can add custom checks for your agency's own disclosure rules.
Is RubricHQ FedRAMP authorized or SOC 2 certified?+
No. RubricHQ does not hold a FedRAMP authorization or a SOC 2 report. We recommend testing with synthetic resident data, which is how simulations work by default — the simulated callers use made-up identities.
Can it test callers who need more time or have trouble hearing?+
Yes. The persona library includes confused senior, long-pause, soft-spoken, and tech-novice callers, and audio metrics measure whether the agent cut them off or left dead air.
Which voice platforms does it work with?+
RubricHQ connects to agents built on Vapi, Retell, LiveKit, and Pipecat, and can call any agent with a phone number. Production calls from other platforms can be sent in through the API for scoring.
How do we keep the agent reliable after launch?+
Schedule test runs nightly, gate prompt changes in CI with the RubricHQ GitHub Action, and send production calls in through the API. Alert rules notify Slack, email, or a webhook when production calls fail a metric you choose.
Test your public sector voice agent today
$0 to start — 200 free credits, no credit card. Connect your agent and run your first batch of test calls in minutes.
More industries