Test before launch
Simulate
Auto-generate scenarios & run calls
Evaluate
Metrics & transcript analysis
Optimize
AI-rewritten prompt improvements
Monitor in production
Monitor
Live production call observability
Co-Pilot
Ask your calls anything
IndustriesPricingAPI documentationAbout usBlogs
Book a demoLog inSign up

Voice AI testing for law firm intake agents

Simulate thousands of prospective client calls, scoring each against your rules for legal-advice boundaries, conflict-check intake, and consultation booking.

No credit card · 200 free credits

Works withVapiRetellLiveKitPipecatElevenLabsOpenAI

Simulate

Simulate the client calls your agent handles

Paste your agent prompt and RubricHQ generates happy paths, edge cases, and adversarial calls for each workflow — then runs them as real voice calls.

New-client intake

Capturing the matter type, key dates, and contact details from callers who often start mid-story.

  • Caller describes an accident out of order
  • Prospect calls about a matter the firm doesn't handle
  • Caller mentions a filing deadline this week
  • Family member calls on behalf of the client

Conflict-check information

Collecting the names of opposing and related parties so the firm can run conflicts before anyone discusses the case.

  • Caller doesn't know the opposing party's full name
  • Opposing party is a company with several names
  • Caller refuses to name the other side
  • Caller starts sharing case details before conflicts are run

Consultation scheduling

Booking consultations with the right attorney or practice group, and handling reschedules and fee questions.

  • Caller wants a consultation today
  • Prospect asks whether the consultation is free
  • Reschedule to an attorney with no availability
  • Caller asks for a specific attorney by name

Outcome metrics

Measure what success means in legal

Every call gets scored, so these become rates you can track across every test batch and every production call.

Intake completion rate

Matter type, contact details, and opposing parties captured on every call.

Consultation booking accuracy

Consultations booked with date, time, and attorney confirmed back to the caller.

Containment rate

Calls completed without a transfer to the front desk or a paralegal.

Advice boundary pass rate

No answer a caller could take as legal advice.

Evaluate

What RubricHQ checks on every call

Each simulated or production call is scored with code-as-judge rules, LLM-as-judge metrics, and audio metrics — so a failure shows up as a failed metric, not a missed client.

No legal advice

Intake agents shouldn't tell callers whether they have a case. RubricHQ flags the calls where yours did.

  • Case-strength questions deflected to an attorney
  • No predicted outcomes or settlement values
  • No promises that the firm will take the matter

Complete, accurate intake

Did the agent capture every party and date the firm needs, and book the right consultation?

  • Opposing and related parties captured for conflicts
  • Urgent deadlines flagged to staff
  • Consultation booked with the right attorney and time
  • Details read back before the call ends

Real caller conditions

Prospective clients call upset, rambling, or from a bad line. Test for them before launch.

  • Distressed and long-winded callers
  • Background noise and dropped audio
  • Latency, dead-air, and interruptions measured per turn

Metrics to score every legal call

Ready to run from the metric gallery

[General] Appointment / Scheduling Accuracy[General] No Misrepresentation or False Claims[General] Primary Issue Resolution[General] Agent Empathy & Tone[General] Professionalism & Language

Custom metrics you can add

+No legal advice given+Opposing parties captured for conflict check+Urgent deadlines escalated to staff
How evaluation works →

Callers to test against

Personas from the built-in library, with multi-language support.

CRChatty RamblerAWAnxious WorrierAEAngry EscalatorIDIndecisive DithererNENon-Native ESL Caller
How simulation works →

From prompt to production in one platform

  1. 01

    Simulate

    Auto-generate scenarios from your prompt and run them as concurrent voice calls.

  2. 02

    Evaluate

    Score every call with code, LLM-as-judge, and audio metrics.

  3. 03

    Optimize

    Diagnose failures and prove prompt fixes before you ship.

  4. 04

    Monitor

    Score live production calls and get alerted in Slack or email when checks fail.

Legal voice AI testing FAQ

Can RubricHQ check that my intake agent doesn't give legal advice?+

Yes. Add a custom metric that defines your firm's advice boundary — for example, 'never tells the caller whether they have a case' — and it is scored on every simulated and production call. Scenarios include callers who push hard for an answer.

Can it check that conflict-check details were captured?+

Yes. Add a custom metric listing the parties your conflicts process needs, and RubricHQ flags every call where the agent booked a consultation without them.

Is RubricHQ SOC 2 certified?+

No. RubricHQ does not currently hold a SOC 2 report. We recommend testing with synthetic client data, which is how simulations work by default — the simulated callers use made-up names and matters.

Which voice platforms does it work with?+

RubricHQ connects to agents built on Vapi, Retell, LiveKit, and Pipecat, and can call any agent with a phone number. Production calls from other platforms can be sent in through the API for scoring.

Can it simulate callers who ramble or get emotional?+

Yes. The persona library includes chatty, anxious, angry, and indecisive callers, and simulations can run in multiple languages.

How much does it cost to try?+

It's free to start: you get 200 credits with no card required.

Test your legal voice agent today

$0 to start — 200 free credits, no credit card. Connect your agent and run your first batch of test calls in minutes.

More industries