Voice AI testing for automotive dealership agents
Simulate thousands of shopper and service calls and score every one for inventory accuracy, APR and fee disclosures, and pressure-free sales conduct.
No credit card · 200 free credits
Simulate
Simulate the calls your dealership agent handles
Paste your agent prompt and RubricHQ generates happy paths, edge cases, and adversarial calls for each workflow — then runs them as real voice calls.
Sales, inventory, and test drives
Vehicle availability, trim and feature questions, and test-drive bookings, including the shopper who has already priced the car elsewhere.
- Shopper asks about a vehicle that already sold
- Caller wants the out-the-door price over the phone
- Test drive booked for a trim the lot doesn't have
- Trade-in value question before any appraisal
Service scheduling
Oil changes, recalls, and repair appointments booked into the right slot, with the right advisor and loaner expectations.
- Owner asks whether a recall applies to their VIN
- Reschedule when the only open slot has no loaner
- Caller describes a warning light and asks if it's safe to drive
- Two vehicles booked on one call
Parts and warranty
Parts availability, pricing, and warranty coverage questions, where the agent should confirm rather than guess.
- Caller asks if a repair is covered under warranty
- Part is on backorder with no firm date
- Aftermarket versus OEM price comparison
- Extended-warranty upsell the caller declines
Outcome metrics
Measure what success means in automotive
Every call gets scored, so these become rates you can track across every test batch and every production call.
Booking accuracy
Test drives and service visits booked with date, time, and vehicle confirmed back to the caller.
Lead capture rate
Shopper name, number, and vehicle interest captured before the call ends.
Containment rate
Calls completed without a transfer to the sales or service desk.
Pricing accuracy rate
Prices, fees, and APRs quoted match the dealer's approved figures.
Evaluate
What RubricHQ checks on every call
Each simulated or production call is scored with code-as-judge rules, LLM-as-judge metrics, and audio metrics — so a failure shows up as a failed metric, not a lost sale or a complaint.
Pricing and financing accuracy
Did the agent quote prices, fees, and financing terms that match your rules, and disclose what it was required to?
- APR and term quoted only as allowed by your script
- Doc, dealer, and add-on fees disclosed when price comes up
- No invented incentives or rebates
Sales conduct
Shoppers hang up on pushy agents. RubricHQ flags false urgency and misleading claims before your customers do.
- No fake scarcity or 'today only' pressure
- No false claims about vehicle history or availability
- A clear 'no' is respected on add-ons
Real caller conditions
Owners call from the road and shoppers call from noisy lots. Test for them before launch.
- Hands-free callers with road noise
- Fast talkers and constant interrupters
- Latency, dead-air, and interruptions measured per turn
Metrics to score every automotive call
Ready to run from the metric gallery
Custom metrics you can add
Callers to test against
Personas from the built-in library, with multi-language support.
From prompt to production in one platform
- 01
Simulate
Auto-generate scenarios from your prompt and run them as concurrent voice calls.
- 02
Evaluate
Score every call with code, LLM-as-judge, and audio metrics.
- 03
Optimize
Diagnose failures and prove prompt fixes before you ship.
- 04
Monitor
Score live production calls and get alerted in Slack or email when checks fail.
Automotive voice AI testing FAQ
Can RubricHQ check that my agent quotes financing correctly?+
Yes. A prebuilt metric checks that APR, term, and payment were stated clearly and consistently, with no bait-and-switch, and another checks fee transparency when price comes up. Add custom checks for your script's required disclosures. You can add custom checks for your own rate and incentive rules.
Does it catch high-pressure sales tactics?+
Yes. The prebuilt high-pressure tactics metric scores how much false urgency or pushing past a caller's 'no' the agent used, and a misrepresentation metric flags claims the agent can't back up.
Which voice platforms does it work with?+
RubricHQ connects to agents built on Vapi, Retell, LiveKit, and Pipecat, and can call any agent with a phone number. Production calls from other platforms can be sent in through the API for scoring.
Can it test service-lane scheduling?+
Yes. Scenarios are generated from your agent prompt, so they cover reschedules, recall questions, loaner requests, and multi-vehicle bookings, and a scheduling-accuracy metric checks that the date, time, and service were stated clearly and confirmed with the caller.
What happens when a test fails?+
RubricHQ diagnoses the failing calls and suggests a rewritten prompt. For Vapi and Retell agents, you can push the new prompt directly, then re-run the same scenarios to confirm the fix.
Can I try it without talking to sales?+
Yes. It is free to start with 200 credits and no card required. Connect your agent, paste its prompt, and run your first batch of simulated calls.
Test your automotive voice agent today
$0 to start — 200 free credits, no credit card. Connect your agent and run your first batch of test calls in minutes.
More industries