Voice AI testing for retail and e-commerce agents
Simulate thousands of shopper calls, scoring each against your rules for order lookups, return eligibility, refund amounts, and safe card handling.
No credit card · 200 free credits
Simulate
Simulate the shopper calls your agent handles
Paste your agent prompt and RubricHQ generates happy paths, edge cases, and adversarial calls for each workflow — then runs them as real voice calls.
Order status and delivery
Where-is-my-order calls, delivery changes, and missing packages, including orders that don't match the caller's account.
- Tracking shows delivered but the package is missing
- Caller gives an order number with one digit wrong
- Address change after the order has shipped
- Split shipment where only half arrived
Returns, exchanges, and refunds
Return eligibility, exchange options, and refund timing, with the agent applying your policy instead of improvising one.
- Return requested one day outside the window
- Refund to a gift card versus the original payment method
- Exchange for a size that is out of stock
- Caller insists an item arrived damaged
Pricing, promos, and payments
Price-match requests, promo codes, and payment updates, where accuracy matters more than friendliness.
- Promo code that expired yesterday
- Caller asks for a price adjustment after a sale starts
- Two discounts that can't be combined
- Customer starts reading a full card number aloud
Outcome metrics
Measure what success means in retail & e-commerce
Every call gets scored, so these become rates you can track across every test batch and every production call.
Containment rate
Calls completed without a transfer to a store associate.
Order-status resolution
Where-is-my-order calls answered with the correct carrier status.
Return completion rate
Eligible returns started with a label issued on the same call.
Refund accuracy
Refund amounts quoted match the order and the return policy.
Evaluate
What RubricHQ checks on every call
Each simulated or production call is scored with code-as-judge rules, LLM-as-judge metrics, and audio metrics — so a failure shows up as a failed metric, not a chargeback or a one-star review.
Order and account security
Did the agent verify the caller before discussing an order, and handle payment details the way your policy requires?
- Verification completed before order or address details are shared
- Full card numbers never repeated back
- Payment updates routed to a secure flow
Policy and pricing accuracy
Invented discounts and wrong refund amounts cost money. RubricHQ flags the calls where your agent made one up.
- Refund amount and timing match your policy
- No promos or price matches promised outside the rules
- Return eligibility applied correctly at the boundary
Real caller conditions
Shoppers call frustrated, distracted, and on the move. Test for them before peak season.
- Angry callers demanding a manager
- Impatient multitaskers and poor mobile connections
- Latency, dead-air, and interruptions measured per turn
Metrics to score every retail & e-commerce call
Ready to run from the metric gallery
Custom metrics you can add
Callers to test against
Personas from the built-in library, with multi-language support.
From prompt to production in one platform
- 01
Simulate
Auto-generate scenarios from your prompt and run them as concurrent voice calls.
- 02
Evaluate
Score every call with code, LLM-as-judge, and audio metrics.
- 03
Optimize
Diagnose failures and prove prompt fixes before you ship.
- 04
Monitor
Score live production calls and get alerted in Slack or email when checks fail.
Retail & E-commerce voice AI testing FAQ
Can RubricHQ check that my agent handles card details safely?+
Yes, as a custom metric. You can add a code-as-judge rule that fails any call where a full card number is repeated back, alongside the prebuilt identity-verification metrics that check the caller was verified before order details were shared.
Is RubricHQ PCI compliant?+
No. RubricHQ does not hold a PCI DSS attestation. Simulations use made-up shopper identities by default, and we recommend keeping real payment data out of test runs.
Can it test promo and refund accuracy?+
Yes. Scenarios are generated from your agent prompt, so they cover expired codes, stacked discounts, and out-of-window returns. A prebuilt metric flags false claims, and custom metrics check refund amounts against your policy.
Which voice platforms does it work with?+
RubricHQ connects to agents built on Vapi, Retell, LiveKit, and Pipecat, and can call any agent with a phone number. Production calls from other platforms can be sent in through the API for scoring.
Can it monitor production calls during peak season?+
Yes. Send production calls in through the API and they are scored with the same metrics as your test runs. Alert rules notify Slack, email, or a webhook when production calls fail a metric you choose. You can also schedule nightly regression runs.
How long does a test run take?+
Calls run in parallel, so a batch finishes far faster than dialing by hand; concurrency depends on your plan. You can also gate deploys in CI with the RubricHQ GitHub Action.
Test your retail & e-commerce voice agent today
$0 to start — 200 free credits, no credit card. Connect your agent and run your first batch of test calls in minutes.
More industries