Voice AI testing for restaurant and QSR ordering agents
Simulate thousands of phone and drive-thru orders with modifiers and background noise, scoring each against your rules for readback accuracy and upsells.
No credit card · 200 free credits
Simulate
Simulate the orders your agent takes
Paste your agent prompt and RubricHQ generates happy paths, edge cases, and adversarial orders for each menu flow — then runs them as real voice calls.
Phone and drive-thru ordering
Full orders from greeting to total, including guests who change their mind or add items after the readback.
- Guest adds a drink after hearing the total
- Order for a group read out all at once
- Item requested that isn't on this location's menu
- Pickup time requested outside store hours
Combos and modifiers
Meal upgrades, substitutions, and special requests captured correctly, not just the base item.
- No onions, extra sauce, and a swapped side on one item
- Combo size change mid-order
- Allergy question about a menu item
- Kids' meal with a substituted toy or drink
Upsell and order readback
Suggesting add-ons once and gracefully, then reading back an order that matches what the guest said.
- Guest declines the upsell and the agent asks again
- Readback skips a modifier the guest requested
- Total quoted doesn't match the items ordered
- Guest interrupts the readback to correct an item
Outcome metrics
Measure what success means in restaurants & qsr
Every call gets scored, so these become rates you can track across every test batch and every production call.
Order accuracy
Items, sizes, and modifiers on the ticket match what the guest asked.
Upsell acceptance
Suggested add-ons the guest accepted, offered once and never pushed.
Order completion rate
Calls that end with a confirmed order sent to the kitchen.
Containment rate
Orders completed without a hand-off to crew.
Evaluate
What RubricHQ checks on every call
Each simulated or production call is scored with code-as-judge rules, LLM-as-judge metrics, and audio metrics — so a failure shows up as a failed metric, not a remade order.
Order accuracy
Did the agent capture every item and modifier, read them back correctly, and quote the right total?
- Every modifier present in the final readback
- Total matches menu prices for the items ordered
- Items and locations that don't exist are never confirmed
Guest experience and upsell
Upsells should add revenue, not friction. RubricHQ flags repeated pitches and allergy answers the agent shouldn't give.
- Upsell offered once and dropped after a no
- Allergy questions answered only from approved menu data
- Tone stays friendly with rushed or rude guests
Noisy, real-world audio
Guests order from cars, busy kitchens, and loud streets. Test for them before lunch rush.
- Road noise, background chatter, and bad mics
- Constant interrupters and fast talkers
- Latency, dead-air, and interruptions measured per turn
Metrics to score every restaurants & qsr call
Ready to run from the metric gallery
Custom metrics you can add
Callers to test against
Personas from the built-in library, with multi-language support.
From prompt to production in one platform
- 01
Simulate
Auto-generate scenarios from your prompt and run them as concurrent voice calls.
- 02
Evaluate
Score every call with code, LLM-as-judge, and audio metrics.
- 03
Optimize
Diagnose failures and prove prompt fixes before you ship.
- 04
Monitor
Score live production calls and get alerted in Slack or email when checks fail.
Restaurants & QSR voice AI testing FAQ
Can RubricHQ check that my agent reads back orders correctly?+
Yes, as a custom metric. You can add a check that every item and modifier the guest asked for appears in the readback, and a code-as-judge rule that the quoted total matches menu prices.
Can it simulate drive-thru noise?+
Yes. The persona library includes noisy-cafe, hands-free driving, bad-mic, and poor-connection callers, and audio metrics measure latency, dead-air, and interruptions on every turn.
Does it work with drive-thru agents, not just phone ordering?+
If your ordering agent is built on Vapi, Retell, LiveKit, or Pipecat, or answers a phone number, RubricHQ can run simulated calls against it over phone or web. Production orders from other systems can be sent in through the API for scoring.
Can it test upsell behavior?+
Yes. Add a custom metric for your upsell policy — for example, 'offer a combo upgrade once and drop it after a no' — and it is scored on every simulated and production call.
Can it test menu changes before they go live?+
Yes. Update your agent prompt, re-run the same scenario batch, and compare results. You can gate deploys in CI with the RubricHQ GitHub Action so a menu update that breaks ordering doesn't ship.
Can it test Spanish-language ordering?+
Yes. Simulations can run in multiple languages, and the persona library includes accented and non-native English callers.
Test your restaurants & qsr voice agent today
$0 to start — 200 free credits, no credit card. Connect your agent and run your first batch of test calls in minutes.
More industries