Test before launch
Simulate
Auto-generate scenarios & run calls
Evaluate
Metrics & transcript analysis
Optimize
AI-rewritten prompt improvements
Monitor in production
Monitor
Live production call observability
Co-Pilot
Ask your calls anything
IndustriesPricingAPI documentationAbout usBlogs
Book a demoLog inSign up
IndustriesRestaurants & QSR

Voice AI testing for restaurant and QSR ordering agents

Simulate thousands of phone and drive-thru orders with modifiers and background noise, scoring each against your rules for readback accuracy and upsells.

No credit card · 200 free credits

Works withVapiRetellLiveKitPipecatElevenLabsOpenAI

Simulate

Simulate the orders your agent takes

Paste your agent prompt and RubricHQ generates happy paths, edge cases, and adversarial orders for each menu flow — then runs them as real voice calls.

Phone and drive-thru ordering

Full orders from greeting to total, including guests who change their mind or add items after the readback.

  • Guest adds a drink after hearing the total
  • Order for a group read out all at once
  • Item requested that isn't on this location's menu
  • Pickup time requested outside store hours

Combos and modifiers

Meal upgrades, substitutions, and special requests captured correctly, not just the base item.

  • No onions, extra sauce, and a swapped side on one item
  • Combo size change mid-order
  • Allergy question about a menu item
  • Kids' meal with a substituted toy or drink

Upsell and order readback

Suggesting add-ons once and gracefully, then reading back an order that matches what the guest said.

  • Guest declines the upsell and the agent asks again
  • Readback skips a modifier the guest requested
  • Total quoted doesn't match the items ordered
  • Guest interrupts the readback to correct an item

Outcome metrics

Measure what success means in restaurants & qsr

Every call gets scored, so these become rates you can track across every test batch and every production call.

Order accuracy

Items, sizes, and modifiers on the ticket match what the guest asked.

Upsell acceptance

Suggested add-ons the guest accepted, offered once and never pushed.

Order completion rate

Calls that end with a confirmed order sent to the kitchen.

Containment rate

Orders completed without a hand-off to crew.

Evaluate

What RubricHQ checks on every call

Each simulated or production call is scored with code-as-judge rules, LLM-as-judge metrics, and audio metrics — so a failure shows up as a failed metric, not a remade order.

Order accuracy

Did the agent capture every item and modifier, read them back correctly, and quote the right total?

  • Every modifier present in the final readback
  • Total matches menu prices for the items ordered
  • Items and locations that don't exist are never confirmed

Guest experience and upsell

Upsells should add revenue, not friction. RubricHQ flags repeated pitches and allergy answers the agent shouldn't give.

  • Upsell offered once and dropped after a no
  • Allergy questions answered only from approved menu data
  • Tone stays friendly with rushed or rude guests

Noisy, real-world audio

Guests order from cars, busy kitchens, and loud streets. Test for them before lunch rush.

  • Road noise, background chatter, and bad mics
  • Constant interrupters and fast talkers
  • Latency, dead-air, and interruptions measured per turn

Metrics to score every restaurants & qsr call

Ready to run from the metric gallery

[General] Primary Issue Resolution[General] No Misrepresentation or False Claims[General] Professionalism & Language[General] Agent Empathy & Tone

Custom metrics you can add

+All modifiers captured in readback+Order total matches menu pricing+Upsell offered at most once+Allergy questions deferred to approved info
How evaluation works →

Callers to test against

Personas from the built-in library, with multi-language support.

NCNoisy Cafe CallerDHDriving Hands-FreeCIConstant InterrupterFTFast TalkerIDIndecisive DithererGBGarbled Bad Mic
How simulation works →

From prompt to production in one platform

  1. 01

    Simulate

    Auto-generate scenarios from your prompt and run them as concurrent voice calls.

  2. 02

    Evaluate

    Score every call with code, LLM-as-judge, and audio metrics.

  3. 03

    Optimize

    Diagnose failures and prove prompt fixes before you ship.

  4. 04

    Monitor

    Score live production calls and get alerted in Slack or email when checks fail.

Restaurants & QSR voice AI testing FAQ

Can RubricHQ check that my agent reads back orders correctly?+

Yes, as a custom metric. You can add a check that every item and modifier the guest asked for appears in the readback, and a code-as-judge rule that the quoted total matches menu prices.

Can it simulate drive-thru noise?+

Yes. The persona library includes noisy-cafe, hands-free driving, bad-mic, and poor-connection callers, and audio metrics measure latency, dead-air, and interruptions on every turn.

Does it work with drive-thru agents, not just phone ordering?+

If your ordering agent is built on Vapi, Retell, LiveKit, or Pipecat, or answers a phone number, RubricHQ can run simulated calls against it over phone or web. Production orders from other systems can be sent in through the API for scoring.

Can it test upsell behavior?+

Yes. Add a custom metric for your upsell policy — for example, 'offer a combo upgrade once and drop it after a no' — and it is scored on every simulated and production call.

Can it test menu changes before they go live?+

Yes. Update your agent prompt, re-run the same scenario batch, and compare results. You can gate deploys in CI with the RubricHQ GitHub Action so a menu update that breaks ordering doesn't ship.

Can it test Spanish-language ordering?+

Yes. Simulations can run in multiple languages, and the persona library includes accented and non-native English callers.

Test your restaurants & qsr voice agent today

$0 to start — 200 free credits, no credit card. Connect your agent and run your first batch of test calls in minutes.

More industries