Test before launch
Simulate
Auto-generate scenarios & run calls
Evaluate
Metrics & transcript analysis
Optimize
AI-rewritten prompt improvements
Monitor in production
Monitor
Live production call observability
Co-Pilot
Ask your calls anything
IndustriesPricingAPI documentationAbout usBlogs
Book a demoLog inSign up
IndustriesTravel & Hospitality

Voice AI testing for travel and hospitality agents

Simulate thousands of guest calls and score every one for booking accuracy, fee transparency, and cancellation policy — before you go live.

No credit card · 200 free credits

Works withVapiRetellLiveKitPipecatElevenLabsOpenAI

Simulate

Simulate the guest calls your agent handles

Paste your agent prompt and RubricHQ generates happy paths, edge cases, and adversarial calls for each workflow — then runs them as real voice calls.

Bookings, changes, and cancellations

New reservations, date changes, and cancellations, including the calls where the guest changes dates, room types, or travelers mid-booking.

  • Guest changes dates onto a sold-out night
  • Cancellation just inside the non-refundable window
  • Caller books for a group with mixed room types
  • Name on the booking doesn't match the caller

Front desk and reservations

Hotel questions about check-in times, amenities, and special requests, answered from policy instead of guesswork.

  • Early check-in request on a full-occupancy day
  • Guest asks whether the resort fee is included
  • Pet, accessibility, or crib request at booking
  • Caller disputes a charge on a past stay

Disruption rebooking and loyalty

Delays, cancellations, and missed connections handled calmly, with loyalty balances and benefits stated correctly.

  • Traveler rebooking after a cancelled flight at midnight
  • Angry caller demanding compensation not in policy
  • Member asks to redeem points for a partial stay
  • Status benefit requested that the tier doesn't include

Outcome metrics

Measure what success means in travel & hospitality

Every call gets scored, so these become rates you can track across every test batch and every production call.

Booking accuracy

Reservations booked or changed with dates and details confirmed back to the guest.

Resolution rate

Share of calls where the guest's request was fully handled.

Containment rate

Calls completed without a transfer to the reservations desk.

Fee disclosure rate

Change fees and cancellation terms stated before the guest confirms.

Evaluate

What RubricHQ checks on every call

Each simulated or production call is scored with code-as-judge rules, LLM-as-judge metrics, and audio metrics — so a failure shows up as a failed metric, not a one-star review.

Booking accuracy and fees

Did the agent confirm the right dates, room, and traveler, and quote every fee before the guest agreed?

  • Dates, room type, and guest count read back correctly
  • Taxes and resort fees disclosed before confirmation
  • No price or availability promised that the system can't honor

Policy and identity

Refund, change, and loyalty answers should match your policy, and booking details should only reach the booking holder.

  • Cancellation and refund terms stated as written
  • Caller verified before itinerary details are shared
  • No compensation or upgrade offered outside policy

Real caller conditions

Travelers call from airports, taxis, and hotel lobbies. Test for them before launch.

  • Accented and non-native English callers
  • Airport noise and unstable mobile connections
  • Latency, dead-air, and interruptions measured per turn

Metrics to score every travel & hospitality call

Ready to run from the metric gallery

[General] Appointment / Scheduling Accuracy[General] Identity Verification Before Sensitive Disclosure[General] No Misrepresentation or False Claims[General] Primary Issue Resolution[General] Agent Empathy & Tone

Custom metrics you can add

+Booking details read back before confirmation+All fees quoted before booking+Refund terms match fare or rate rules+Loyalty balance stated correctly
How evaluation works →

Callers to test against

Personas from the built-in library, with multi-language support.

AEAngry EscalatorIMImpatient MultitaskerIDIndecisive DithererDVDemanding VIPNCNoisy Cafe CallerUMUnstable MobileBEBritish English Caller
How simulation works →

From prompt to production in one platform

  1. 01

    Simulate

    Auto-generate scenarios from your prompt and run them as concurrent voice calls.

  2. 02

    Evaluate

    Score every call with code, LLM-as-judge, and audio metrics.

  3. 03

    Optimize

    Diagnose failures and prove prompt fixes before you ship.

  4. 04

    Monitor

    Score live production calls and get alerted in Slack or email when checks fail.

Travel & Hospitality voice AI testing FAQ

Can RubricHQ check that my agent quotes the right fees?+

Yes. Prebuilt metrics flag misrepresentation and false claims, and you can add a custom metric for your own rules — for example, 'resort fee and taxes are stated before the booking is confirmed' — scored on every call.

Can it test disruption and rebooking calls?+

Yes. Describe your rebooking policy in the agent prompt and RubricHQ generates scenarios for cancellations, delays, and missed connections, including callers who are angry, impatient, or asking for compensation outside policy.

Is RubricHQ PCI certified?+

No. RubricHQ does not hold a PCI DSS certification. Simulated callers use made-up identities by default, and we recommend testing payment flows with synthetic data rather than real cards.

Which voice platforms does it work with?+

RubricHQ connects to agents built on Vapi, Retell, LiveKit, and Pipecat, and can call any agent with a phone number. Production calls from other platforms can be sent in through the API for scoring.

Can it simulate international travelers?+

Yes. Simulations can run in multiple languages, and the persona library includes British, Australian, Indian, and non-native English callers, plus noisy and poor-connection conditions.

How do I catch problems after launch?+

Send production calls in through the API and they are scored with the same metrics as your tests. Alert rules notify Slack, email, or a webhook when production calls fail a metric you choose, and RubricHQ can diagnose the failure and suggest a prompt rewrite.

Test your travel & hospitality voice agent today

$0 to start — 200 free credits, no credit card. Connect your agent and run your first batch of test calls in minutes.

More industries