Test before launch
Simulate
Auto-generate scenarios & run calls
Evaluate
Metrics & transcript analysis
Optimize
AI-rewritten prompt improvements
Monitor in production
Monitor
Live production call observability
Co-Pilot
Ask your calls anything
IndustriesPricingAPI documentationAbout usBlogs
Book a demoLog inSign up
IndustriesLogistics

Voice AI testing for logistics and delivery agents

Simulate thousands of shipper, recipient, and driver calls, scoring each against your rules for ETA accuracy, address verification, and rescheduling.

No credit card · 200 free credits

Works withVapiRetellLiveKitPipecatElevenLabsOpenAI

Simulate

Simulate the calls your agent handles

Paste your agent prompt and RubricHQ generates happy paths, edge cases, and adversarial calls for each workflow — then runs them as real voice calls.

Delivery status (WISMO)

Where-is-my-order calls answered from tracking data, including the ones where the tracking number is wrong or the status is stale.

  • Tracking number read out with a digit missing
  • Package shows delivered but the caller doesn't have it
  • Caller asks for an exact delivery time window
  • Shipment split across multiple deliveries

Missed, damaged, and address changes

Redelivery, damage claims, and address updates handled with the right checks before anything changes.

  • Recipient requests redelivery on a weekend
  • Caller reports a damaged package and wants a refund
  • Address change requested by someone other than the recipient
  • Hold-at-location request for an in-transit package

Driver and dispatch lines

Drivers calling in from the road with delays, access problems, and exceptions that need a dispatcher.

  • Driver can't access a gated delivery address
  • Vehicle breakdown mid-route with packages on board
  • Driver reports a refused delivery
  • Hands-free caller on a highway with road noise

Outcome metrics

Measure what success means in logistics

Every call gets scored, so these become rates you can track across every test batch and every production call.

WISMO containment

Where-is-my-order calls answered without a transfer to support.

Resolution rate

Share of calls where the caller's delivery issue was fully resolved.

Reschedule accuracy

Missed deliveries rebooked with the new window confirmed back to the caller.

ETA accuracy

Delivery estimates match live tracking data, never guessed.

Evaluate

What RubricHQ checks on every call

Each simulated or production call is scored with code-as-judge rules, LLM-as-judge metrics, and audio metrics — so a failure shows up as a failed metric, not a lost package.

Status and ETA accuracy

Did the agent report what the tracking data says, or did it promise a delivery time it can't back up?

  • Status and ETA match the shipment record
  • No delivery guarantee made outside policy
  • Tracking and order numbers read back correctly

Verification before changes

Address changes and redirects should only happen for a verified sender or recipient.

  • Caller verified before any address change
  • No shipment details shared with an unverified caller
  • Damage and loss claims routed to the right team

Real caller conditions

Drivers call from cabs and loading docks; recipients call from doorsteps. Test for them before launch.

  • Road noise, hands-free, and dropped audio
  • Fast talkers and constant interrupters
  • Latency, dead-air, and interruptions measured per turn

Metrics to score every logistics call

Ready to run from the metric gallery

[General] Identity Verification Before Sensitive Disclosure[General] Identity Verification Method Used[General] Primary Issue Resolution[General] No Misrepresentation or False Claims[General] Appointment / Scheduling Accuracy

Custom metrics you can add

+ETA matches tracking data+Address change read back before saving+Driver exception escalated to dispatch+Damage claim details captured
How evaluation works →

Callers to test against

Personas from the built-in library, with multi-language support.

FCFrustrated CustomerIMImpatient MultitaskerFTFast TalkerCIConstant InterrupterDHDriving Hands-FreePCPoor ConnectionEWEvasive Won't Verify
How simulation works →

From prompt to production in one platform

  1. 01

    Simulate

    Auto-generate scenarios from your prompt and run them as concurrent voice calls.

  2. 02

    Evaluate

    Score every call with code, LLM-as-judge, and audio metrics.

  3. 03

    Optimize

    Diagnose failures and prove prompt fixes before you ship.

  4. 04

    Monitor

    Score live production calls and get alerted in Slack or email when checks fail.

Logistics voice AI testing FAQ

Can RubricHQ check my agent's delivery status answers?+

Yes. Add a custom metric that compares what the agent said against the shipment data it was given — for example, 'the ETA stated matches the tracking record' — and it is scored on every simulated and production call.

Can it test address-change fraud attempts?+

Yes. RubricHQ generates adversarial scenarios such as a caller who isn't the recipient trying to redirect a package, and prebuilt metrics check that identity was verified before anything sensitive was shared.

Can it simulate drivers calling from the road?+

Yes. The persona library includes hands-free drivers, poor connections, unstable mobile audio, and fast talkers, and audio metrics measure latency, dead-air, and interruptions on every turn.

Which voice platforms does it work with?+

RubricHQ connects to agents built on Vapi, Retell, LiveKit, and Pipecat, and can call any agent with a phone number. Production calls from other platforms can be sent in through the API for scoring.

How long does a test run take?+

Calls run in parallel, so a batch finishes far faster than dialing by hand; concurrency depends on your plan. You can also schedule runs nightly or gate deploys in CI with the RubricHQ GitHub Action.

Does it support multilingual customers and drivers?+

Yes. Simulations can run in multiple languages, and the persona library includes accented and non-native English callers.

Test your logistics voice agent today

$0 to start — 200 free credits, no credit card. Connect your agent and run your first batch of test calls in minutes.

More industries