3 questions from the package
From the RFI round. The first shows part of the guide each question carries; the workbook adds follow-ups, how to verify the answer, a priority and a weight.
1. For each stack layer (transcription, model inference, speech synthesis, telephony), state whether you operate it on your own infrastructure or call a third-party provider.
Why it matters. Every layer run by a separate provider adds a network hop to each turn. A third-party outage or slowdown in any one layer raises reply latency on every call, and the buyer cannot address it directly.
Good answer
- Gives a layer-by-layer answer that names each layer as self-operated or third-party
- States whether the third-party providers are fixed, or vary by region, plan or configuration
- Explains where the layers are hosted relative to each other, such as the same data center, the same region or across regions
Red flags
- Describes the stack as a single proprietary pipeline but will not break it down by layer
- Says provider choice is an internal matter subject to change, with no list
- Answers for one layer and stays silent on the others
2. Describe the signals your product uses to decide that a caller has finished speaking and the agent should reply, such as silence duration, intonation or the words spoken so far.
Why it matters. If a fixed silence timer alone decides the end of a turn, the agent either cuts in on callers who pause mid-sentence or waits after every reply. Callers hear both as talking to a machine.
3. At a conversation step where the agent is listening for a spoken answer, can a caller respond by telephone keypad instead?
Why it matters. Callers in noisy places, callers with speech the model struggles with, and callers who do not want to say an account number aloud need a keypad path. Without one, those calls fail or transfer to a human.
Capability areas
End-to-end latency (10)
Time from the caller finishing speech to the agent's audible reply, measured over real telephony across transcription, model inference, synthesis and carrier legs, how latency behaves under load and on degraded connections, and which stack layers the vendor runs itself, calls from third parties, or lets the buyer replace with its own providers. Excludes how turns are detected, which is covered in Turn-taking and barge-in.
Turn-taking and barge-in (10)
End-of-turn detection, interruption handling, recovery after the caller talks over the agent, handling of silence and filler speech, and tuning controls for these behaviors. Excludes raw recognition accuracy.
Speech recognition and audio robustness (11)
Transcription accuracy on accented speech, background noise, low-bitrate mobile audio and domain vocabulary, and capture of names, addresses and alphanumeric strings. Includes DTMF input as an alternative. Excludes synthesis and language coverage of the agent's own voice.
Voice output, persona and spoken language (9)
Synthesized voice quality, pronunciation control for names and terms, reading of numbers, dates, amounts and codes, persona and tone configuration, and spoken-language detection and switching within a call. Excludes the languages of the administration interface.
Telephony, routing and concurrency (12)
SIP trunking and carrier connectivity, number provisioning and porting, coexistence with an existing IVR, ACD or CCaaS platform, outbound calling with answering-machine and voicemail detection, call recording at the media layer, concurrent-call capacity at peak and behavior at the limit, and where calls go when the agent platform is unavailable. Excludes the pricing of concurrency, which is covered by the commercial module, and outbound consent rules, which are covered in Regulated call controls.
Escalation and human handoff (12)
Conditions that trigger escalation, caller-requested transfer, warm and cold transfer mechanics into the buyer's contact-center routing, the context delivered to the human agent's desktop, queue behavior when no human is available or the center is closed, callback offers, and whether the caller has to repeat information.
Caller identification and authentication (10)
How the agent verifies a caller before a consequential action: caller ID matching, knowledge-based checks, one-time passcodes, voice biometrics, step-up rules per action type, and the behavior on failed verification. Excludes the agent's general action permissions, which are covered by the ai-agentic-autonomy module.
In-call data access and action execution (10)
Reading account records and executing writes during a live call, such as scheduling, address changes and order lookups, and how the agent handles backend timeouts, errors and partial completion mid-conversation. Excludes the general API and SSO surface, which is covered by the integration module.
Conversation design and agent building (10)
How agents are defined and edited (structured flows, prompts or both), grounding in knowledge sources, guardrails on scope, versioning, staged publishing and rollback, who can make changes, and reuse of one agent definition across voice and text channels. Excludes testing tooling.
Pre-production testing and simulation (9)
Replaying our recorded calls through the agent, simulated callers with noise, accents and interruptions, regression suites run before each change, and test runs over real telephony rather than a browser. Excludes monitoring of live traffic.
Production monitoring and call analytics (10)
Call transcripts and recordings for review, containment, transfer and abandonment reporting by call type, per-call latency and component usage traces, QA scoring workflows for contact-center staff, and alerting on degradation. Excludes model governance and change disclosure.
Regulated call controls (10)
In-call mechanisms for regulated conversations: recording-consent and AI-disclosure prompts, pausing or masking during payment card capture, redaction of sensitive data in transcripts and logs, and outbound calling consent and time-of-day rules. Excludes the vendor's certifications and attestations, which are covered by the compliance-certifications module.
Questions about this package
How many AI Voice Agents & IVR Replacement RFP questions are there?
123 solution questions in 12 capability areas: 26 for the RFI, 63 for the RFP and 34 deep-dive questions for the finalists. The workbook adds 100 due-diligence questions on security, integration, implementation and exit.
What comes with each question?
Why it matters, good-answer signals, red flags, follow-up questions, how to verify the answer (a demo step, a test or a document), and a suggested priority and weight for scoring.
Can I edit the questions?
Yes. The workbook is an ordinary Excel file. Change, add or remove questions, and change the weights; the scorecard recalculates.
Which license do I need?
The Enterprise License covers any number of evaluations inside one organization. The Consultancy License covers use with any number of clients. Neither allows reselling or republishing the questions.