Decision page
Voice client intake stack for Legal
What should a legal team use for a voice agent that collects initial client-intake information?
ZBS editorial starting point. Start with Vapi for calls, Fish Audio for speech and Claude API for a bounded intake script with immediate escalation.
The firm receives structured intake notes while advice, engagement and emergencies remain with people.
Editorial starting point
Managed voice intake
A concrete starting configuration that keeps source facts, model output and operational authority separate.
Choose this when: The firm needs after-hours intake and accepts reviewed external processing.
-
Connect telephony and agent turns. The official documentation exposes a provider-composition path for phone agents.
Limit: Recording, retention, telephony and every configured provider remain separate boundaries.
Evidence: source 1
-
Speak questions and confirmations. The speech API can be evaluated and replaced independently of the rest of the agent.
Limit: Pronunciation, language quality, latency and full-call cost require direct measurement.
Evidence: source 1
-
Follow the approved intake schema. The external API is a practical bounded reasoning layer with documented behavior.
Limit: The model must not replace source-of-truth data, deterministic policy or human authority.
Evidence: source 1
Private / local
Controlled reasoning path
Keep parsing, retrieval or model inference in controlled infrastructure while retaining the same source-of-truth and approval rules.
Choose this when: Sensitive inputs cannot be sent to an external model API and the team can operate the additional infrastructure.
-
Run the realtime workflow. The framework documents a provider-flexible realtime agent runtime.
Limit: Self-operation transfers telephony, scaling and recovery work to the team.
Evidence: source 1
-
Transcribe callers locally. It provides an open speech-recognition path that can run in controlled infrastructure.
Limit: Phone audio, accents, names and realtime performance require direct testing.
Evidence: source 1
-
Intake
vLLM
source backed inference
Serve a local intake model. It provides a documented self-operated model-serving layer.
Limit: Serving a model does not prove its task accuracy, safe tool use or secure operation.
Evidence: source 1
-
Synthesize responses locally. The public repository provides a self-operated speech path.
Limit: Licensing, model files, latency and language quality need separate qualification.
Evidence: source 1
Budget alternative
Lower-cost external model path
Keep the workflow and source integration explicit while evaluating a lower-cost model candidate on the same acceptance set.
Choose this when: External processing is acceptable and measured model spend is a leading constraint.
-
Calls
Vapi
source backed inference
Keep orchestration managed. The official documentation exposes a provider-composition path for phone agents.
Limit: Recording, retention, telephony and every configured provider remain separate boundaries.
Evidence: source 1
-
Produce structured intake notes. It is a concrete lower-cost external model candidate for the same acceptance set.
Limit: Price alone is not task fitness; output structure, languages, availability and data terms need testing.
Evidence: source 1
-
Stream the spoken script. The speech API can be evaluated and replaced independently of the rest of the agent.
Limit: Pronunciation, language quality, latency and full-call cost require direct measurement.
Evidence: source 1
Trade-offs that change the choice
Implementation path
1. Write permitted questions, consent and emergency escalation.
2. Test conflicts, urgency, silence and requests for advice.
3. Require human review before matter-system entry.
4. Measure missing fields, unsafe advice and failed transfers.
Known limits
The agent cannot decide representation.
Recording and consent rules need jurisdictional review.
No product on this page is a universal winner; the configuration still needs a task-specific acceptance test.
EU and US routes stay consolidated with Global until evidence changes the answer.
Validate this stack on your data
A recommendation is a starting point. Practice Lab can test the same workflow on representative inputs, constraints and failure cases.
Request a real-data evaluation