ZBS Index What actually exists in applied AI, with the source next to it

Decision page

Voice client intake stack for Legal

What should a legal team use for a voice agent that collects initial client-intake information?

ZBS editorial starting point. Start with Vapi for calls, Fish Audio for speech and Claude API for a bounded intake script with immediate escalation.

The firm receives structured intake notes while advice, engagement and emergencies remain with people.

Editorial starting point

Managed voice intake

A concrete starting configuration that keeps source facts, model output and operational authority separate.

Choose this when: The firm needs after-hours intake and accepts reviewed external processing.

  1. Calls Vapi observed

    Connect telephony and agent turns. The official documentation exposes a provider-composition path for phone agents.

    Limit: Recording, retention, telephony and every configured provider remain separate boundaries.

    Evidence: source 1

  2. Voice Fish Audio TTS observed

    Speak questions and confirmations. The speech API can be evaluated and replaced independently of the rest of the agent.

    Limit: Pronunciation, language quality, latency and full-call cost require direct measurement.

    Evidence: source 1

  3. Intake Claude API observed

    Follow the approved intake schema. The external API is a practical bounded reasoning layer with documented behavior.

    Limit: The model must not replace source-of-truth data, deterministic policy or human authority.

    Evidence: source 1

Private / local

Controlled reasoning path

Keep parsing, retrieval or model inference in controlled infrastructure while retaining the same source-of-truth and approval rules.

Choose this when: Sensitive inputs cannot be sent to an external model API and the team can operate the additional infrastructure.

  1. Calls LiveKit Agents source backed inference

    Run the realtime workflow. The framework documents a provider-flexible realtime agent runtime.

    Limit: Self-operation transfers telephony, scaling and recovery work to the team.

    Evidence: source 1

  2. Speech input faster-whisper source backed inference

    Transcribe callers locally. It provides an open speech-recognition path that can run in controlled infrastructure.

    Limit: Phone audio, accents, names and realtime performance require direct testing.

    Evidence: source 1

  3. Intake vLLM source backed inference

    Serve a local intake model. It provides a documented self-operated model-serving layer.

    Limit: Serving a model does not prove its task accuracy, safe tool use or secure operation.

    Evidence: source 1

  4. Voice Fish Speech source backed inference

    Synthesize responses locally. The public repository provides a self-operated speech path.

    Limit: Licensing, model files, latency and language quality need separate qualification.

    Evidence: source 1

Budget alternative

Lower-cost external model path

Keep the workflow and source integration explicit while evaluating a lower-cost model candidate on the same acceptance set.

Choose this when: External processing is acceptable and measured model spend is a leading constraint.

  1. Calls Vapi source backed inference

    Keep orchestration managed. The official documentation exposes a provider-composition path for phone agents.

    Limit: Recording, retention, telephony and every configured provider remain separate boundaries.

    Evidence: source 1

  2. Intake DeepSeek API source backed inference

    Produce structured intake notes. It is a concrete lower-cost external model candidate for the same acceptance set.

    Limit: Price alone is not task fitness; output structure, languages, availability and data terms need testing.

    Evidence: source 1

  3. Voice Fish Audio TTS source backed inference

    Stream the spoken script. The speech API can be evaluated and replaced independently of the rest of the agent.

    Limit: Pronunciation, language quality, latency and full-call cost require direct measurement.

    Evidence: source 1

Community check

Do you agree with this starting stack?

This is a reader opinion about the whole editorial recommendation, not evidence that the stack is objectively good. Votes never change it automatically.

Loading reader votes…

Voting needs JavaScript. The recommendation and every source above remain available without it.

Trade-offs that change the choice

ConstraintPrimaryPrivate / localBudget
Data boundary The named managed APIs receive only the fields explicitly sent to them Reasoning stays controlled; source systems may remain externalLower cost does not make external processing private
Operational load Lower: managed components with explicit integration points Highest: serving, retrieval and recovery are yoursModerate: custom workflow plus external APIs
Decision authority Risky writes and low-confidence cases require a deterministic or human gate The same gate is required regardless of hostingLower model price does not relax the approval rule

Implementation path

1. Write permitted questions, consent and emergency escalation.

2. Test conflicts, urgency, silence and requests for advice.

3. Require human review before matter-system entry.

4. Measure missing fields, unsafe advice and failed transfers.

Known limits

The agent cannot decide representation.

Recording and consent rules need jurisdictional review.

No product on this page is a universal winner; the configuration still needs a task-specific acceptance test.

EU and US routes stay consolidated with Global until evidence changes the answer.

Validate this stack on your data

A recommendation is a starting point. Practice Lab can test the same workflow on representative inputs, constraints and failure cases.

Request a real-data evaluation

Sources

  1. DeepSeek API documentation — DeepSeek, observed , trust tier 2.
  2. Fish Audio text-to-speech API — Fish Audio, observed , trust tier 2.
  3. LiveKit Agents documentation — LiveKit, observed , trust tier 2.
  4. Vapi introduction — Vapi, observed , trust tier 2.
  5. vLLM documentation — vLLM, observed , trust tier 2.
  6. Fish Speech repository — Fish Audio, observed , trust tier 3.
  7. faster-whisper repository — SYSTRAN, observed , trust tier 3.
  8. ZBS Index solution-stack editorial synthesis — ZBS Index, observed , trust tier 7.
  9. Claude API overview — Anthropic, observed , trust tier 2.