ZBS Index What actually exists in applied AI, with the source next to it

mcp server

Speech AI - Pronunciation, STT & TTS

Pronunciation scoring, speech-to-text, and text-to-speech for language learning

Description as published by the maintainer. Source

  • version 2.3.0
  • active
  • voice
  • evaluation

active — Most recent push to the repository was 2026-08-03. Dashed tags are derived by ZBS Index from the published description, not stated by the maintainer.

What this server can do

10 functions, named and described by the server itself. Parameter names are shown because they say more about what a function does than its name usually does.

assess_pronunciation(text, audio_base64, audio_format)
Assess English pronunciation quality from audio. Scores pronunciation at four levels: overall, sentence, word, and phoneme. Each score is 0-100. Phonemes are returned in both IPA and ARPAbet notation. Sub-300ms inference latency. Args: audio_base64: Base64-encoded audio data. Supports WAV, MP3, OGG, and WebM formats. text: The reference English text that the speaker was expected to read aloud. audio_format: Audio format hint — one of 'wav', 'mp3', 'ogg', 'webm'. Defaults to 'wav'. Returns: dict with keys: - overallScore (int 0-100): Overall pronunciation quality - sentenceScore (int 0-100): Sentence-level fluency and accuracy - words (list): Per-word scores, each containing: - word (str): The word - score (int 0-100): Word pronunciation score - phonemes (list): Per-phoneme scores with IPA/ARPAbet notation - decodedTranscript (str): What the model heard (ASR transcript) - transcript (str): Reference text - confidence (float 0-1): Scoring confidence - warnings (list[str]): Quality warnings if any - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.) Required: audio_base64, text.
check_pronunciation_service
Check if the pronunciation assessment service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the scoring model is loaded - version (str): API version
check_stt_service
Check if the speech-to-text service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the STT model is loaded - version (str): API version
check_tts_service
Check if the text-to-speech service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the TTS model is loaded - version (str): API version
check_whisper_service
Check if the Whisper STT Pro service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the Whisper model is loaded - diarizeLoaded (bool): Whether the diarization pipeline is loaded - version (str): API version - modelName (str): Whisper model name (e.g. 'large-v3-turbo')
get_phoneme_inventory
Get the full phoneme inventory supported by the pronunciation scorer. Returns a list of all English phonemes the engine can assess, including ARPAbet symbol, IPA equivalent, example word, and phoneme category (vowel, consonant, diphthong). Returns: list of dicts, each with keys: - arpabet (str): ARPAbet symbol (e.g. 'AA', 'TH') - ipa (str): IPA notation - example (str): Example word containing the phoneme - category (str): vowel, consonant, or diphthong
list_tts_voices
List all available text-to-speech voices with metadata. Returns: dict with keys: - voices (list): Available voices, each with id, name, gender, accent, grade - defaultVoice (str): Default voice ID
synthesize_speech(text, speed, voice)
Generate natural speech audio from English text. Produces high-quality speech with 12 English voices. Returns base64-encoded WAV audio (16-bit PCM, 24kHz mono) along with metadata. Available voices: - af_heart (default), af_bella, af_nicole, af_sarah, af_sky (American female) - am_adam, am_michael (American male) - bf_emma, bf_isabella (British female) - bm_george, bm_lewis, bm_daniel (British male) Args: text: English text to synthesize (1-5000 characters). voice: Voice ID. See list above. Defaults to 'af_heart'. speed: Speed multiplier from 0.5 to 2.0 (default: 1.0). Returns: dict with keys: - audio_base64 (str): Base64-encoded WAV audio (16-bit PCM, 24kHz) - duration_ms (str): Audio duration in milliseconds - voice (str): Voice ID used - text_length (str): Input text character count - processing_ms (str): Synthesis time in milliseconds Required: text.
transcribe_audio(audio_base64, audio_format, include_timestamps)
Transcribe audio to text with word-level timestamps. Converts spoken English audio into text with optional word-level timestamps and per-word confidence scores. Args: audio_base64: Base64-encoded audio data (WAV, MP3, OGG, FLAC, WebM). audio_format: Audio format hint. Auto-detected from magic bytes if omitted. include_timestamps: Whether to include word-level timing (default: true). Returns: dict with keys: - text (str): Full decoded transcript - words (list): Per-word results with timestamps, each containing: - word (str): The transcribed word - start (float): Start time in seconds - end (float): End time in seconds - confidence (float 0-1): Word-level confidence - audioDurationMs (int): Audio duration in milliseconds - metadata (dict): Processing time, audio length, model version - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.) Required: audio_base64.
transcribe_audio_pro(diarize, language, audio_base64)
Transcribe audio with Whisper Large V3 Turbo — multilingual STT. Supports 99 languages with automatic language detection, word-level timestamps, per-word confidence scores, and optional speaker diarization (identifies who spoke each word). Best-in-class WER (~2%). Args: audio_base64: Base64-encoded audio (WAV, MP3, OGG, FLAC, WebM). language: Language code. Auto-detected if omitted. Supports 99 languages. diarize: Enable speaker diarization (default: false). When true, each word includes a speaker label (e.g. SPEAKER_00, SPEAKER_01). Returns: dict with keys: - text (str): Full decoded transcript - words (list): Per-word results with timestamps, each containing: - word (str), start (float), end (float), confidence (float 0-1) - speaker (str|null): Speaker label when diarize=true - speakers (dict|null): Speaker info with count and labels - audioDurationMs (int): Audio duration in milliseconds - metadata (dict): Processing time, language, languageProbability - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.) Required: audio_base64.

Last successful function declaration observed on . Source: https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp. We list what the server declared; we do not call any of these functions.

Endpoint status observed on . Source: https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp.

Signals

These are separate measurements of different things. They are deliberately not combined into one score, because a popularity number that mixes website traffic with saves and stars cannot be checked or acted on.

Signal Value What it measures Window Observed Source
GitHub stars 2 Number of GitHub accounts that bookmarked this repository since it was created. It is a bookmark count, not installs, not active users and not quality. cumulative, all time GitHub
Last commit 2026-08-03 Date of the most recent push to any branch. This is the strongest cheap indicator of whether the project is still maintained. point in time GitHub
Open issues 24 Open issues plus open pull requests, as GitHub counts them together. A high number can mean an active project or an abandoned one. as of fetch GitHub
Latest published version 2.3.0 Latest version string the maintainer published to the registry. as of fetch Model Context Protocol
Registry record last updated 2026-03-05 When the registry record was last updated by its maintainer. point in time Model Context Protocol
License MIT Licence GitHub detected in the repository. Detection can be wrong; the LICENSE file is authoritative. as of fetch GitHub
First listed in the MCP Registry 2026-03-05 Date this server was first published to the official MCP Registry. Not a usage or quality measure. point in time Model Context Protocol
repository status active The repository exists on GitHub and is not archived. This says nothing about how recently it was worked on. as of fetch GitHub
mcp tools declared 10 tools Number of functions the server itself declared when asked to list them. This is what the server offers an agent, not a measure of how well any of them work. as of probe apim-ai-apis.azure-api.net
mcp endpoint status ok The server listed 10 functions when asked. as of probe apim-ai-apis.azure-api.net

Where to get it

Related, by what their authors tagged them

  • Image Tools - Background Removal, Upscaling & Face Restoration — last commit 2026-08-03, shares api-examples, language-learning, pronunciation
    Background removal, 4x upscaling, and face restoration via GPU
  • NLP Tools - Sentiment, NER, Toxicity & Language Detection — last commit 2026-08-03, shares api-examples, language-learning, pronunciation
    Toxicity, sentiment, NER, PII detection, and language identification tools
  • io.github.chicogong/ffvoice — last commit 2026-05-19, shares speaker-diarization, speech-to-text
    Offline speech-to-text & speaker diarization MCP server: transcribe audio on-device, no cloud
  • com.brainiall/tts — last commit 2026-08-03, shares pt-br, text-to-speech
    Hosted pay-per-use TTS: 54 neural voices, 9 languages incl. Brazilian Portuguese. $10 free credits.
  • io.github.anzy-renlab-ai/pronounce — last commit 2026-07-28, shares pronunciation
    How engineers actually pronounce developer jargon — 1848+ sourced entries (kubectl, nginx, GIF).
  • Speko AI — last commit 2026-08-05, shares speech-to-text, text-to-speech
    Manage Speko voice-AI agents, sessions, calls, phone numbers, knowledge bases, evals, and docs.
  • Perfex CRM — last commit 2026-08-04, shares api-examples
    Read and write a self-hosted Perfex CRM from an AI agent: 148 permission-filtered tools.
  • Torify — Japan Locale APIs for AI Agents — last commit 2026-05-29, shares api-examples
    39 Japanese locale APIs — wareki, NTA invoice, 法人番号, postal, romanization, kanji-kana (Workers AI).
  • Nummeropslag — last commit 2026-08-04, shares api-examples
    Privacy-first Danish caller ID, CVR company lookup, operator data, and community spam signals.
  • dev.waxberry/live-translate-mcp — last commit 2026-06-17, shares speech-to-text
    MCP server for local speech translation (EN ↔ 中文) via Whisper + Claude + Piper

These share tags the maintainers applied themselves, such as api-examples, language-learning, pronunciation, pt-br. Common tags like "mcp" or "ai" are ignored for this: agreeing with six hundred other projects is not a similarity.

This is not a recommendation and not a test result. It is a map of what the authors said their work is about.

How the author describes it

Topics the maintainer set on GitHub: ai-agents, api-examples, language-learning, mcp, pronunciation, pt-br, speaker-diarization, speech-ai, speech-to-text, synthetic-testing, text-to-speech, webvtt.

This record as data

Every field on this page, with its source and observation date, is in the catalog JSON. Fetch the whole kind at once instead of parsing this HTML.

GET /api/v1/entries/mcp_server.json

Sources

  1. fasuizu-br/speech-ai-examples on GitHub — GitHub, observed , trust tier 3.
  2. Tools declared by the MCP server at https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp — apim-ai-apis.azure-api.net, observed , trust tier 1.
  3. Official MCP Registry — Model Context Protocol, observed , trust tier 1.