mcp server
Speech AI - Pronunciation, STT & TTS
Pronunciation scoring, speech-to-text, and text-to-speech for language learning
Description as published by the maintainer. Source
- version 2.3.0
- active
- voice
- evaluation
active — Most recent push to the repository was 2026-08-03. Dashed tags are derived by ZBS Index from the published description, not stated by the maintainer.
What this server can do
10 functions, named and described by the server itself. Parameter names are shown because they say more about what a function does than its name usually does.
assess_pronunciation(text, audio_base64, audio_format)- Assess English pronunciation quality from audio. Scores pronunciation at four levels: overall, sentence, word, and phoneme. Each score is 0-100. Phonemes are returned in both IPA and ARPAbet notation. Sub-300ms inference latency. Args: audio_base64: Base64-encoded audio data. Supports WAV, MP3, OGG, and WebM formats. text: The reference English text that the speaker was expected to read aloud. audio_format: Audio format hint — one of 'wav', 'mp3', 'ogg', 'webm'. Defaults to 'wav'. Returns: dict with keys: - overallScore (int 0-100): Overall pronunciation quality - sentenceScore (int 0-100): Sentence-level fluency and accuracy - words (list): Per-word scores, each containing: - word (str): The word - score (int 0-100): Word pronunciation score - phonemes (list): Per-phoneme scores with IPA/ARPAbet notation - decodedTranscript (str): What the model heard (ASR transcript) - transcript (str): Reference text - confidence (float 0-1): Scoring confidence - warnings (list[str]): Quality warnings if any - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.) Required: audio_base64, text.
check_pronunciation_service- Check if the pronunciation assessment service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the scoring model is loaded - version (str): API version
check_stt_service- Check if the speech-to-text service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the STT model is loaded - version (str): API version
check_tts_service- Check if the text-to-speech service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the TTS model is loaded - version (str): API version
check_whisper_service- Check if the Whisper STT Pro service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the Whisper model is loaded - diarizeLoaded (bool): Whether the diarization pipeline is loaded - version (str): API version - modelName (str): Whisper model name (e.g. 'large-v3-turbo')
get_phoneme_inventory- Get the full phoneme inventory supported by the pronunciation scorer. Returns a list of all English phonemes the engine can assess, including ARPAbet symbol, IPA equivalent, example word, and phoneme category (vowel, consonant, diphthong). Returns: list of dicts, each with keys: - arpabet (str): ARPAbet symbol (e.g. 'AA', 'TH') - ipa (str): IPA notation - example (str): Example word containing the phoneme - category (str): vowel, consonant, or diphthong
list_tts_voices- List all available text-to-speech voices with metadata. Returns: dict with keys: - voices (list): Available voices, each with id, name, gender, accent, grade - defaultVoice (str): Default voice ID
synthesize_speech(text, speed, voice)- Generate natural speech audio from English text. Produces high-quality speech with 12 English voices. Returns base64-encoded WAV audio (16-bit PCM, 24kHz mono) along with metadata. Available voices: - af_heart (default), af_bella, af_nicole, af_sarah, af_sky (American female) - am_adam, am_michael (American male) - bf_emma, bf_isabella (British female) - bm_george, bm_lewis, bm_daniel (British male) Args: text: English text to synthesize (1-5000 characters). voice: Voice ID. See list above. Defaults to 'af_heart'. speed: Speed multiplier from 0.5 to 2.0 (default: 1.0). Returns: dict with keys: - audio_base64 (str): Base64-encoded WAV audio (16-bit PCM, 24kHz) - duration_ms (str): Audio duration in milliseconds - voice (str): Voice ID used - text_length (str): Input text character count - processing_ms (str): Synthesis time in milliseconds Required: text.
transcribe_audio(audio_base64, audio_format, include_timestamps)- Transcribe audio to text with word-level timestamps. Converts spoken English audio into text with optional word-level timestamps and per-word confidence scores. Args: audio_base64: Base64-encoded audio data (WAV, MP3, OGG, FLAC, WebM). audio_format: Audio format hint. Auto-detected from magic bytes if omitted. include_timestamps: Whether to include word-level timing (default: true). Returns: dict with keys: - text (str): Full decoded transcript - words (list): Per-word results with timestamps, each containing: - word (str): The transcribed word - start (float): Start time in seconds - end (float): End time in seconds - confidence (float 0-1): Word-level confidence - audioDurationMs (int): Audio duration in milliseconds - metadata (dict): Processing time, audio length, model version - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.) Required: audio_base64.
transcribe_audio_pro(diarize, language, audio_base64)- Transcribe audio with Whisper Large V3 Turbo — multilingual STT. Supports 99 languages with automatic language detection, word-level timestamps, per-word confidence scores, and optional speaker diarization (identifies who spoke each word). Best-in-class WER (~2%). Args: audio_base64: Base64-encoded audio (WAV, MP3, OGG, FLAC, WebM). language: Language code. Auto-detected if omitted. Supports 99 languages. diarize: Enable speaker diarization (default: false). When true, each word includes a speaker label (e.g. SPEAKER_00, SPEAKER_01). Returns: dict with keys: - text (str): Full decoded transcript - words (list): Per-word results with timestamps, each containing: - word (str), start (float), end (float), confidence (float 0-1) - speaker (str|null): Speaker label when diarize=true - speakers (dict|null): Speaker info with count and labels - audioDurationMs (int): Audio duration in milliseconds - metadata (dict): Processing time, language, languageProbability - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.) Required: audio_base64.
Last successful function declaration observed on . Source: https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp. We list what the server declared; we do not call any of these functions.
Endpoint status observed on . Source: https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp.
Signals
These are separate measurements of different things. They are deliberately not combined into one score, because a popularity number that mixes website traffic with saves and stars cannot be checked or acted on.
| Signal | Value | What it measures | Window | Observed | Source |
|---|---|---|---|---|---|
| GitHub stars | 2 | Number of GitHub accounts that bookmarked this repository since it was created. It is a bookmark count, not installs, not active users and not quality. | cumulative, all time | GitHub | |
| Last commit | 2026-08-03 | Date of the most recent push to any branch. This is the strongest cheap indicator of whether the project is still maintained. | point in time | GitHub | |
| Open issues | 24 | Open issues plus open pull requests, as GitHub counts them together. A high number can mean an active project or an abandoned one. | as of fetch | GitHub | |
| Latest published version | 2.3.0 | Latest version string the maintainer published to the registry. | as of fetch | Model Context Protocol | |
| Registry record last updated | 2026-03-05 | When the registry record was last updated by its maintainer. | point in time | Model Context Protocol | |
| License | MIT | Licence GitHub detected in the repository. Detection can be wrong; the LICENSE file is authoritative. | as of fetch | GitHub | |
| First listed in the MCP Registry | 2026-03-05 | Date this server was first published to the official MCP Registry. Not a usage or quality measure. | point in time | Model Context Protocol | |
| repository status | active | The repository exists on GitHub and is not archived. This says nothing about how recently it was worked on. | as of fetch | GitHub | |
| mcp tools declared | 10 tools | Number of functions the server itself declared when asked to list them. This is what the server offers an agent, not a measure of how well any of them work. | as of probe | apim-ai-apis.azure-api.net | |
| mcp endpoint status | ok | The server listed 10 functions when asked. | as of probe | apim-ai-apis.azure-api.net |
Where to get it
Related, by what their authors tagged them
-
Image Tools - Background Removal, Upscaling & Face Restoration
— last commit 2026-08-03, shares api-examples, language-learning, pronunciation
Background removal, 4x upscaling, and face restoration via GPU
-
NLP Tools - Sentiment, NER, Toxicity & Language Detection
— last commit 2026-08-03, shares api-examples, language-learning, pronunciation
Toxicity, sentiment, NER, PII detection, and language identification tools
-
io.github.chicogong/ffvoice
— last commit 2026-05-19, shares speaker-diarization, speech-to-text
Offline speech-to-text & speaker diarization MCP server: transcribe audio on-device, no cloud
-
com.brainiall/tts
— last commit 2026-08-03, shares pt-br, text-to-speech
Hosted pay-per-use TTS: 54 neural voices, 9 languages incl. Brazilian Portuguese. $10 free credits.
-
io.github.anzy-renlab-ai/pronounce
— last commit 2026-07-28, shares pronunciation
How engineers actually pronounce developer jargon — 1848+ sourced entries (kubectl, nginx, GIF).
-
Speko AI
— last commit 2026-08-05, shares speech-to-text, text-to-speech
Manage Speko voice-AI agents, sessions, calls, phone numbers, knowledge bases, evals, and docs.
-
Perfex CRM
— last commit 2026-08-04, shares api-examples
Read and write a self-hosted Perfex CRM from an AI agent: 148 permission-filtered tools.
-
Torify — Japan Locale APIs for AI Agents
— last commit 2026-05-29, shares api-examples
39 Japanese locale APIs — wareki, NTA invoice, 法人番号, postal, romanization, kanji-kana (Workers AI).
-
Nummeropslag
— last commit 2026-08-04, shares api-examples
Privacy-first Danish caller ID, CVR company lookup, operator data, and community spam signals.
-
dev.waxberry/live-translate-mcp
— last commit 2026-06-17, shares speech-to-text
MCP server for local speech translation (EN ↔ 中文) via Whisper + Claude + Piper
These share tags the maintainers applied themselves, such as api-examples, language-learning, pronunciation, pt-br. Common tags like "mcp" or "ai" are ignored for this: agreeing with six hundred other projects is not a similarity.
This is not a recommendation and not a test result. It is a map of what the authors said their work is about.
How the author describes it
Topics the maintainer set on GitHub: ai-agents, api-examples, language-learning, mcp, pronunciation, pt-br, speaker-diarization, speech-ai, speech-to-text, synthetic-testing, text-to-speech, webvtt.
This record as data
Every field on this page, with its source and observation date, is in the catalog JSON. Fetch the whole kind at once instead of parsing this HTML.
GET /api/v1/entries/mcp_server.json