mcp server
Islam West Africa Collection (IWAC)
Read-only access to the Islam West Africa Collection via Hugging Face datasets.
Description as published by the maintainer. Source
- version 1.5.1
- active
active — Most recent push to the repository was 2026-08-07.
What this server can do
34 functions, named and described by the server itself. Parameter names are shown because they say more about what a function does than its name usually does.
fetch(id)- Retrieve the full text and metadata of one IWAC item by an id returned from `search` (format '<category>:<number>', e.g. 'articles:28576'). Returns {id, title, text, url, metadata}: `text` is the item's OCR / abstract / transcription / description, `url` is the canonical islam.zmo.de link to cite, and `metadata` holds the remaining fields (author, date, country, newspaper, AI sentiment, …). Categories: articles, publications, references, documents, index, audiovisual, images. Required: id.
get_article(keyword, article_id, max_excerpts, context_chars)- Get one article (by id): full metadata, the AI abstract (description_ai), AI sentiment, and OCR text. Pass a `keyword` to get ~2000-char excerpts around each match instead of the full (capped) OCR. Required: article_id.
get_audiovisual(audiovisual_id)- Get one audiovisual record by id, including creator/publisher, media URL, duration, medium, subjects, places, language, source, and IWAC URL. Required: audiovisual_id.
get_collection_stats- Overall statistics for every IWAC subset, including `fulltext_coverage` — how many items in each subset actually carry searchable full text in this public dataset. Read that before treating any keyword count as a full-text census.
get_cooccurrence(field, top_n, subset, country, date_to, keyword, subject, date_from, newspaper)- How often the top values of a multi-valued field appear on the SAME item — a subject/place co-mention matrix. Answers 'what is X discussed alongside' without reading anything: the pair counts are the structure of the tagging. Returns the top values, the full symmetric matrix (diagonal = each value's own count) and the strongest pairs.
get_country_comparison- Compare article counts, newspaper counts, date ranges, and gpt-5-6-luna polarity across countries.
get_document(keyword, document_id, max_excerpts, context_chars)- Get one archival document (by id): full metadata, AI description, and OCR text. Pass a `keyword` to get ~2000-char excerpts around each match instead of the full (capped) OCR — useful for long documents. Required: document_id.
get_field_distribution(field, top_n, subset, country, date_to, keyword, subject, date_from, newspaper, over_time)- Rank the values of one multi-valued field across a filtered set — the direct way to answer 'which places does this coverage name most', 'who signs these articles', 'what subjects dominate'. Pipe-joined fields (subject, spatial, author, language, country) are split, so an article tagged 'Prière|Ramadan' counts once for each. Optional over_time adds the per-year share of items that carry ANY value for the field, which is how you see e.g. bylines appearing as the press professionalises. Required: field.
get_image(image_id)- Get one photograph by id: title, photographer, capture date, place and coordinates, subjects, rights, the IIIF manifest, and the full-resolution `image_url`. The server returns URLs, not image bytes. Required: image_id.
get_index_entry(entry_id)- Get full details of an index entry by id (raw dataset columns, French names — Titre, Prénom, Coordonnées…). Required: entry_id.
get_lexical_metrics(top_n, country, date_to, keyword, subject, group_by, date_from, newspaper)- Readability, lexical richness and length of the press text, averaged by year, newspaper or country. `Lisibilite_OCR` is a French readability score (higher = easier); `Richesse_Lexicale_OCR` is MATTR, a moving-average type-token ratio that is ALREADY length-robust — do not normalise it by word count or bin it by length. Readability is computed against a French lexicon, so non-French items are excluded from that metric (and counted in readability_excluded) rather than reported as unreadable; MATTR and word count need no lexicon and cover everything. Only items whose full text ships in this public dataset carry these columns at all.
get_newspaper_stats(country)- Per-newspaper article counts and date ranges.
get_place_distribution(top_n, subset, country, date_to, keyword, subject, date_from, newspaper)- Places named by a filtered set of items, joined to the index's authority records so each carries coordinates where the index has them. Use this rather than get_field_distribution when the question is geographic — where coverage clusters — and the plain ranking when it is not. Only `Lieux` index entries are geocoded (555 of 683); persons, organisations and events carry no coordinates and never will, and any named place with no index entry comes back under `ungeocoded` rather than being dropped.
get_publication_fulltext(keyword, max_excerpts, context_chars, publication_id)- Full OCR text of a publication, optionally returning ~2000-char excerpts around keyword matches (accent-insensitive; capped — see match_count vs excerpts_returned). Required: publication_id.
get_reference(reference_id)- Full bibliographic record for one academic reference (by id), including the complete abstract (present for ~51% of references), subjects, DOI/URL, and host-work details (book, volume, issue, pages). Required: reference_id.
get_semantic_map(limit, subset, country, date_to, keyword, subject, color_by, date_from, newspaper)- A 2-D scatter of a filtered set, projected from the stored 768-dimension embeddings by PCA. Shows which items sit near each other in meaning — where a set splits into distinct strands and where it is one cloud. Read `explained_variance` before drawing any conclusion: with 768 dimensions the first two components usually carry a modest share, and a scatter explaining 6% of the variance is a much weaker claim than one explaining 40%. This is PCA, not UMAP: it spreads the broadest axes of variation and flattens fine cluster structure, so it is not comparable to the semantic landscapes on islam.zmo.de. Needs no API key — the vectors are a column in the dataset — but only items whose full text ships are embedded at all. NOTE the payload scales with `limit`: a point cloud is a chart, not something a text-only client can read, so for those the useful part is the explained-variance summary rather than the coordinates. Keep `limit` low unless a chart is going to be drawn.
get_sentiment_distribution(model, country, subject, newspaper)- Aggregate AI polarity, centrality and subjectivity across a filter set. Three models scored the corpus independently — gpt-5-6-luna, mistral-small-2603, deepseek-v4-flash-0731 — so model:"all" returns each one's distribution plus how often they AGREE. Treat disagreement as a fact about the judgement rather than noise: in a set where the three models split on polarity, no single model's number should be quoted alone. All three scales are ordinal French labels; subjectivity is much the weakest and ships a caveat to quote with it. Articles were scored whether or not their full text ships, so these shares are not subject to the OCR coverage limit; compare scored_by_all against total_articles for the residual gap (the ~51 non-francophone articles are unscored by design).
get_similar_items(id, limit, subset, min_score)- The items nearest to a given one in meaning, by cosine similarity over the stored embeddings. Answers 'what else is like this' without a keyword — it finds pieces on the same event or theme that share no vocabulary. A neighbour above ~0.85 is usually the same story reprinted or lightly rewritten, which is how to spot syndication in this corpus; 0.6-0.8 is 'same subject, different piece'. Needs no API key: the item's own vector is a column, so nothing has to be embedded at request time. This is per-item, NOT the corpus-wide near-duplicate sweep — that is an all-pairs job and belongs offline. Required: id.
get_temporal_distribution(subset, country, date_to, keyword, subject, calendar, group_by, date_from, newspaper, granularity)- Counts of matching items per year (or month) — the direct way to chart coverage trends over time instead of paging through search results. Defaults to articles; also works on publications, references, documents, audiovisual, and images. Accepts the same filters as the corresponding search_* tool (keyword = ONE substring over the subset's text fields, country, newspaper/series, subject, date range). Optional group_by=country|newspaper returns one distribution per group. Items dated only to a year keep a bare-year key even at month granularity; undated items are counted in undated_count, never dropped silently. Set calendar=hijri to bucket by the Islamic (Umm al-Qura) calendar instead — with granularity=lunar_month this collapses every year into the twelve lunar months, which is the ONLY way to see observance-driven coverage (Ramadan, Dhu al-Hijja/hajj, Shawwal/Korité): the lunar year drifts ~11 days against the Gregorian, so a Gregorian axis smears each observance across all twelve months. Hijri buckets need a full YYYY-MM-DD, so items dated only to a year or month are reported in imprecise_date_count.
get_topic_distribution(top_n, subset, country, date_to, keyword, subject, min_prob, date_from, newspaper, over_time)- How a filtered set distributes across the precomputed LDA topics, each labelled by its top terms (articles carry 30 topics and are ~99.5% classified; references have their own 33-topic model and only ~46% carry an assignment, so read its `classified` against `total_matches`). Topics are assigned offline over the full text, so they describe what a piece is ABOUT rather than which words it contains — use this instead of keyword counting to map a corpus. Optional over_time returns per-year counts for the leading topics. min_prob keeps only articles where the topic is at least that dominant (mean assignment probability is 0.34, so 0.5 is already a strong filter).
list_audiovisual(limit, offset, country)- List audiovisual materials (Nigerian recordings, incl. Hausa/Arabic content).
list_locations(limit, offset, country)- List lieux from the IWAC index, sorted by frequency (most-referenced first). The optional 'country' filter selects entries that APPEAR IN records from that country (mentioned-in, not located-in), ranked by collection-wide 'frequency' — so foreign and cross-border entries can appear. Nigeria returns none here (index frequency is computed from articles + publications + references, which have no Nigerian items — Nigeria is audiovisual only).
list_periodicals(country)- List the Islamic periodical/series titles in the publications subset, with issue counts and year ranges. Use the returned newspaper value as the `newspaper` filter on search_publications.
list_persons(limit, offset, country)- List personnes from the IWAC index, sorted by frequency (most-referenced first). The optional 'country' filter selects entries that APPEAR IN records from that country (mentioned-in, not located-in), ranked by collection-wide 'frequency' — so foreign and cross-border entries can appear. Nigeria returns none here (index frequency is computed from articles + publications + references, which have no Nigerian items — Nigeria is audiovisual only).
list_subjects(limit, offset)- List sujets from the IWAC index, sorted by frequency (most-referenced first).
search(limit, query)- Search the Islam West Africa Collection across newspaper articles, Islamic publications, archival documents, academic references, audiovisual recordings, photographs, and the authority index (persons/places/organisations/events/subjects). Pass ONE concept or name — e.g. 'Tijaniyya', 'laïcité', 'Sheikh Gumi', 'pèlerinage'. Matching is accent- and case-insensitive; a multi-word query requires every word to appear somewhere in the item, so prefer a single concept per call. Write query strings and concept keywords in French for press/publication/document/index discovery even when the user's report language is not French. Academic references are multilingual, so try French and English title/abstract terms when relevant; metadata/filter labels remain French. Use the French transliteration of Islamic terms (Tabaski not 'Eid al-Adha', charia not 'sharia', Maouloud not 'Mawlid'). Returns {results:[{id,title,url,category}], ranking}; each result's `category` names its subset and the `ranking` field documents the ordering. Pass an id to `fetch` to read the full text. For filtered queries (by country, date, or newspaper) use the search_* tools instead. Required: query.
search_articles(limit, offset, country, date_to, keyword, subject, date_from, newspaper, hijri_year, hijri_month, with_description)- Search IWAC newspaper articles by keyword (title + OCR + AI abstracts, French and English), country, newspaper, subject, and date range. Use French concept keywords regardless of the user's report language. Matching is accent- and case-insensitive.
search_audiovisual(limit, medium, offset, country, keyword, subject, language)- Search audiovisual materials by keyword and metadata. Keyword matches title, creator, publisher, subject, spatial, language, source, and AI description where present.
search_by_sentiment(limit, offset, country, subject, polarity, centrality, subjectivity)- Filter articles by gpt-5-6-luna sentiment labels (accent/case-insensitive exact match). One model's reading, not a consensus — two other models scored the same articles and often disagree; get_sentiment_distribution with model:"all" shows by how much. `subjectivity` is much the weakest of the three scales, so treat a set selected on it as a lead to read rather than as a finding.
search_documents(limit, offset, country, keyword)- Search the small archival-documents subset (~26 items: Islamic association reports, flyers, project documents — mostly Burkina Faso). Use French concept keywords regardless of the user's report language. Most have OCR text and an AI description. Call with no arguments to list all.
search_images(limit, offset, country, creator, date_to, keyword, spatial, subject, date_from)- Search the IWAC photographs (30 items: mosques, radio stations, schools, signage and street scenes documented during fieldwork). Keyword matches title, creator, subject, place and the rare caption. Each result carries `image_url` (the full-resolution file), `coordinates` ('lat, lng' where known) and the canonical IWAC page. Call with no arguments to list all. Captions are almost never present, so prefer subject/place filters over keywords, or semantic_search_images when it is enabled.
search_index(limit, offset, keyword, index_type)- Search the IWAC authority index (persons, places, organisations, events, subjects) by name. Accent/case-insensitive. Required: keyword.
search_publications(limit, offset, country, date_to, keyword, subject, date_from, newspaper, hijri_year, hijri_month)- Search Islamic publications (periodical issues, books). `keyword` matches title, subject, table of contents, and full OCR text (TOC hits come back as matching_toc_entries); use French concept keywords regardless of the user's report language. Filter by newspaper/series, subject, country and year. Use list_periodicals to discover series titles, and get_publication_fulltext for keyword excerpts from a single issue.
search_references(limit, author, offset, country, date_to, keyword, subject, language, date_from, reference_type)- Search academic references (journal articles, book chapters, theses, books, reports) by keyword and metadata. `keyword` is a single substring match over title + abstract, so search ONE term per call (combined terms like 'pèlerinage Mecque' miss results). References are multilingual: try French and English title/abstract keywords when relevant; metadata/filter values such as `reference_type` and `language` use French labels. Results include a short abstract snippet — use get_reference for the full abstract and bibliographic detail.
Last successful function declaration observed on . Source: https://islam.zmo.de/mcp/. We list what the server declared; we do not call any of these functions.
Endpoint status observed on . Source: https://islam.zmo.de/mcp/.
Signals
These are separate measurements of different things. They are deliberately not combined into one score, because a popularity number that mixes website traffic with saves and stars cannot be checked or acted on.
| Signal | Value | What it measures | Window | Observed | Source |
|---|---|---|---|---|---|
| GitHub stars | 1 | Number of GitHub accounts that bookmarked this repository since it was created. It is a bookmark count, not installs, not active users and not quality. | cumulative, all time | GitHub | |
| Last commit | 2026-08-07 | Date of the most recent push to any branch. This is the strongest cheap indicator of whether the project is still maintained. | point in time | GitHub | |
| Open issues | 3 | Open issues plus open pull requests, as GitHub counts them together. A high number can mean an active project or an abandoned one. | as of fetch | GitHub | |
| Latest published version | 1.5.1 | Latest version string the maintainer published to the registry. | as of fetch | Model Context Protocol | |
| Registry record last updated | 2026-08-05 | When the registry record was last updated by its maintainer. | point in time | Model Context Protocol | |
| License | MIT | Licence GitHub detected in the repository. Detection can be wrong; the LICENSE file is authoritative. | as of fetch | GitHub | |
| First listed in the MCP Registry | 2026-08-05 | Date this server was first published to the official MCP Registry. Not a usage or quality measure. | point in time | Model Context Protocol | |
| repository status | active | The repository exists on GitHub and is not archived. This says nothing about how recently it was worked on. | as of fetch | GitHub | |
| mcp tools declared | 34 tools | Number of functions the server itself declared when asked to list them. This is what the server offers an agent, not a measure of how well any of them work. | as of probe | islam.zmo.de | |
| mcp endpoint status | ok | The server listed 34 functions when asked. | as of probe | islam.zmo.de |
Where to get it
Related, by what their authors tagged them
-
io.github.caelum29/calibre-mcp
— last commit 2026-08-03, shares claude-desktop, huggingface
Calibre ebook library server: search, read content, curate metadata, semantic search, gated writes.
-
TrainTools
— last commit 2026-07-22, shares huggingface
Recommend paper-backed diagnostics for PyTorch and Hugging Face training problems.
-
io.github.Embassy-of-the-Free-Mind/source-library
— last commit 2026-08-08, shares digital-humanities
Search 22,000+ rare pre-modern texts with AI English translations, summaries, and 73K+ images.
-
Source Library
— last commit 2026-08-08, shares digital-humanities
Search 15K rare pre-modern texts translated to English: philosophy, religion, science, literature.
-
Proton Drive CLI MCP
— last commit 2026-06-12, shares claude-desktop, mcpb
Manage Proton Drive from MCP clients through Proton's official CLI without exposing credentials.
-
app.reassign/reassign
— last commit 2026-07-25, shares mcpb
Reassign: a circular 24-hour calendar and time-tracking copilot with ADHD-friendly scheduling.
-
ContrastAPI
— last commit 2026-08-04, shares mcpb
55 tools, 7 Resources, Sigma rules, email SPF/DMARC, MITRE, CVE/KEV, risk_score. No key.
-
io.github.dizzlkheinz/ynab-mcpb
— last commit 2026-07-29, shares mcpb
Local-first YNAB MCP for reconciliation, receipt splitting, safe cleanup, and spending analysis
-
Help Scout MCP Server
— last commit 2026-08-01, shares mcpb
Search Help Scout conversations, customers, organizations, threads, and inboxes with AI assistants
-
ai.aliengiraffe/spotdb
— last commit 2026-08-05, shares duckdb
Ephemeral data sandbox for AI workflows with guardrails and security
These share tags the maintainers applied themselves, such as claude-desktop, huggingface, digital-humanities, mcpb. Common tags like "mcp" or "ai" are ignored for this: agreeing with six hundred other projects is not a similarity.
This is not a recommendation and not a test result. It is a map of what the authors said their work is about.
How the author describes it
Topics the maintainer set on GitHub: african-studies, claude-desktop, digital-archive, digital-humanities, duckdb, huggingface, islam, islamic-studies, mcp-server, mcpb, model-context-protocol, research-data, typescript, west-africa.
This record as data
Every field on this page, with its source and observation date, is in the catalog JSON. Fetch the whole kind at once instead of parsing this HTML.
GET /api/v1/entries/mcp_server.json