mcp server
GoldenMatch
Find duplicate records in 30 seconds. Zero-config entity resolution, 97.2% F1 out of the box.
Description as published by the maintainer. Source
- version 3.5.0
- active
active — Most recent push to the repository was 2026-08-06.
What this server can do
77 functions, named and described by the server itself. Parameter names are shown because they say more about what a function does than its name usually does.
add_correction(id_a, id_b, path, reason, dataset, decision, cluster_id, field_name, matchkey_name, original_value, corrected_value)- Add a Learning Memory correction. Two shapes: - pair-level: decision='approve' or 'reject', requires id_a + id_b - field-level (v1.18.2+): decision='field_correct', requires cluster_id + field_name + corrected_value Source is 'agent' with trust=0.5 (lower than human steward 1.0). Pair (id_a, id_b) is canonicalized to (min, max) before storage. Required: decision, dataset.
agent_approve_reject(id_a, id_b, reason, decision, job_name, decided_by)- Approve or reject a review queue pair Required: job_name, id_a, id_b, decision, decided_by.
agent_compare_strategies(encoding, filename, file_path, file_content, ground_truth, ground_truth_name, ground_truth_content)- Compare ER strategies on your data
agent_deduplicate(config, encoding, filename, file_path, file_content, exclude_columns)- Run full ER pipeline with confidence gating and reasoning
agent_explain_cluster(cluster_id)- Explain why records are in the same cluster Required: cluster_id.
agent_explain_pair(exact, fuzzy, record_a, record_b)- Natural language explanation for a record pair Required: record_a, record_b.
agent_match_sources(config, file_a, file_b, encoding, file_a_name, file_b_name, file_a_content, file_b_content, exclude_columns)- Match two files with intelligent strategy selection
agent_review_queue(job_name)- Get borderline pairs awaiting approval Required: job_name.
analyze_blocking(limit, sample_size, target_block_size)- Diagnose blocking on the loaded dataset: returns ranked blocking key candidates with block counts, max block size, total candidate comparisons, and estimated recall. Use it to explain why matching is slow or produces too many candidate pairs.
analyze_data(encoding, filename, file_path, file_content)- Profile data, detect domain, recommend ER strategy
auto_configure(encoding, filename, file_path, constraints, file_content, exclude_columns)- Run AutoConfigController on a CSV; return the committed GoldenMatchConfig (incl. negative_evidence / Path Y when chosen) plus telemetry — stop_reason, health, decision trace, indicator column priors. Programmatic equivalent of `goldenmatch autoconfig`.
certify_recall(encoding, filename, file_path, file_content)- Estimate match RECALL without ground truth (unsupervised). Treats each auto-configured matchkey/pass as a decorrelated system and uses capture-recapture over their overlaps to estimate how many true matches were missed. Returns a point estimate (a safe lower bound additionally needs a small labelled audit; see `goldenmatch evaluate --certify --audit-out`). Needs >=3 decorrelated systems.
compare_clusters(encoding, clusters_a_name, clusters_a_path, clusters_b_name, clusters_b_path, clusters_a_content, clusters_b_content)- Compare two ER clustering outcomes on the same dataset without ground truth (CCMS): classifies each cluster as unchanged / merged / partitioned / overlapping and returns the Talburt-Wang Index. Both inputs are JSON cluster files (as written by export-style output).
config_weaknesses(phrasing, max_findings)- Diagnose weaknesses in the loaded run's auto-config: columns admitted that shouldn't be (source/provenance labels, per-row IDs), oversized or shared-value blocks, null sinks, low-signal matchkeys, and over-merging. Returns ranked findings, each with a plain-English explanation + a concrete fix, plus a one-paragraph summary.
controller_telemetry- Return the AutoConfigController telemetry from the most recent `auto_configure` or `agent_deduplicate` call in this MCP session. Same JSON shape as the web /api/v1/controller/telemetry endpoint.
create_domain(name, scope, signals, stop_words, brand_patterns, attribute_patterns, identifier_patterns)- Create a custom domain extraction rulebook. Define patterns for a specific data domain (medical devices, automotive parts, real estate, etc.). Required: name, signals.
dedupe(top_k, record)- Alias for `find_duplicates`. Find duplicate matches for a record. Provide field values to search against the loaded dataset. Required: record.
documents_ingest(model, paths, schema, backend, out_path, drop_empty)- Extract records from documents (PDF/image) against a target schema into rows ready for dedupe_df. Returns records + an ingest report. Required: paths, schema.
documents_suggest_schema(model, backend, sample_path)- Propose a target extraction schema (JSON) from a sample document image/PDF. Required: sample_path.
evaluate(col_a, col_b, ground_truth_path)- Score the loaded run against ground-truth pairs. Loads a ground-truth CSV (id_a,id_b columns) and returns precision, recall, and F1 for the current clustering. Required: ground_truth_path.
explain_cluster(cluster_id)- Alias for `agent_explain_cluster`. Explain why records are in the same cluster Required: cluster_id.
explain_match(record_a, record_b)- Explain why two records match or don't match. Shows per-field score breakdown. Required: record_a, record_b.
explain_pair(record_a, record_b)- Alias for `explain_match`. Explain why two records match or don't match. Shows per-field score breakdown. Required: record_a, record_b.
explain_routing(n_rows, cluster, driver_mem_gb, estimated_pair_count)- Human-readable explanation of why each stage is routed the way it is, with the driver-RAM projection that drove it. Required: n_rows, estimated_pair_count.
export_results(format, output_path)- Export matching results to a file (CSV or JSON). Required: output_path.
find_duplicates(top_k, record)- Find duplicate matches for a record. Provide field values to search against the loaded dataset. Required: record.
fix_quality(domain, encoding, filename, fix_mode, file_path, output_path, file_content)- Run GoldenCheck scan and apply fixes to a CSV file. Returns the fixed data summary and a manifest of all fixes applied. Requires goldencheck: pip install goldenmatch[quality]
get_cluster(cluster_id)- Get details of a specific cluster: all member records and their field values. Required: cluster_id.
get_golden_record(cluster_id)- Get the merged golden (canonical) record for a cluster. Required: cluster_id.
get_stats- Get dataset statistics: record count, cluster count, match rate, cluster sizes.
identity_audit(path, actor, limit, dataset)- Export the append-only identity audit log in commit order: every event with actor / trust / timestamp / reason, so a reviewer can reconstruct exactly which actor changed what, when, and why. Optionally filtered by dataset / actor.
identity_audit_seal(path, actor, dataset)- Anchor the append-only audit log with a tamper-evidence seal: a chained sha256 root over every event since the last seal. Cheap and idempotent (a no-op when nothing new has been logged). Run it periodically (or after a batch of stewardship actions) so the history becomes provably untampered. Optionally scoped to a dataset. Publish/mirror the returned root_hash to make tampering detectable by an external party.
identity_audit_verify(path, dataset)- Verify the append-only audit log against its seal chain. Replays the per-event content hashes and the seal roots to detect content edits, deletion, reordering, and insertion of any sealed event. Returns {ok, events_checked, seals_checked} plus the ids of any content mismatches / broken seals / missing sealed events. Optionally scoped to a dataset.
identity_claim(path, actor, trust, reason, entity_id, record_id)- Claim a record into an identity, moving it out of any prior entity ('this record belongs to that identity'). Emits a provenance-stamped `claimed` event on both the gaining and losing entities. Required: entity_id, record_id.
identity_conflicts(path, dataset)- List evidence edges marked `conflicts_with`.
identity_history(path, limit, entity_id)- Return the temporal event log for an identity. Required: entity_id.
identity_list(path, limit, offset, status, dataset)- List identities, optionally filtered by dataset/status.
identity_merge(path, actor, trust, reason, keep_entity_id, absorb_entity_id)- Manually merge two identities. All records from `absorb_entity_id` are reassigned to `keep_entity_id`. The merge events are stamped with `actor`/`trust` provenance so the audit log records who merged these and on what authority. Required: keep_entity_id, absorb_entity_id.
identity_profile(path, entity_id)- MDM profile of one entity: record count + per-source breakdown, golden record, confidence, conflict count, canonical version (structural-event count), and first/last activity. Returns {found: false} when no such entity exists. Required: entity_id.
identity_resolve(path, record_id)- Resolve a record_id to its durable identity. Returns the full identity view (members, evidence edges, recent events) or null when no identity exists for that record. Required: record_id.
identity_resolve_conflict(path, actor, apply, trust, reason, dataset, resolution, record_a_id, record_b_id)- Adjudicate a `conflicts_with` pair: 'same' keeps the entity intact, 'distinct' splits the second record out into a new identity, 'defer' only logs. Records a durable mediation verdict + event with actor/trust provenance, and stops the conflict re-surfacing in the open-conflicts queue. Required: record_a_id, record_b_id, resolution.
identity_show(path, entity_id, event_limit)- Fetch the full detail of one identity by entity_id: its member records, evidence edges, and recent event log. Returns {found: false} when no such entity exists. Required: entity_id.
identity_split(path, actor, trust, reason, entity_id, record_ids)- Split a subset of records off an identity into a brand-new identity. The original keeps the remaining records. The split events carry `actor`/`trust` provenance. Required: entity_id, record_ids.
identity_stats(path, dataset)- Graph-level summary / health stats: entities by status, total records, records-per-entity distribution, conflict total, source mix, and the largest entities. Optionally scoped to a dataset.
identity_worklist(path, limit, dataset, weak_confidence)- Prioritized steward worklist: active entities needing attention (open conflicts and/or confidence below weak_confidence), highest conflict count first.
incremental(config, encoding, base_file, threshold, new_records, base_file_name, new_records_name, base_file_content, new_records_content)- Match a batch of new records against an existing base dataset (without re-running the whole base). Returns matched (new_row_id, base_row_id, score) pairs plus counts. Auto-configures from the base file if no config is given.
learn_thresholds(path, matchkey_name)- Force a MemoryLearner pass over accumulated corrections. Returns the list of LearnedAdjustments produced (matchkey_name, threshold, sample_size, learned_at). Requires >= 10 corrections per matchkey before threshold tuning fires; otherwise returns an empty list.
lineage(max_pairs, output_dir, natural_language)- Field-level provenance for the loaded run: for each scored pair, the per-field scores that produced the match, plus cluster id. Optionally write a lineage JSON to a directory.
lint_routing(env, n_rows, cluster, driver_mem_gb, estimated_pair_count)- Flag config/env overrides that force a slow path (e.g. CLUSTERING_THRESHOLD=0 when the edge set fits driver RAM). ERROR at scale; would_refuse mirrors the runtime guard. Required: n_rows, estimated_pair_count.
list_clusters(limit, min_size)- List duplicate clusters found in the dataset. Returns cluster IDs, sizes, and member counts.
list_corrections(path, dataset)- List stored Learning Memory corrections, optionally filtered by dataset. Returns id_a, id_b, decision, source, trust, reason, matchkey_name, dataset, original_score, created_at.
list_domains- List available domain extraction rulebooks (built-in + user-defined).
list_plugins(category)- List all registered goldenmatch plugins by category. Includes the 22 v1.18.2 predefined plugins (numeric/format/business/aggregation) plus any user-registered plugins via entry-points or PluginRegistry.register_*(). Each entry includes name, source (builtin or user), category, and the first line of the merge docstring.
list_runs(output_dir)- List previous dedupe/match runs (for rollback) from the run log.
match(top_k, record, threshold)- Alias for `match_record`. Match a single record against the loaded dataset in real-time. Paste a record's fields and instantly see if it matches any existing record. Uses the configured matchkeys, scorers, and thresholds. Example: {"name": "John Smith", "email": "john@test.com", "zip": "10001"} Required: record.
match_record(top_k, record, threshold)- Match a single record against the loaded dataset in real-time. Paste a record's fields and instantly see if it matches any existing record. Uses the configured matchkeys, scorers, and thresholds. Example: {"name": "John Smith", "email": "john@test.com", "zip": "10001"} Required: record.
memory_export(path, dataset)- Return all corrections as a list of dicts (CSV-shaped). Caller is responsible for writing the file. Optionally filter by dataset.
memory_import(path, corrections)- Import corrections from a list of dicts (the exact shape memory_export returns). Upserts into the store: higher trust wins, same trust = latest wins. Returns the count imported. Required: corrections.
memory_stats(path)- Return Learning Memory status: total correction count, last learn time, and current learned adjustments. Cheap; safe for status checks.
plan_routing(n_rows, cluster, driver_mem_gb, estimated_pair_count)- Project per-stage distributed routing (scoring/clustering/golden) for a given data shape + cluster. Pure; no controller run. Required: n_rows, estimated_pair_count.
pprl_auto_config(use_llm, security_level)- Analyze the loaded dataset and recommend optimal PPRL (privacy-preserving record linkage) configuration. Returns recommended fields, bloom filter parameters, threshold, and explanation.
pprl_link(fields, file_a, file_b, encoding, threshold, file_a_name, file_b_name, file_a_content, file_b_content, security_level)- Run privacy-preserving record linkage between two parties' data. Computes bloom filters, matches records without sharing raw data. Specify fields, threshold, and security level. Required: fields.
profile- Alias for `profile_data`. Get data quality profile: column types, null rates, unique counts, sample values.
profile_data- Get data quality profile: column types, null rates, unique counts, sample values.
retrieve_similar(k, model, query, column, filters, encoding, filename, file_path, threshold, file_content)- Semantic retrieval (#1089): return the records in a CSV most similar to a free-text query, ranked by cosine similarity. Embeds the chosen column and the query with the zero-config in-house embedder (no cloud/torch by default) and runs ANN search. The read side of the RAG entity-canonicalization epic -- fetch candidate records by query without running a full dedupe. Required: query, column.
review_config- Run the config healer over the loaded dataset: analyze the dedupe run and return ranked, self-verified suggestions for improving the matching config (thresholds, scorers, negative evidence, blocking). Each suggestion carries an id, kind, target, rationale, and a machine-applicable patch. Requires the native kernel (pip install goldenmatch[native]); returns an empty list otherwise.
rollback(run_id, output_dir)- Undo a previous run by DELETING its output files (looked up by run_id in the run log). Destructive: removes the files that run wrote. Use list_runs first to find the run_id. Required: run_id.
run_transforms(encoding, filename, file_path, output_path, file_content)- Run GoldenFlow data transforms on a CSV file. Normalizes phone numbers (E.164), dates (ISO), categorical spelling, and Unicode issues. Returns a manifest of transforms applied. Requires goldenflow: pip install goldenmatch[transform]
scan_quality(domain, encoding, filename, file_path, file_content)- Run GoldenCheck data quality scan on a CSV file. Returns issues found (encoding errors, Unicode problems, format violations) without applying fixes. Requires goldencheck: pip install goldenmatch[quality]
schema_match(file_a, file_b, encoding, min_score, file_a_name, file_b_name, file_a_content, file_b_content)- Auto-map columns between two files with different schemas. Returns proposed (col_a, col_b) mappings with a confidence score and method (synonym / name_sim / composite). Useful before matching two sources.
sensitivity(sweep, config, encoding, filename, file_path, sample_size, file_content)- Parameter-sensitivity analysis: sweep one or more config parameters across a range and report how stable the clustering is at each value (CCMS unchanged %). Use it to find robust thresholds. Auto-configures the file if no config is given. Required: sweep.
shatter_cluster(cluster_id)- Break an entire cluster into individual records. All members become singletons. Use when a cluster is completely wrong. Required: cluster_id.
suggest_config(bad_merges)- Analyze bad merges and suggest config changes. Provide examples of incorrect merges (pairs that should NOT have matched) and GoldenMatch will identify which fields/thresholds to tighten. Example: [{"record_a": {...}, "record_b": {...}, "reason": "different people"}] Required: bad_merges.
suggest_pprl(encoding, filename, file_path, file_content)- Check if data needs privacy-preserving matching
test_domain(domain_name, sample_size)- Test a domain extraction rulebook against sample records. Shows what features would be extracted from the loaded data. Required: domain_name.
unmerge_record(record_id)- Remove a record from its cluster. The record becomes a singleton. Remaining cluster members are re-clustered using stored pair scores. Use this to fix bad merges. Required: record_id.
upload_dataset(encoding, filename, file_content)- Upload a local file's bytes to the server and get back a server-side path to reuse across other tools (analyze_data, auto_configure, agent_deduplicate, ...). No hosting needed. Send base64 (default) or raw text via `encoding`. Uploaded files are ephemeral scratch, reaped after GOLDENMATCH_MCP_UPLOAD_TTL (default 24h); re-upload if you need a path older than that. Max size GOLDENMATCH_MCP_MAX_UPLOAD_BYTES (default 64MB) -- above it, pass a public http(s) URL as file_path instead. Required: file_content, filename.
Last successful function declaration observed on . Source: https://goldenmatch-mcp-production.up.railway.app/mcp/. We list what the server declared; we do not call any of these functions.
Endpoint status observed on . Source: https://goldenmatch-mcp-production.up.railway.app/mcp/.
Signals
These are separate measurements of different things. They are deliberately not combined into one score, because a popularity number that mixes website traffic with saves and stars cannot be checked or acted on.
| Signal | Value | What it measures | Window | Observed | Source |
|---|---|---|---|---|---|
| GitHub stars | 128 | Number of GitHub accounts that bookmarked this repository since it was created. It is a bookmark count, not installs, not active users and not quality. | cumulative, all time | GitHub | |
| Last commit | 2026-08-06 | Date of the most recent push to any branch. This is the strongest cheap indicator of whether the project is still maintained. | point in time | GitHub | |
| Open issues | 11 | Open issues plus open pull requests, as GitHub counts them together. A high number can mean an active project or an abandoned one. | as of fetch | GitHub | |
| Latest published version | 3.5.0 | Latest version string the maintainer published to the registry. | as of fetch | Model Context Protocol | |
| Registry record last updated | 2026-07-18 | When the registry record was last updated by its maintainer. | point in time | Model Context Protocol | |
| License | MIT | Licence GitHub detected in the repository. Detection can be wrong; the LICENSE file is authoritative. | as of fetch | GitHub | |
| First listed in the MCP Registry | 2026-07-18 | Date this server was first published to the official MCP Registry. Not a usage or quality measure. | point in time | Model Context Protocol | |
| repository status | active | The repository exists on GitHub and is not archived. This says nothing about how recently it was worked on. | as of fetch | GitHub | |
| mcp tools declared | 77 tools | Number of functions the server itself declared when asked to list them. This is what the server offers an agent, not a measure of how well any of them work. | as of probe | goldenmatch-mcp-production.up.railway.app | |
| mcp endpoint status | ok | The server listed 77 functions when asked. | as of probe | goldenmatch-mcp-production.up.railway.app |
Where to get it
Related, by what their authors tagged them
-
GoldenAnalysis
— last commit 2026-08-06, shares data-cleaning, data-engineering, data-matching
Read-only cross-cutting analysis, metrics, and reporting across the Golden Suite.
-
GoldenCheck
— archived, last commit 2026-05-01, shares data-engineering, data-quality, polars
Auto-discover validation rules from data — scan, profile, health-score. No rules to write.
-
GoldenFlow
— archived, last commit 2026-05-01, shares data-cleaning, data-engineering, data-quality
Standardize, reshape, and normalize messy data — CSV, Excel, Parquet, S3, databases.
-
InferMap
— archived, last commit 2026-05-01, shares data-engineering, data-quality, fuzzy-matching
Map messy columns to a known schema — 7 scorers, domain dictionaries, F1 0.84. Zero config.
-
io.github.GCTRL-TECH/gctrl
— last commit 2026-08-06, shares entity-resolution, knowledge-graph
Governed graph-native agent memory: knowledge extraction, fusion, hybrid RAG, scoped access tokens.
-
io.github.GeiserX/duplicacy-mcp
— last commit 2026-08-01, shares deduplication
MCP server for Duplicacy — monitor backup status and Prometheus metrics
-
GoldenPipe
— archived, last commit 2026-05-01, shares data-engineering, data-quality, polars
One command to validate, transform, and deduplicate — chain GoldenCheck + Flow + Match.
-
datadiffer
— last commit 2026-07-28, shares data-engineering, data-quality
Read-only table diffs with segment attribution: which slice of your data changed.
-
TrustyData
— last commit 2026-07-13, shares data-quality
French address quality, geocoding & routing from official data (BAN, INSEE, OpenStreetMap).
-
io.github.agenson-horrowitz/agent-output-guard
— last commit 2026-04-05, shares data-quality
Validate and verify data from other agents before acting on it. Zero LLM costs.
These share tags the maintainers applied themselves, such as data-cleaning, data-engineering, data-matching, data-quality. Common tags like "mcp" or "ai" are ignored for this: agreeing with six hundred other projects is not a similarity.
This is not a recommendation and not a test result. It is a map of what the authors said their work is about.
How the author describes it
Topics the maintainer set on GitHub: data-cleaning, data-engineering, data-matching, data-quality, deduplication, entity-resolution, fellegi-sunter, fuzzy-matching, knowledge-graph, llm, master-data-management, mcp-server, polars, pprl, python, record-linkage, rust, splink, typescript, zero-config.
This record as data
Every field on this page, with its source and observation date, is in the catalog JSON. Fetch the whole kind at once instead of parsing this HTML.
GET /api/v1/entries/mcp_server.json