mcp server
CompletionKit
Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Description as published by the maintainer. Source
- version 1.0.0
- active
- evaluation
active — Most recent push to the repository was 2026-08-05. Dashed tags are derived by ZBS Index from the published description, not stated by the maintainer.
What this server can do
54 functions, named and described by the server itself. Parameter names are shown because they say more about what a function does than its name usually does.
agreements_create(note, run_id, verdict, metric_id, created_by, response_id, corrected_score)- Upsert an agreement for (run, response, metric, created_by). Verdict is one of agree, disagree, borderline. corrected_score (1..5) is required when verdict is 'disagree'. Required: run_id, response_id, metric_id, verdict.
agreements_list(run_id, metric_id, created_by, response_id)- List agreements. Filter by run_id, response_id, metric_id, or created_by.
datasets_create(name, csv_data, tag_names)- Create a dataset with CSV data. First row is the header. Two column names are recognized specially: "expected_output" is each row's answer key (ground truth) given to the judge and to checks that compare against the row's expected value, and "actual_output" is a pre-made output to score in a prompt-less run. Both are overridable per run (expected_column / output_column). Every column is also available to the prompt as a variable. Required: name, csv_data.
datasets_create_from_url(url, name, tag_names)- Create a dataset by downloading CSV from a URL instead of inlining it. Use this for large datasets: pass a public http(s) URL and the server fetches the CSV directly, so the data never has to pass through the tool-call arguments. The URL is SSRF-checked and the download is capped at 10MB. First row is the header; the "expected_output" (answer key) and "actual_output" (pre-made output) columns are recognized specially, overridable per run. Required: name, url.
datasets_delete(id)- Delete a dataset Required: id.
datasets_get(id)- Get a dataset by ID Required: id.
datasets_list- List all datasets
datasets_update(id, name, csv_data, tag_names)- Update a dataset Required: id.
judges_compare(metric_id, metric_version_a_id, metric_version_b_id)- Compare two versions of one metric's agreement stats side by side. Requires metric_id, metric_version_a_id, and metric_version_b_id (both versions must belong to that metric). Unavailable for check metrics. Required: metric_id, metric_version_a_id, metric_version_b_id.
judges_replay(name, metric_id, dataset_id, judge_model, output_column)- Create a scoring run for the current judge over a dataset's existing outputs (wraps runs_create with prompt_id omitted and output_column supplied). This only sets up the run; call runs_generate to actually re-judge the outputs so you can compare against human verdicts. Required: name, metric_id, dataset_id, judge_model.
metric_groups_create(name, tag_names, metric_ids, description)- Create a metric group Required: name.
metric_groups_delete(id)- Delete a metric group Required: id.
metric_groups_get(id)- Get a metric group by ID Required: id.
metric_groups_list- List all metric groups
metric_groups_update(id, name, tag_names, metric_ids, description)- Update a metric group Required: id.
metric_versions_dismiss(metric_version_id)- Destroy a draft MetricVersion (use for either source: 'edit' or source: 'suggestion'). Published versions are refused — to demote a published version, publish a different one as current instead. Required: metric_version_id.
metric_versions_list(metric_id)- List every MetricVersion (drafts + published) for a metric, newest first. Each row carries version_number, state, source, current flag, and timestamps. Required: metric_id.
metric_versions_publish(metric_version_id)- Publish a MetricVersion as the live version of its metric. Works for both 'draft → published' and 'revert to an older published version → current'. Transactionally flips current, demotes peers, and writes the version's instruction + rubric_bands back onto the metric so the judge grades against it. Required: metric_version_id.
metrics_create(name, tag_names, instruction, metric_type, check_config, rubric_bands)- Create a metric with evaluation criteria. For a deterministic check set metric_type:"check" and check_config. Per-kind required keys: value (contains/not_contains/equals), pattern (regex), json_path+expected (json_path_equals), min and/or max (length_bounds); valid_json takes no extra keys. target_path is required when target is json_path. For contains, not_contains, and equals, set compare_to:"expected" to grade against each row's own expected_output (ground truth) instead of a constant value (drop value); add expected_path to dig into the expected value when it is JSON. Required: name.
metrics_delete(id)- Delete a metric Required: id.
metrics_get(id)- Get a metric by ID Required: id.
metrics_list- List all metrics
metrics_suggest_variants(count, model, metric_id)- Ask the model to rewrite the metric's judge instruction in N variants targeted at the recent disagreements. Each variant is saved as a draft MetricVersion with source="suggestion". Returns the persisted drafts. Stripe-metering hooks fire via ActiveSupport::Notifications under completion_kit.judge_suggestion.generated. Required: metric_id.
metrics_update(id, name, tag_names, instruction, metric_type, check_config, rubric_bands)- Update a metric. For a deterministic check set metric_type:"check" and check_config. Per-kind required keys: value (contains/not_contains/equals), pattern (regex), json_path+expected (json_path_equals), min and/or max (length_bounds); valid_json takes no extra keys. target_path is required when target is json_path. For contains, not_contains, and equals, set compare_to:"expected" to grade against each row's own expected_output (ground truth) instead of a constant value (drop value); add expected_path to dig into the expected value when it is JSON. Required: id.
promptfoo_import(config)- Import a promptfooconfig.yaml. Creates a prompt, a dataset from the test vars, and metrics from the assert blocks (llm-rubric/g-eval become judge metrics; contains/equals/regex/is-json become deterministic check metrics). Returns a summary of what mapped and what was skipped and why; nothing is dropped silently. Required: config.
prompts_create(name, template, llm_model, tag_names, description)- Create a prompt Required: name, template, llm_model.
prompts_delete(id)- Delete a prompt Required: id.
prompts_get(id)- Get a prompt by ID Required: id.
prompts_list- List all prompts
prompts_publish(id)- Publish a prompt version, making it the current version Required: id.
prompts_suggest_improvement(run_id)- Suggest an improved version of a prompt, grounded in a run's test results and judge feedback. Analyzes the run's responses, scores, and reviews, then returns reasoning plus a rewritten template (preserving {{variables}}) and persists it as a Suggestion. Requires a run that has a prompt (not a scoring-only run). Required: run_id.
prompts_update(id, name, template, llm_model, tag_names, description)- Update a prompt. If the prompt already has runs, this creates a new DRAFT version (current=false) rather than editing in place or publishing — promote it with prompts_publish — so an agent's edits don't go live without a gate. If it has no runs, it is updated in place. Required: id.
provider_credentials_create(api_key, provider, api_version, api_endpoint)- Create a provider credential Required: provider, api_key.
provider_credentials_delete(id)- Delete a provider credential Required: id.
provider_credentials_get(id)- Get a provider credential by ID (API key is not exposed) Required: id.
provider_credentials_list- List all provider credentials (API keys are not exposed)
provider_credentials_update(id, api_key, provider, api_version, api_endpoint)- Update a provider credential Required: id.
responses_get(id, run_id)- Get a specific response Required: run_id, id.
responses_list(sort, limit, fields, offset, run_id, status, max_score, min_score)- List responses for a run, in row order. Returns {total, limit, offset, returned, responses}. Defaults to 50 rows because full payloads are large: use "fields" to drop the bodies, "min_score"/"max_score" to isolate low scorers, and sort "score_asc" to read the worst rows first. For per-metric averages of the whole run use runs_get instead of aggregating here. Required: run_id.
runs_create(name, prompt_id, tag_names, dataset_id, max_tokens, metric_ids, judge_model, temperature, output_column, expected_column, metric_group_id, judge_temperature)- Create a run. Omit prompt_id and provide output_column to score existing outputs by grading a pre-existing dataset column instead of generating new ones. Required: name.
runs_delete(id)- Delete a run Required: id.
runs_generate(id)- Start a run. Required for every run, including score-only runs (no prompt): generates responses with the prompt when there is one, otherwise copies the graded dataset column and grades it. Required: id.
runs_get(id)- Get a run by ID, including "metric_averages": a per-metric breakdown with each metric's average score (or pass rate for checks), how many rows it graded, and how many scored low. Use this to find the metric dragging a prompt down without listing responses. Required: id.
runs_list- List all runs
runs_regrade(id)- Re-grade a run's existing responses with its currently attached metrics, without regenerating. Use after attaching or editing metrics on an already-generated run. Required: id.
runs_rerun(id)- Create and start a fresh copy of a run with the same prompt, dataset, metrics, and settings. Use when the judge changed and you want a clean run instead of mixing versions. Required: id.
runs_retry_failures(id, only)- Re-run only the failed responses of a run, optionally limited to specific response ids via "only". Required: id.
runs_update(id, name, tag_names, dataset_id, max_tokens, metric_ids, judge_model, temperature, output_column, expected_column, metric_group_id, judge_temperature)- Update a run Required: id.
tags_create(name)- Create a tag. Color is auto-assigned. Required: name.
tags_delete(id)- Delete a tag. Removes the tag from every linked metric, prompt, run, and dataset. Required: id.
tags_get(id)- Get a tag by ID Required: id.
tags_list- List all tags
tags_update(id, name)- Rename a tag. Required: id.
usage_get- Get this organization's plan usage and limits for the current billing period: runs and prompt fetches used, their limits, how many remain, and when the period resets. Call this to pre-check quota before starting runs. Runs are hard-blocked once the run limit is reached (with a small grace band), so a run over the limit will fail with run_limit_reached.
Last successful function declaration observed on . Source: https://completionkit.com/mcp. We list what the server declared; we do not call any of these functions.
Endpoint status observed on . Source: https://completionkit.com/mcp.
Signals
These are separate measurements of different things. They are deliberately not combined into one score, because a popularity number that mixes website traffic with saves and stars cannot be checked or acted on.
| Signal | Value | What it measures | Window | Observed | Source |
|---|---|---|---|---|---|
| GitHub stars | 1 | Number of GitHub accounts that bookmarked this repository since it was created. It is a bookmark count, not installs, not active users and not quality. | cumulative, all time | GitHub | |
| Last commit | 2026-08-05 | Date of the most recent push to any branch. This is the strongest cheap indicator of whether the project is still maintained. | point in time | GitHub | |
| Open issues | 11 | Open issues plus open pull requests, as GitHub counts them together. A high number can mean an active project or an abandoned one. | as of fetch | GitHub | |
| Latest published version | 1.0.0 | Latest version string the maintainer published to the registry. | as of fetch | Model Context Protocol | |
| Registry record last updated | 2026-07-18 | When the registry record was last updated by its maintainer. | point in time | Model Context Protocol | |
| First listed in the MCP Registry | 2026-07-18 | Date this server was first published to the official MCP Registry. Not a usage or quality measure. | point in time | Model Context Protocol | |
| repository status | active | The repository exists on GitHub and is not archived. This says nothing about how recently it was worked on. | as of fetch | GitHub | |
| mcp tools declared | 54 tools | Number of functions the server itself declared when asked to list them. This is what the server offers an agent, not a measure of how well any of them work. | as of probe | completionkit.com | |
| mcp endpoint status | ok | The server listed 54 functions when asked. | as of probe | completionkit.com |
Where to get it
Related, by what their authors tagged them
-
PasteHTML
— last commit 2026-08-05, shares rails, ruby
Publish, update, and organize HTML and Markdown pastes on PasteHTML as the authorized user.
-
Lumen
— last commit 2026-06-07, shares llm-as-judge
Self-hostable agentic-AI LMS: catalog, RAG tutor, FSRS reviews, AI authoring, ingest.
-
Rails AI Context
— last commit 2026-08-01, shares rails, ruby
38 MCP tools give AI agents live Rails schema, routes, models, and conventions.
-
Citadel
— last commit 2026-08-06, shares llm-evaluation
Encrypted-first embedded database with vector search and agent memory, exposed as MCP tools
-
io.github.ahmedEid1/forgejudge
— last commit 2026-07-25, shares llm-evaluation
Open eval leaderboard + CI gate for autonomous coding agents (solve, score, trace).
-
ai.testiv/mcp
— last commit 2026-08-04, shares ruby
Local-first visual regression for AI agents: verdicts, diff images, explain_snapshot. No API key.
-
MCP GSC CPG
— last commit 2026-05-04, shares rails
A2A MCP & CPG Rails: HBPC (Hygiene + Beauty Personal Care) procurement and ESG.
-
llmtrim
— last commit 2026-08-05, shares llmops, openai, prompt-engineering
MCP server and proxy that compresses LLM prompts, tool output, and replies to cut token cost.
-
Distil
— last commit 2026-08-08, shares llmops, openai
Reversibly compress tool outputs to recoverable handles; expand to exact original bytes on demand.
-
Unified AI System
— last commit 2026-08-02, shares llmops, prompt-engineering
Terminal-first AI gateway with nine governed MCP tools for Codex, Cursor, Cline; no-key local path.
These share tags the maintainers applied themselves, such as rails, ruby, llm-as-judge, llm-evaluation. Common tags like "mcp" or "ai" are ignored for this: agreeing with six hundred other projects is not a similarity.
This is not a recommendation and not a test result. It is a map of what the authors said their work is about.
How the author describes it
Topics the maintainer set on GitHub: anthropic, evaluation-framework, evaluation-metrics, llm, llm-as-judge, llm-eval, llm-evaluation, llm-evaluation-framework, llm-evaluation-metrics, llmops, mcp, ollama, openai, prompt-engineering, prompt-testing, rails, rails-engine, ruby, ruby-on-rails.
This record as data
Every field on this page, with its source and observation date, is in the catalog JSON. Fetch the whole kind at once instead of parsing this HTML.
GET /api/v1/entries/mcp_server.json