ZBS Index What actually exists in applied AI, with the source next to it

mcp server

CompletionKit

Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.

Description as published by the maintainer. Source

  • version 1.0.0
  • active
  • evaluation

active — Most recent push to the repository was 2026-08-05. Dashed tags are derived by ZBS Index from the published description, not stated by the maintainer.

What this server can do

54 functions, named and described by the server itself. Parameter names are shown because they say more about what a function does than its name usually does.

agreements_create(note, run_id, verdict, metric_id, created_by, response_id, corrected_score)
Upsert an agreement for (run, response, metric, created_by). Verdict is one of agree, disagree, borderline. corrected_score (1..5) is required when verdict is 'disagree'. Required: run_id, response_id, metric_id, verdict.
agreements_list(run_id, metric_id, created_by, response_id)
List agreements. Filter by run_id, response_id, metric_id, or created_by.
datasets_create(name, csv_data, tag_names)
Create a dataset with CSV data. First row is the header. Two column names are recognized specially: "expected_output" is each row's answer key (ground truth) given to the judge and to checks that compare against the row's expected value, and "actual_output" is a pre-made output to score in a prompt-less run. Both are overridable per run (expected_column / output_column). Every column is also available to the prompt as a variable. Required: name, csv_data.
datasets_create_from_url(url, name, tag_names)
Create a dataset by downloading CSV from a URL instead of inlining it. Use this for large datasets: pass a public http(s) URL and the server fetches the CSV directly, so the data never has to pass through the tool-call arguments. The URL is SSRF-checked and the download is capped at 10MB. First row is the header; the "expected_output" (answer key) and "actual_output" (pre-made output) columns are recognized specially, overridable per run. Required: name, url.
datasets_delete(id)
Delete a dataset Required: id.
datasets_get(id)
Get a dataset by ID Required: id.
datasets_list
List all datasets
datasets_update(id, name, csv_data, tag_names)
Update a dataset Required: id.
judges_compare(metric_id, metric_version_a_id, metric_version_b_id)
Compare two versions of one metric's agreement stats side by side. Requires metric_id, metric_version_a_id, and metric_version_b_id (both versions must belong to that metric). Unavailable for check metrics. Required: metric_id, metric_version_a_id, metric_version_b_id.
judges_replay(name, metric_id, dataset_id, judge_model, output_column)
Create a scoring run for the current judge over a dataset's existing outputs (wraps runs_create with prompt_id omitted and output_column supplied). This only sets up the run; call runs_generate to actually re-judge the outputs so you can compare against human verdicts. Required: name, metric_id, dataset_id, judge_model.
metric_groups_create(name, tag_names, metric_ids, description)
Create a metric group Required: name.
metric_groups_delete(id)
Delete a metric group Required: id.
metric_groups_get(id)
Get a metric group by ID Required: id.
metric_groups_list
List all metric groups
metric_groups_update(id, name, tag_names, metric_ids, description)
Update a metric group Required: id.
metric_versions_dismiss(metric_version_id)
Destroy a draft MetricVersion (use for either source: 'edit' or source: 'suggestion'). Published versions are refused — to demote a published version, publish a different one as current instead. Required: metric_version_id.
metric_versions_list(metric_id)
List every MetricVersion (drafts + published) for a metric, newest first. Each row carries version_number, state, source, current flag, and timestamps. Required: metric_id.
metric_versions_publish(metric_version_id)
Publish a MetricVersion as the live version of its metric. Works for both 'draft → published' and 'revert to an older published version → current'. Transactionally flips current, demotes peers, and writes the version's instruction + rubric_bands back onto the metric so the judge grades against it. Required: metric_version_id.
metrics_create(name, tag_names, instruction, metric_type, check_config, rubric_bands)
Create a metric with evaluation criteria. For a deterministic check set metric_type:"check" and check_config. Per-kind required keys: value (contains/not_contains/equals), pattern (regex), json_path+expected (json_path_equals), min and/or max (length_bounds); valid_json takes no extra keys. target_path is required when target is json_path. For contains, not_contains, and equals, set compare_to:"expected" to grade against each row's own expected_output (ground truth) instead of a constant value (drop value); add expected_path to dig into the expected value when it is JSON. Required: name.
metrics_delete(id)
Delete a metric Required: id.
metrics_get(id)
Get a metric by ID Required: id.
metrics_list
List all metrics
metrics_suggest_variants(count, model, metric_id)
Ask the model to rewrite the metric's judge instruction in N variants targeted at the recent disagreements. Each variant is saved as a draft MetricVersion with source="suggestion". Returns the persisted drafts. Stripe-metering hooks fire via ActiveSupport::Notifications under completion_kit.judge_suggestion.generated. Required: metric_id.
metrics_update(id, name, tag_names, instruction, metric_type, check_config, rubric_bands)
Update a metric. For a deterministic check set metric_type:"check" and check_config. Per-kind required keys: value (contains/not_contains/equals), pattern (regex), json_path+expected (json_path_equals), min and/or max (length_bounds); valid_json takes no extra keys. target_path is required when target is json_path. For contains, not_contains, and equals, set compare_to:"expected" to grade against each row's own expected_output (ground truth) instead of a constant value (drop value); add expected_path to dig into the expected value when it is JSON. Required: id.
promptfoo_import(config)
Import a promptfooconfig.yaml. Creates a prompt, a dataset from the test vars, and metrics from the assert blocks (llm-rubric/g-eval become judge metrics; contains/equals/regex/is-json become deterministic check metrics). Returns a summary of what mapped and what was skipped and why; nothing is dropped silently. Required: config.
prompts_create(name, template, llm_model, tag_names, description)
Create a prompt Required: name, template, llm_model.
prompts_delete(id)
Delete a prompt Required: id.
prompts_get(id)
Get a prompt by ID Required: id.
prompts_list
List all prompts
prompts_publish(id)
Publish a prompt version, making it the current version Required: id.
prompts_suggest_improvement(run_id)
Suggest an improved version of a prompt, grounded in a run's test results and judge feedback. Analyzes the run's responses, scores, and reviews, then returns reasoning plus a rewritten template (preserving {{variables}}) and persists it as a Suggestion. Requires a run that has a prompt (not a scoring-only run). Required: run_id.
prompts_update(id, name, template, llm_model, tag_names, description)
Update a prompt. If the prompt already has runs, this creates a new DRAFT version (current=false) rather than editing in place or publishing — promote it with prompts_publish — so an agent's edits don't go live without a gate. If it has no runs, it is updated in place. Required: id.
provider_credentials_create(api_key, provider, api_version, api_endpoint)
Create a provider credential Required: provider, api_key.
provider_credentials_delete(id)
Delete a provider credential Required: id.
provider_credentials_get(id)
Get a provider credential by ID (API key is not exposed) Required: id.
provider_credentials_list
List all provider credentials (API keys are not exposed)
provider_credentials_update(id, api_key, provider, api_version, api_endpoint)
Update a provider credential Required: id.
responses_get(id, run_id)
Get a specific response Required: run_id, id.
responses_list(sort, limit, fields, offset, run_id, status, max_score, min_score)
List responses for a run, in row order. Returns {total, limit, offset, returned, responses}. Defaults to 50 rows because full payloads are large: use "fields" to drop the bodies, "min_score"/"max_score" to isolate low scorers, and sort "score_asc" to read the worst rows first. For per-metric averages of the whole run use runs_get instead of aggregating here. Required: run_id.
runs_create(name, prompt_id, tag_names, dataset_id, max_tokens, metric_ids, judge_model, temperature, output_column, expected_column, metric_group_id, judge_temperature)
Create a run. Omit prompt_id and provide output_column to score existing outputs by grading a pre-existing dataset column instead of generating new ones. Required: name.
runs_delete(id)
Delete a run Required: id.
runs_generate(id)
Start a run. Required for every run, including score-only runs (no prompt): generates responses with the prompt when there is one, otherwise copies the graded dataset column and grades it. Required: id.
runs_get(id)
Get a run by ID, including "metric_averages": a per-metric breakdown with each metric's average score (or pass rate for checks), how many rows it graded, and how many scored low. Use this to find the metric dragging a prompt down without listing responses. Required: id.
runs_list
List all runs
runs_regrade(id)
Re-grade a run's existing responses with its currently attached metrics, without regenerating. Use after attaching or editing metrics on an already-generated run. Required: id.
runs_rerun(id)
Create and start a fresh copy of a run with the same prompt, dataset, metrics, and settings. Use when the judge changed and you want a clean run instead of mixing versions. Required: id.
runs_retry_failures(id, only)
Re-run only the failed responses of a run, optionally limited to specific response ids via "only". Required: id.
runs_update(id, name, tag_names, dataset_id, max_tokens, metric_ids, judge_model, temperature, output_column, expected_column, metric_group_id, judge_temperature)
Update a run Required: id.
tags_create(name)
Create a tag. Color is auto-assigned. Required: name.
tags_delete(id)
Delete a tag. Removes the tag from every linked metric, prompt, run, and dataset. Required: id.
tags_get(id)
Get a tag by ID Required: id.
tags_list
List all tags
tags_update(id, name)
Rename a tag. Required: id.
usage_get
Get this organization's plan usage and limits for the current billing period: runs and prompt fetches used, their limits, how many remain, and when the period resets. Call this to pre-check quota before starting runs. Runs are hard-blocked once the run limit is reached (with a small grace band), so a run over the limit will fail with run_limit_reached.

Last successful function declaration observed on . Source: https://completionkit.com/mcp. We list what the server declared; we do not call any of these functions.

Endpoint status observed on . Source: https://completionkit.com/mcp.

Signals

These are separate measurements of different things. They are deliberately not combined into one score, because a popularity number that mixes website traffic with saves and stars cannot be checked or acted on.

Signal Value What it measures Window Observed Source
GitHub stars 1 Number of GitHub accounts that bookmarked this repository since it was created. It is a bookmark count, not installs, not active users and not quality. cumulative, all time GitHub
Last commit 2026-08-05 Date of the most recent push to any branch. This is the strongest cheap indicator of whether the project is still maintained. point in time GitHub
Open issues 11 Open issues plus open pull requests, as GitHub counts them together. A high number can mean an active project or an abandoned one. as of fetch GitHub
Latest published version 1.0.0 Latest version string the maintainer published to the registry. as of fetch Model Context Protocol
Registry record last updated 2026-07-18 When the registry record was last updated by its maintainer. point in time Model Context Protocol
First listed in the MCP Registry 2026-07-18 Date this server was first published to the official MCP Registry. Not a usage or quality measure. point in time Model Context Protocol
repository status active The repository exists on GitHub and is not archived. This says nothing about how recently it was worked on. as of fetch GitHub
mcp tools declared 54 tools Number of functions the server itself declared when asked to list them. This is what the server offers an agent, not a measure of how well any of them work. as of probe completionkit.com
mcp endpoint status ok The server listed 54 functions when asked. as of probe completionkit.com

Where to get it

Related, by what their authors tagged them

  • PasteHTML — last commit 2026-08-05, shares rails, ruby
    Publish, update, and organize HTML and Markdown pastes on PasteHTML as the authorized user.
  • Lumen — last commit 2026-06-07, shares llm-as-judge
    Self-hostable agentic-AI LMS: catalog, RAG tutor, FSRS reviews, AI authoring, ingest.
  • Rails AI Context — last commit 2026-08-01, shares rails, ruby
    38 MCP tools give AI agents live Rails schema, routes, models, and conventions.
  • Citadel — last commit 2026-08-06, shares llm-evaluation
    Encrypted-first embedded database with vector search and agent memory, exposed as MCP tools
  • io.github.ahmedEid1/forgejudge — last commit 2026-07-25, shares llm-evaluation
    Open eval leaderboard + CI gate for autonomous coding agents (solve, score, trace).
  • ai.testiv/mcp — last commit 2026-08-04, shares ruby
    Local-first visual regression for AI agents: verdicts, diff images, explain_snapshot. No API key.
  • MCP GSC CPG — last commit 2026-05-04, shares rails
    A2A MCP & CPG Rails: HBPC (Hygiene + Beauty Personal Care) procurement and ESG.
  • llmtrim — last commit 2026-08-05, shares llmops, openai, prompt-engineering
    MCP server and proxy that compresses LLM prompts, tool output, and replies to cut token cost.
  • Distil — last commit 2026-08-08, shares llmops, openai
    Reversibly compress tool outputs to recoverable handles; expand to exact original bytes on demand.
  • Unified AI System — last commit 2026-08-02, shares llmops, prompt-engineering
    Terminal-first AI gateway with nine governed MCP tools for Codex, Cursor, Cline; no-key local path.

These share tags the maintainers applied themselves, such as rails, ruby, llm-as-judge, llm-evaluation. Common tags like "mcp" or "ai" are ignored for this: agreeing with six hundred other projects is not a similarity.

This is not a recommendation and not a test result. It is a map of what the authors said their work is about.

How the author describes it

Topics the maintainer set on GitHub: anthropic, evaluation-framework, evaluation-metrics, llm, llm-as-judge, llm-eval, llm-evaluation, llm-evaluation-framework, llm-evaluation-metrics, llmops, mcp, ollama, openai, prompt-engineering, prompt-testing, rails, rails-engine, ruby, ruby-on-rails.

This record as data

Every field on this page, with its source and observation date, is in the catalog JSON. Fetch the whole kind at once instead of parsing this HTML.

GET /api/v1/entries/mcp_server.json

Sources

  1. homemade-software-inc/completion-kit on GitHub — GitHub, observed , trust tier 3.
  2. Tools declared by the MCP server at https://completionkit.com/mcp — completionkit.com, observed , trust tier 1.
  3. Official MCP Registry — Model Context Protocol, observed , trust tier 1.