ZBS Index What actually exists in applied AI, with the source next to it

mcp server

io.github.cyanheads/evals-mcp-server

Author verifiable eval records through a draft→review→revise→submit loop with enforced graders.

Description as published by the maintainer. Source

  • version 0.1.2
  • active
  • evaluation

active — Most recent push to the repository was 2026-07-30. Dashed tags are derived by ZBS Index from the published description, not stated by the maintainer.

Signals

These are separate measurements of different things. They are deliberately not combined into one score, because a popularity number that mixes website traffic with saves and stars cannot be checked or acted on.

Signal Value What it measures Window Observed Source
GitHub stars 1 Number of GitHub accounts that bookmarked this repository since it was created. It is a bookmark count, not installs, not active users and not quality. cumulative, all time GitHub
Last commit 2026-07-30 Date of the most recent push to any branch. This is the strongest cheap indicator of whether the project is still maintained. point in time GitHub
Open issues 3 Open issues plus open pull requests, as GitHub counts them together. A high number can mean an active project or an abandoned one. as of fetch GitHub
Latest published version 0.1.2 Latest version string the maintainer published to the registry. as of fetch Model Context Protocol
Registry record last updated 2026-06-28 When the registry record was last updated by its maintainer. point in time Model Context Protocol
License Apache-2.0 Licence GitHub detected in the repository. Detection can be wrong; the LICENSE file is authoritative. as of fetch GitHub
First listed in the MCP Registry 2026-06-28 Date this server was first published to the official MCP Registry. Not a usage or quality measure. point in time Model Context Protocol
repository status active The repository exists on GitHub and is not archived. This says nothing about how recently it was worked on. as of fetch GitHub

Where to get it

Related, by what their authors tagged them

  • io.github.ahmedEid1/forgejudge — last commit 2026-07-25, shares evaluation
    Open eval leaderboard + CI gate for autonomous coding agents (solve, score, trace).
  • io.github.archonics/mcp-audit — last commit 2026-05-14, shares evaluation
    Free context-engineering audits for AI agents. BYOK Anthropic key. Top-3 findings per scan.
  • AgentTrust — last commit 2026-04-27, shares evaluation
    Quality verification for AI agents and MCP servers. 6-axis scoring, adversarial probes.
  • io.github.botzrDev/dreamd — last commit 2026-08-05, shares evaluation
    Local-first, cross-harness memory for AI coding agents.
  • com.clauxel.safetyreplay/safetyreplay-mcp — last commit 2026-05-19, shares evals
    Paid remote MCP for safety replays, policy gates, eval receipts, and release evidence.
  • Lumen — last commit 2026-06-07, shares evals
    Self-hostable agentic-AI LMS: catalog, RAG tutor, FSRS reviews, AI authoring, ingest.
  • ControlKeel — last commit 2026-08-06, shares evals
    Governed MCP workflows with policy validation, findings tracking, and review gates.
  • io.github.auxiliar-ai/auxiliar-mcp — archived, last commit 2026-07-11, shares evals
    Eval-backed discovery for the auxiliar.ai gateway — the best web-access provider per job, measured.
  • Palimpsest — censorship and model-eval observatory — last commit 2026-08-06, shares evals
    Live internet-censorship signals and tamper-evident, pre-registered, hash-chained AI model evals.
  • io.github.cyanheads/anime-mcp-server — last commit 2026-07-30, shares cyanheads
    Search anime/manga, franchise watch order, schedule, characters, rankings, studio filmography.

These share tags the maintainers applied themselves, such as evaluation, evals, cyanheads. Common tags like "mcp" or "ai" are ignored for this: agreeing with six hundred other projects is not a similarity.

This is not a recommendation and not a test result. It is a map of what the authors said their work is about.

Also from cyanheads

How the author describes it

Topics the maintainer set on GitHub: cyanheads, evals, evaluation, grader, inspect-ai, lm-eval-harness, mcp, model-context-protocol, rlvr, typescript, verifiable-rewards.

Bring your own setup

We take apart real AI setups every week and show what broke, what cost too much, and what the trace actually said. If you run agents on real work, that is where the useful conversation is.

Join ZBS AI Practice Lab

Sources

  1. cyanheads/evals-mcp-server on GitHub — GitHub, observed , trust tier 3.
  2. Official MCP Registry — Model Context Protocol, observed , trust tier 1.