Evaluation
Everything here was classified as evaluation by keyword match against the maintainer's own description, so treat the grouping as a starting point rather than a verdict.
The list is ordered by the most recent commit, not by stars. A popular project that stopped in 2024 is not a better answer than a smaller one shipped last week.
Listed, not yet verified by us (60)
Published to the registry, but we have not yet checked its repository. Treat the entry as the maintainer’s claim only.
-
OpenClaw Consensus
— v0.1.1
9-LLM consensus + disagreement scoring + cheapest-route picks to fight hallucinations.
-
PaperBanana-CN
— v2.0.1
Generate and evaluate academic diagrams and plots with independent model connections.
-
Mixpeek
— v0.1.1
Give your agent eyes. Search video, images, and audio via the Mixpeek multimodal retrieval API.
-
karst
— v0.2.9
Code context for AI dev tools: pack-scoped, cited code retrieval over MCP.
-
io.github.Morous-Dev/engram-cc
— v0.1.4
Universal AI coding assistant memory — session handoff, SLM compression, and semantic retrieval.
-
AgentVault
— v0.2.1
Discover, verify, and connect to AI agents with E2E encrypted messaging and trust scoring.
-
io.github.naur9n/arbitrum-transaction-preflight
— v4.0.0
Paid Arbitrum transaction simulation and risk scoring for wallets and AI agents.
-
QueryPilot
— v0.1.1
Safe SQL gateway for AI agents: SELECT-only validation, access policies, masking, audit trail, evals
-
io.github.Nicodemus941/health-api
— v1.0.6
Clinical AI for lab analysis, biomarker scoring, health narratives, and risk insights.
-
io.github.NikitaDatar/vectordecisions-mcp-server
— v1.0.3
AI Governance MCP Server - GATRI trust scoring, kill-switch, EU AI Act compliance for Claude
-
io.github.nipunkhanderia/golden-dataset-mcp
— v0.1.2
Version-controlled golden datasets and RAG evaluation, no API key needed.
-
io.github.noahgift/ruchy-mcp
— v3.67.0
MCP server for Ruchy: code analysis, scoring, linting, formatting, and transpilation tools
-
VulnFeed
— v0.3.7
Dependency vulnerability scanner with EPSS scoring. 9 MCP tools. Free tier + x402.
-
io.github.nqzai/kakunin
— v0.2.3
X.509 identity, risk scoring, and audit logging for AI agents. MiCA + EU AI Act compliant.
-
io.github.nuance-dev/rival
— v1.0.2
Query AI model benchmarks, pricing, and comparisons from rival.tips
-
Pensiata - Bulgarian Pension Fund Analytics
— v1.0.0
Bulgarian pension fund analytics — NAV data, metrics, rankings, and benchmarks.
-
io.github.omarkeshk/council-ai
— v1.0.0
Multi-LLM council MCP: parallel frontier models, consensus scoring, verdict-first code review
-
io.github.omega-memory/core
— v0.9.3
Persistent memory for AI coding agents. #1 on LongMemEval. Local-first.
-
io.github.opcastil11/mcp-prowl
— v0.1.1
Discover, evaluate, and compare SaaS APIs. Agent-verified scores and metrics.
-
OpenTrain
— v0.2.2
Hire human data labelers, RLHF annotators, and evaluators from your coding agent.
-
Serena MCP: the IDE for your agent
— v1.5.3
A powerful toolkit for coding, providing semantic retrieval and editing capabilities.
-
io.github.OtherVibes/mcp-as-a-judge
— v0.3.20
MCP as a Judge: a behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations
-
io.github.Pak209/alphalabs-intelligence
— v1.0.0
Live trading-pipeline intelligence for AI agents: signal scoring, calibration, recorded outcomes.
-
io.github.parallelromb/smara
— v2.0.2
Persistent cross-platform memory for AI agents with Ebbinghaus decay scoring.
-
io.github.penny4nonsense/mcp-scholaris
— v0.1.4
Academic paper search and retrieval via arXiv, Semantic Scholar, PubMed, Unpaywall.
-
Commit — Supply Chain Risk Scoring
— v1.36.0
Supply chain risk scoring for npm, PyPI, Cargo, and Go. 9 tools. Behavioral signals.
-
Congressional Documents
— v0.1.0
Congressional Documents — full-text search and retrieval over the official
-
Dart Kr
— v0.1.0
DART — Korea's Data Analysis, Retrieval and Transfer System.
-
Data Centrevaldeloire
— v0.1.0
Centre-Val de Loire Open Data (data.centrevaldeloire.fr) — OpenDataSoft MCP.
-
Epss
— v0.1.0
FIRST.org EPSS (Exploit Prediction Scoring System) MCP.
-
Exa
— v0.1.0
Exa MCP — neural/semantic web search + content retrieval (exa.ai)
-
Lichess
— v0.1.0
Lichess public API: players, ratings, eval, tablebase, opening explorer
-
Solitaire for Agents
— v1.5.2
Identity infrastructure for AI agents. Evolving persona, session continuity, self-tuning retrieval.
-
io.github.PrinceGabriel-lgtm/freshcontext
— v0.3.17
Freshness-aware AI retrieval with 21 MCP tools for timestamped, decay-ranked live signals.
-
Barevalue
— v1.0.3
Submit podcast orders, check status, and manage webhooks via Barevalue editing API.
-
AgentTrust
— v0.1.5
MCP Server for Agent Reputation & Trust Scoring
-
io.github.Rakesh1002/reposcout
— v0.1.1
Discover, enrich, and rank GitHub repos for an objective with weighted, evidence-based scoring.
-
io.github.rascal-3/chainanalyzer-mcp
— v1.1.1
Multi-chain AML risk scoring, sanctions screening, and tx tracing across BTC, ETH, POL, AVAX, SOL.
-
Onto
— v1.5.0
Clean Markdown and AI-readability scoring for any URL. Built for AI agents.
-
io.github.rchanllc/picdefenseio-mcp-server
— v1.0.1
Image risk scoring, EXIF, reverse-image backlinks, and image content detection via PicDefense.io.
-
Realmint
— v1.0.0
Agent-native scoring, search and routing for tokenized real-world assets across multi-chain.
-
RecourseOS
— v0.1.12
Consequence evaluation for AI agents. Check recoverability before destructive actions.
-
Rival Regulatory Toolkit
— v0.1.5
Read-only regulatory source retrieval, citation tracing, authority metadata, and recent filings.
-
DeepSeek FR MCP
— v0.1.0
Independent French workspace for evaluating DeepSeek-V3-class chat and DeepSeek-R1 reasoning.
-
io.github.rog0x/perf
— v1.0.2
Benchmark, memory, Big O analysis for AI agents
-
io.github.rogertheunissenmerge-oss/mcp-server
— v1.1.2
Wine pairing intelligence with 7 tools. Deterministic scoring, sommelier-calibrated.
-
RTFM
— v0.8.0
The open retrieval layer for AI agents — index code, docs, data. Search via MCP.
-
io.github.Rswcf/deepviews
— v1.0.1
Free financial data: company analysis, DCF, comps, benchmarks, screening. No API key.
-
operant-mcp
— v0.1.0
Read-only MCP server for the OPERANT AI operating-agent calibration benchmark.
-
io.github.salted-butter-joshua/dify-mcp
— v0.1.0
MCP server exposing Dify knowledge base retrieval to Cursor and other MCP clients
-
io.github.sathergate/searchcraft
— v0.1.1
Full-text search for Next.js. BM25 scoring and fuzzy matching, no external service.
-
ScoutScore
— v0.1.2
Trust scoring for AI agents. Check scores, fidelity, and flags for x402 services.
-
io.github.sebastienrousseau/iso20022-readiness-suite-mcp
— v0.0.2
Orchestration MCP server: ISO 20022 readiness scoring, remediation, bank-response simulation.
-
io.github.sedis-ab/mcp
— v1.1.5
Read-only MCP tools for the Sedis PartnerAPI v2 (Bolagsanalys + Fastighetsbenchmark).
-
io.github.sendblue-api/sendblue-browser-mcp
— v0.2.3
Drive a stealth-patched Chromium daemon: navigate, screenshot, eval JS, attach over CDP.
-
Sevalla
— v1.0.0
Official Sevalla MCP — full PaaS API access through just 2 tools.
-
io.github.sharozdawa/content-optimizer
— v1.0.0
SERP-based content scoring and optimization with 7 SEO categories.
-
io.github.shea256/aiiq-mcp
— v0.2.0
Query AI IQ (aiiq.org) model IQ, rankings, benchmarks, and methodology.
-
Axcess — Design Accessibility Evaluation
— v0.2.0
Evaluates UI designs for WCAG accessibility issues automated scanners miss. Paid via x402 on Base.
-
Global Talent AI — UK Global Talent Visa studio
— v2.3.0
UK Global Talent Visa handbook, endorser criteria, cost calculator, draft scoring. Free, no auth.
Page 5 of 7
How this page is ordered
Entries are grouped by whether anyone is still working on them, using the date of the most recent push to the repository. They are not ordered by stars, because a star is a bookmark somebody left once and never took back.
Where we have not checked an entry yet, it says so rather than being mixed in with the verified ones.