Evaluation
Everything here was classified as evaluation by keyword match against the maintainer's own description, so treat the grouping as a starting point rather than a verdict.
The list is ordered by the most recent commit, not by stars. A popular project that stopped in 2024 is not a better answer than a smaller one shipped last week.
Listed, not yet verified by us (55)
Published to the registry, but we have not yet checked its repository. Treat the entry as the maintainer’s claim only.
-
Tokenomic-Pulse
— v1.0.0
DeFi protocol simulation model scoring tokenomic structures.
-
io.github.hifriendbot/cogmemai
— v3.17.0
95.10% LongMemEval (highest published). Encrypted persistent memory for Ai coding assistants.
-
io.github.hongnoul/hwatu
— v0.6.0
Verification browser for coding agents: headless WebKit, DOM eval, screenshots, pixel-diff.
-
RAGScore
— v0.8.6
Generate QA datasets & evaluate RAG systems with failure diagnosis. Any LLM.
-
SEAR
— v0.1.0
Benchmark-first release surface with a read-only MCP endpoint and operator CLI.
-
VulnFeed
— v0.3.3
Dependency vulnerability scanner with EPSS scoring. 9 MCP tools. Free tier + x402.
-
io.github.infino-ai/mcp-server
— v0.7.0
Keyword, vector, hybrid, and SQL retrieval over data on object storage, for AI agents.
-
Eval Engine API
— v0.1.0
Pay-per-call AI evaluation MCP server. Score LLM outputs against benchmark rubrics via Workers AI.
-
io.github.ipezygj/evalgate
— v0.4.1
Statistical checks an agent runs before trusting an AI eval number (is #1 real, judge bias, more).
-
io.github.ipezygj/numguard
— v0.1.1
Verify a number before an agent asserts it — evals, backtests (Deflated Sharpe), signed receipts.
-
io.github.iris-eval/mcp-server
— v0.4.4
The agent eval standard for MCP. Score every agent output for quality, safety, and cost.
-
Memory MCP Server
— v1.7.0
SQLite-backed MCP server for persistent memory, full-text retrieval, and graph traversal.
-
io.github.jackmmaher/ukdatapi
— v3.0.0
UK government data intelligence with proprietary scoring from 400+ sources
-
io.github.jacobsd32-cpu/djd-agent-score
— v1.0.1
Reputation scoring for AI agent wallets on Base. Trust scores, fraud checks, x402.
-
io.github.jaimenbell/rag-mcp
— v0.1.0
Minimal RAG-over-a-corpus MCP retrieval: search_knowledge returns cited chunks. Local embeddings.
-
Agent Pulse
— v1.4.2
Reproducible benchmarks and reliability evidence for agent tools.
-
io.github.jbrand0n/market-intelligence-mcp
— v0.1.2
Recession probability, capital rotation scoring, and economic data API
-
IA-QA — 130+ QA & Dev Tools for AI Agents
— v1.0.0
130+ QA & dev tools for AI agents: prompt injection, RAG testing, VLM eval, guardrails. Free.
-
jDataMunch MCP
— v1.31.1
Tabular data retrieval. Index CSV/Excel, query rows, aggregate. 99%+ savings vs raw file reads.
-
io.github.JKHeadley/moltbridge
— v0.1.5
Agent network intelligence for AI agents. Trust scoring, broker discovery, and Ed25519 identity.
-
signet-eval
— v3.11.0
Deterministic policy enforcement and MCP management for AI agent tool calls.
-
io.github.JubaKitiashvili/context-mem
— v3.2.1
Persistent memory for AI agents — 98%+ retrieval recall, 99% token savings, 44 tools
-
io.github.Judgment-Pack/judgment-pack
— v0.15.0
Offline JPS validator, conformance tester, and experimental evaluator over stdio MCP; keyless.
-
io.github.kakunin-ai/kakunin
— v0.2.4
X.509 identity, risk scoring, and audit logging for AI agents. MiCA + EU AI Act compliant.
-
io.github.karlmehta/trustmodel-mcp
— v0.2.0
Score any AI for trust across 10 dimensions; evaluate, monitor & govern LLMs and agents.
-
io.github.keshrath/agent-knowledge
— v1.9.7
Cross-session memory for AI agents - knowledge graph, scoring, semantic search
-
Judges Panel
— v3.129.9
45 judges that evaluate AI-generated code for security, cost, and quality with built-in AST.
-
A2ABench
— v1.0.1
Public benchmark where agents submit Q&A answers and get scored on a leaderboard.
-
Argus Retrieval
— v1.6.2
Multi-provider search broker for AI agents: 14 providers, 12-step extraction, retrieval workflows.
-
io.github.khan-ashifur/hooklayer
— v1.1.0
Viral-content intelligence for AI agents — 7 read-only MCP tools, evidence-layer scoring.
-
Enterprise Internal Knowledge Base: Production-Ready RAG + MCP
— v0.1.0
Production-ready RAG + MCP demo: eval-in-CI merge gate, Langfuse traces, structure-aware chunking.
-
io.github.Kind-ling/twig
— v0.4.2
MCP server quality scoring and optimization for the agent economy.
-
io.github.KlossKarl/loom
— v0.2.1
Local-first personal knowledge base. Cited answers from your vault via vector + graph retrieval.
-
io.github.krutftw/fetchmux
— v0.1.0
Self-hosted retrieval router for AI agents: budgets, provider routing, route receipts.
-
io.github.KryptosAI/mcp-observatory
— v1.36.1
MCP security scanner. CI-native testing, attack simulation, health scoring, and SARIF.
-
AgentLore
— v1.0.0
AI-verified knowledge base with trust scoring, temporal facts, and skill cards.
-
XRPL Wallet Risk Score
— v1.3.0
XRPL wallet risk scoring. Behavioral tags, sub-scores, graph analysis. Pay per call in XRP.
-
io.github.LamboPoewert/madeonsol
— v1.16.0
Solana memecoin intelligence: KOL trades, deployer reputation, token & wallet scoring
-
io.github.LanceRoylo/mcp
— v0.1.2
AEO scoring, llms.txt audit, and agent-readiness checks for AI agents.
-
io.github.lazymac2x/ai-eval
— v1.0.0
Cloudflare Workers MCP server: ai-eval
-
io.github.LeandroPG19/cuba-memorys
— v0.21.0
Persistent memory MCP server. 29 tools, BM25+MMR+OOD retrieval, tamper-evident audit, code graph.
-
io.github.lennney/mcp-slim-guard
— v0.1.1-alpha.1
76% fewer MCP tokens in our standard benchmark. Same upstream call. Exact recovery.
-
io.github.let-sunny/canicode
— v0.5.2
Analyze Figma designs for dev & AI readiness. 39 rules, scoring, HTML reports.
-
Lore Context
— v0.6.0-alpha.1
Governed AI-agent memory, Evidence Ledger traces, evals, and portable context tools.
-
io.github.lorenzo-cambiaghi/lynx
— v1.7.5
100% local MCP server for semantic code search: AST chunking, hybrid retrieval, code knowledge graph
-
Dense Knowledge
— v1.2.0
Local-first MCP memory server for persistent LLM knowledge with BM25 retrieval.
-
io.github.martinX308/synplex
— v1.0.0
Inventory health scoring, landed cost calculation, and quick diagnosis tools for Shopify merchants.
-
Espresso MCP
— v0.2.0
Find great espresso cafes worldwide with curated data and transparent quality scoring.
-
Yagami AI Context Firewall
— v0.7.3
Govern model, retrieval, memory, and tool access for AI applications and agents.
-
Medical RAG MCP - Semantic search and retrieval of high-quality clinical medical data
— v1.0.0
Medical RAG: semantic search for clinical guidelines, drug interactions, diagnoses & EHR data.
-
Medical RAG - Clinical Decision Support & Healthcare Knowledge Retrieval
— v1.0.1
Medical RAG: semantic search for clinical guidelines, drug interactions, diagnoses & EHR data.
-
Context Engine
— v0.1.0
Compress logs, retrieval chunks, and code context into structured LLM-ready signal.
-
io.github.meltingpixelsai/drainbrain
— v1.0.2
AI-powered Solana token rug pull detection with ML ensemble scoring and honeypot detection.
-
io.github.menantonio83-hue/tnt-house-risk-data-api
— v1.0.0
Solana token risk-scoring MCP server for AI trading agents with insider wallet cluster detection.
-
io.github.MetriLLM/metrillm
— v0.2.6
Benchmark local LLM models — speed, quality & hardware fitness verdict from any MCP client
Repository gone or archived (5)
The registry still lists these, but the linked repository returns 404 or the owner archived it. Shown so you do not spend time discovering that yourself.
-
Compliance Auditor MCP
— repository gone, v0.1.0
City hiring-compliance MCP server with regulation search and full audit risk scoring.
-
Project Brain
— repository gone, v1.4.1
Read-only source retrieval across cataloged GitHub repositories
-
AI Pricing Hub
— repository gone, v1.0.0
Source-backed AI model pricing, rankings, history, and benchmark data.
-
TrustLayer
— repository gone, v1.0.0
Agent reputation scoring: trust scores, Sybil detection, cross-chain identity for 133K+ agents
-
Omnis Venture Intelligence MCP
— repository gone, v1.0.0
Venture intelligence for autonomous agents with discovery, scoring, and workspace automation.
Page 4 of 7
How this page is ordered
Entries are grouped by whether anyone is still working on them, using the date of the most recent push to the repository. They are not ordered by stars, because a star is a bookmark somebody left once and never took back.
Where we have not checked an entry yet, it says so rather than being mixed in with the verified ones.