ZBS Index What actually exists in applied AI, with the source next to it

Evaluation

Everything here was classified as evaluation by keyword match against the maintainer's own description, so treat the grouping as a starting point rather than a verdict.

The list is ordered by the most recent commit, not by stars. A popular project that stopped in 2024 is not a better answer than a smaller one shipped last week.

Worked on in the last 90 days (58)

Someone pushed a commit recently. This is the shortlist worth trying first.

  • OpenChainBench — last commit 2026-08-08, v1.0.0
    Live, neutral benchmarks for public RPC latency, oracles, bridges, perp DEX, and prediction markets.
  • io.github.getmcpm/cli — last commit 2026-08-08, v0.1.3
    MCP package manager with trust scoring. Search, install, and audit MCP servers.
  • Obsidian Hybrid Search — last commit 2026-08-08, v0.13.25
    Search Obsidian vaults with hybrid full-text, fuzzy, semantic, and graph retrieval.
  • io.github.cammac-creator/ibanforge — last commit 2026-08-06, v1.4.3
    Pre-payout IBAN screening for AI agents: validation, sanctions, Swiss clearing, risk scoring
  • Trust Score API — last commit 2026-08-06, v1.1.0
    Trust scoring for domains, wallets, APIs. SSL+DNS+WHOIS+headers. Score 0-100.
  • io.github.cdeust/hypermnesia-mcp — last commit 2026-08-06, v4.17.1
    Persistent memory for Claude — 36 cited neuroscience mechanisms, local-first, hybrid retrieval.
  • Palimpsest — censorship and model-eval observatory — last commit 2026-08-06, v1.3.1
    Live internet-censorship signals and tamper-evident, pre-registered, hash-chained AI model evals.
  • io.github.gvasile29/qai-consultant-mcp — last commit 2026-08-06, v3.3.1
    Keyless local MCP server for QA: standards retrieval, effort estimation, doc review, test analysis.
  • CPersona — last commit 2026-08-06, v2.5.3
    Persistent AI memory in one SQLite file: 3-layer hybrid search, confidence scoring, 29 tools.
  • io.github.Atharva-Jayappa/blast-scope — last commit 2026-08-06, v0.6.0
    Contextual blast-radius scoring for shell commands an AI agent is about to run
  • Speko AI — last commit 2026-08-05, v1.0.12
    Manage Speko voice-AI agents, sessions, calls, phone numbers, knowledge bases, evals, and docs.
  • io.github.dingdawg/dingdawg-sales-agent — last commit 2026-08-05, v2.0.6
    AI prospect research, outreach drafting, lead scoring & pipeline tracking.
  • io.github.dingdawg/dingdawg-shield — last commit 2026-08-05, v1.0.4
    AI security scanning and trust scoring. Stack-specific threat models. Free local scan.
  • Arsenal Decision Engine — last commit 2026-08-05, v2.0.1
    R_net LP risk evaluator for DeFAI agents. IL + Breakeven Corridor O(1). L402 Lightning paywall.
  • CompletionKit — last commit 2026-08-05, v1.0.0
    Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
  • Nautilus Compass — last commit 2026-08-05, v1.6.2
    Drift-aware cross-agent memory · MCP/A2A · LongMemEval-S 56.6% · 1/15 Zep cost
  • io.github.aa0101181514/tw-legal-rag — last commit 2026-08-05, v1.2.2
    Citation-guarded retrieval over 22M Taiwan court judgments and administrative interpretations
  • io.github.enzoemir1/leadpipe-mcp — last commit 2026-08-05, v1.4.0
    AI lead qualification: ICP filter, 0-100 scoring, Hunter.io enrichment, HubSpot/Pipedrive export.
  • Entity Enricher — last commit 2026-08-05, v2.0.0
    Multi-LLM entity enrichment: schemas, single/batch enrichment, fusion, model benchmarks.
  • hubmesh — last commit 2026-08-05, v0.4.1
    Deterministic multi-hop graph retrieval for RAG. Zero LLM calls in the query path.
  • Sema — last commit 2026-08-04, v1.34.2
    MCP tools for Sema — eval, compile, build, format, and docs for a Lisp with LLM primitives.
  • ERP Report Engine — last commit 2026-08-04, v0.7.0
    Read-only MCP for the SQL database behind an ERP: provably read-only guard, public benchmark.
  • Gateco — last commit 2026-08-04, v1.8.1
    Permission-aware retrieval for AI systems: policy-enforced access to organizational knowledge.
  • io.engineeringleaders/elc-toolkit — last commit 2026-08-04, v1.0.0
    Leadership-ratio benchmark, partnership ROI builder, community-launch readiness test. ELC data.
  • Agent^Rider — last commit 2026-08-04, v1.0.0
    Signed agent identity, trust scoring, credit economy, and social layer for AI agents.
  • Evidrift — last commit 2026-08-04, v0.4.0
    Catch TypeScript API and OpenAPI contract drift in AI-generated code, then revalidate it in CI.
  • io.github.AInoAKARI/keymaster-mcp — last commit 2026-08-03, v1.0.2
    Read-only runtime secret retrieval from HashiCorp Vault via Keymaster for autonomous AI agents.
  • io.github.geeks-accelerator/inbed — last commit 2026-08-03, v1.0.0
    AI agent dating — personality matching, compatibility scoring, and real conversations on inbed.ai
  • io.github.Garl-Protocol/agent-trust — last commit 2026-08-03, v1.4.3
    Tamper-evident action receipts, trust scoring & capability tokens for AI agents. 29 MCP tools.
  • Speech AI - Pronunciation, STT & TTS — last commit 2026-08-03, v2.3.0
    Pronunciation scoring, speech-to-text, and text-to-speech for language learning
  • io.github.CSOAI-ORG/lead-scoring-ai-mcp — last commit 2026-08-01, v1.0.4
    lead-scoring-ai-mcp MCP server by MEOK AI Labs
  • io.github.hbhqq9/bde-score — last commit 2026-08-01, v1.27.1
    7-factor stock scoring MCP server. US/HK/CN, 74 stocks. Free + Premium (USDC/Base). x402 ready.
  • io.github.Apex-Foundation/copilot-mcp — last commit 2026-07-30, v0.11.5
    Web3 founder diligence: code audit, jurisdiction, fund matching, portfolio, scoring.
  • io.github.cyanheads/protein-mcp-server — last commit 2026-07-30, v1.0.3
    MCP Server for 3D protein structural data retrieval & analysis from RCSB PDB, PDBe, and UniProt.
  • io.github.cyanheads/evals-mcp-server — last commit 2026-07-30, v0.1.2
    Author verifiable eval records through a draft→review→revise→submit loop with enforced graders.
  • coach.marian/eng-leadership-toolkit — last commit 2026-07-26, v1.2.1
    Engineering leadership benchmarks, 1:1 playbooks, developer value calculator. 3,400+ sessions.
  • Okareo — last commit 2026-07-25, vpublic-0.0.43
    Simulation, evaluation and monitoring for voice agents.
  • io.github.ahmedEid1/forgejudge — last commit 2026-07-25, v0.1.1
    Open eval leaderboard + CI gate for autonomous coding agents (solve, score, trace).
  • Keyword Research API — last commit 2026-07-19, v1.1.0
    SEO keyword research with Google Suggest, intent scoring, long-tail discovery. x402.
  • io.github.gabrielmahia/mkopo-mcp — last commit 2026-07-18, v0.1.3
    💳 mkopo-mcp — Alternative Credit Scoring MCP Server
  • Password Strength Analyzer API — last commit 2026-07-18, v1.1.0
    Password strength scoring with entropy, crack time estimation, common password detection. x402.
  • PuzzleTide — last commit 2026-07-17, v0.2.1
    Word search, crossword, and sudoku generators, printable PDF worksheets, and verifiable LLM evals.
  • io.github.codespar/mcp-clearsale — last commit 2026-07-16, v0.2.0-alpha.3
    MCP server for ClearSale — Brazilian fraud prevention, order risk scoring, device fingerprinting
  • io.github.codespar/mcp-konduto — last commit 2026-07-16, v0.2.0-alpha.3
    MCP server for Konduto — Brazilian fraud prevention: order risk scoring, device intel, lists
  • io.github.codespar/mcp-legiti — last commit 2026-07-16, v0.2.0-alpha.3
    MCP server for Legiti — Brazilian fraud prevention: real-time order evaluation, chargeback feedback
  • io.github.codespar/mcp-rd-station — last commit 2026-07-16, v0.2.2
    MCP server for RD Station — contacts, events, funnels, deals, segmentations, lead scoring, webhooks
  • io.github.codespar/mcp-sift — last commit 2026-07-16, v0.2.0-alpha.3
    MCP server for Sift — global ML fraud detection: real-time risk scoring, decisions, event ingestion
  • io.github.AIDataNordic/food-recipe-mcp — last commit 2026-07-14, v1.0.1
    Semantic search across 50,000+ food recipes with hybrid retrieval and reranking.
  • Agent Readiness Auditor — last commit 2026-07-12, v0.4.7
    Audits any website for AI agent readiness and safety, scoring prompt injection risk and llms.txt.
  • io.github.cyanheads/gutenberg-mcp-server — last commit 2026-07-11, v0.1.6
    MCP server for Project Gutenberg — 75,000+ public-domain ebooks with full plain-text retrieval.
  • NDI-MCP-Server — last commit 2026-07-02, v1.0.2
    Commercial real estate deal search, comp lookup, and scoring for the Northeast US
  • io.github.cyanheads/calculator-mcp-server — last commit 2026-07-02, v0.4.0
    Evaluate, simplify, and differentiate mathematical expressions.
  • io.github.grossiweb/toolroute — last commit 2026-07-01, v0.2.2
    Route AI agent tasks to the best MCP server and LLM, scored on 132+ real benchmark executions.
  • io.github.comil27/solrisk-mcp — last commit 2026-06-28, v1.0.0
    Solana token risk (rug / honeypot / dump) scoring for AI agents — remote MCP server.
  • ATOM Pricing Intelligence — last commit 2026-06-23, v1.1.2
    The Global Price Benchmark for AI Inference. 1,600+ SKUs, 40+ vendors, 14 price indexes.
  • ai.responsibleailabs/rail-score — last commit 2026-06-22, v1.1.1
    Responsible-AI guardrails for agents: scoring with policy, injection & PII detection, DPDP.
  • io.github.Gareth1953/agent-services-mcp — last commit 2026-06-22, v0.2.1
    MCP server: x402-paid & free tools for AI agents — provenance, quality scoring, action audit.
  • ai.byteask/embedded-docs — last commit 2026-06-20, v1.0.1
    Page-cited retrieval for embedded docs, datasheets, MISRA, CMSIS, and RTOS references.

Repository gone or archived (2)

The registry still lists these, but the linked repository returns 404 or the owner archived it. Shown so you do not spend time discovering that yourself.

Page 1 of 7

How this page is ordered

Entries are grouped by whether anyone is still working on them, using the date of the most recent push to the repository. They are not ordered by stars, because a star is a bookmark somebody left once and never took back.

Where we have not checked an entry yet, it says so rather than being mixed in with the verified ones.