ZBS Index What actually exists in applied AI, with the source next to it

Evaluation

Everything here was classified as evaluation by keyword match against the maintainer's own description, so treat the grouping as a starting point rather than a verdict.

The list is ordered by the most recent commit, not by stars. A popular project that stopped in 2024 is not a better answer than a smaller one shipped last week.

Worked on in the last 90 days (25)

Someone pushed a commit recently. This is the shortlist worth trying first.

  • io.github.Daichi-Kudo/llm-advisor — last commit 2026-06-18, v0.4.5
    Real-time LLM/VLM benchmarks, pricing, and recommendations. 300+ models, 5 sources.
  • io.github.Bichev/agentradar — last commit 2026-06-17, v0.1.3
    Trust scoring, scam detection, and EAS attestations for ERC-8004 + x402 agents on Base.
  • Andru — Operational Empathy for B2B — last commit 2026-06-15, v1.5.0
    28 signed tools for founders + PE/VC operators: ICP scoring, deal classification, market signals.
  • Amber — last commit 2026-06-15, v1.1.1
    Long-term memory for AI assistants. Hybrid retrieval, query expansion, auto-topics.
  • Cert Atlas — last commit 2026-06-13, v1.2.3
    Search 1,580 certification exam blueprints: domains, scoring, prerequisites, compare. No API key.
  • io.github.erorund/hvilkenai-mcp — last commit 2026-06-11, v1.0.0
    Daily Scandinavian AI benchmark — Norwegian, Swedish, Danish. 12+ models tested.
  • io.github.apasztetnik/el-buen-agente-mcp — last commit 2026-06-11, v2.8.1
    La guía 'El Buen Agente' como 18 tools para evaluar, mejorar y construir agentes LLM. En español.
  • OntoRamp Decision Intelligence — last commit 2026-06-10, v1.0.1
    Log, evaluate, and ground AI decisions against authority context. Returns PASS, WARN, or BLOCK.
  • io.github.AutomateLab-tech/ai-seo — last commit 2026-06-08, v0.1.1
    MCP server for AI-SEO auditing: schema, robots.txt, llms.txt, citation scoring, and AEO/GEO rewrites
  • io.github.Autostackup/sales — last commit 2026-06-08, v0.1.6
    BANT + MEDDIC qualification, ICP scoring, pipeline forecasting & outreach for Claude.
  • io.github.harshitAgr/codabench-mcp — last commit 2026-06-01, v0.1.1
    MCP server for the Codabench REST API — drives a full ML-benchmark participant workflow.
  • io.github.DingDawg-dev/dingdawg-compliance — last commit 2026-05-26, v2.0.4
    EU AI Act + Colorado AI Act compliance scoring. 87/100 in 60 seconds. Free local scan.
  • io.github.dingdawg/dingdawg-compliance — last commit 2026-05-26, v2.0.11
    EU AI Act + Colorado AI Act compliance scoring. 87/100 in 60 seconds. Free local scan.
  • com.clauxel.evalscopebench/evalscopebench-mcp — last commit 2026-05-25, v1.0.0
    A paid remote MCP for AI SDK benchmark dashboard, built to return verdicts, receipts, usage logs, an
  • com.clauxel.futureagievals/futureagievals-mcp — last commit 2026-05-25, v1.0.0
    A paid remote MCP for AI SDK eval dashboard, built to return verdicts, receipts, usage logs, and aud
  • XFMS — Model Source — last commit 2026-05-23, v0.4.0
    Pick the right LLM for any task. Ranked shortlist with rationale across 8 evaluators.
  • Bawbel Scanner — last commit 2026-05-23, v1.1.0
    Scan MCP servers and skill files for AVE vulnerabilities. Conformance scoring and threat intel.
  • com.clauxel.geminiupgradeqa/geminiupgradeqa-mcp — last commit 2026-05-20, v1.0.0
    Remote MCP for Gemini upgrade evals, prompt regressions, output diffs, and eval receipts.
  • com.clauxel.safetyreplay/safetyreplay-mcp — last commit 2026-05-19, v1.0.0
    Paid remote MCP for safety replays, policy gates, eval receipts, and release evidence.
  • Agentic Platform — last commit 2026-05-19, v1.1.0
    Free MCP tools: the only MCP linter, health checks, cost estimation, and trust evaluation.
  • Zipp — last commit 2026-05-17, v1.0.2
    Multi-language crypto news with editorial sentiment + importance scoring; cites original publisher.
  • BIGHUB — last commit 2026-05-16, v0.2.1
    Decision learning for AI agent actions. Evaluate, score, decide, and learn from outcomes.
  • io.github.AIAppsAPI/adaptive-recall — last commit 2026-05-14, v1.0.1
    Adaptive memory system with cognitive scoring, knowledge graph, and self-improving ML.
  • io.github.bch1212/bizintel — last commit 2026-05-13, v0.1.0
    Local business intel for AI agents: audits, lead scoring, tech stack, prospecting.
  • com.mcparmory/posthog — last commit 2026-05-12, v1.0.5
    Capture analytics events, evaluate feature flags, and manage projects

Exists, but quiet (14)

No commit in the last 90 days. Often fine for something small and finished, risky for anything you need fixed.

  • io.github.dadang11/cryptoiz — last commit 2026-04-29, v4.16.17
    Solana DEX whale intelligence: alpha, divergence, phase scoring, BTC regime. USDC payments via x402.
  • AgentTrust — last commit 2026-04-27, v0.1.2
    Quality verification for AI agents and MCP servers. 6-axis scoring, adversarial probes.
  • io.github.akkylab/litra-paper-search — last commit 2026-04-19, v1.0.1
    MCP server for Litra.ai – AI-powered academic paper search with relevance scoring and summarization
  • Open Brain — last commit 2026-04-17, v1.1.0
    Graph-structured MCP memory server. 37.2% LongMemEval. Auto dedup, themes, decay, synthesis.
  • UK Legislation MCP Server from MCPBundles — last commit 2026-04-11, v1.0.0
    Search UK Acts, Statutory Instruments, and legislation with full text retrieval
  • io.github.ebenezer-isaac/llmconveyors — last commit 2026-04-08, v0.2.1
    53 tools for LLM Conveyors: job hunting, B2B sales, ATS scoring, resume tools.
  • io.github.AlexanderLawson17/revettr-mcp — last commit 2026-04-02, v0.2.3
    Counterparty risk scoring for agentic commerce via x402 micropayments.
  • StressZero MCP - Burnout Risk Scoring — last commit 2026-03-24, v1.0.1
    Burnout risk scoring across 3 dimensions (physical, emotional, effectiveness)
  • FeedOracle Stablecoin Risk — last commit 2026-03-19, v1.0.0
    7-signal stablecoin risk scoring. 13 tools, 28+ tokens, SAFE/CAUTION/AVOID verdicts.
  • Maximum Sats — last commit 2026-03-10, v2.0.0
    Bitcoin AI + Nostr WoT scoring (12 tools). L402 pay-per-call. 50 free WoT/day.
  • io.github.BigJai/tokennuke — last commit 2026-03-02, v1.3.0
    Code indexing MCP: 15 tools, 10 languages, hybrid search, call graphs, O(1) retrieval.
  • io.github.BigJai/codemunch-pro — last commit 2026-03-02, v1.2.0
    Code indexing MCP: 15 tools, 10 languages, hybrid search, call graphs, O(1) retrieval.
  • Bulwark — last commit 2026-02-25, v0.2.0
    AI agent governance: content scanning, audit logs, policy evaluation, session management.
  • ai.smithery/magenie33-quality-dimension-generator — last commit 2025-09-22, v1.0.0
    Generate tailored quality criteria and scoring guides from your task descriptions. Refine objectiv…

Listed, not yet verified by us (17)

Published to the registry, but we have not yet checked its repository. Treat the entry as the maintainer’s claim only.

  • ai.agentrapay/agentra — v1.0.0
    Identity oracle and trust layer for autonomous AI agents. Bidirectional KYA and trust scoring.
  • APIThreshold — v0.1.0
    Quality scoring and progressive gates for AI-generated API tests. Stripe/Twilio profiles. Free tier.
  • Axon NeuroAutomata — v0.29.0
    Protein analysis: ESM-2/ESMC embeddings, mutation scoring, landscape scans, ESMFold structure.
  • ai.childadhd/library — v1.0.0
    Clinician-reviewed library on ADHD in children — evaluation, treatment, school, parenting.
  • ai.childpsychiatry/library — v1.0.0
    Clinician-reviewed library on child psychiatric evaluation and medication decision-making.
  • Decision Log — v0.1.0
    Append-only decisions with provenance, supersession, retrieval, and audited MCP actions.
  • ai.echoloc/company-technographics — v1.0.0
    Search 760K+ companies by technographics with direction of change: adopting, replacing, evaluating
  • ai.factori/mcp — v1.0.1
    Real-world location intelligence: foot traffic, trade areas, demographics, site scoring, and more.
  • ai.multi-turn/enacta — v1.0.0
    Long-term memory for AI agents: durable records, observable retrieval, governed context assembly.
  • Scite — v1.0.0
    Ground answers in scientific literature. Search full text, evaluate trust, access full-text articles
  • ai.teenadhd/library — v1.0.0
    Clinician-reviewed library on ADHD in teens — accommodations, executive function, evaluations.
  • The Aggregate — LLM benchmark aggregate — v1.0.1
    Fused LLM rankings: one IRT/Elo scale across ~5,000 public benchmark leaderboards, updated daily.
  • ai.velarion/company-intelligence — v0.1.0
    Exec comp benchmarking, say-on-pay risk, and governance cards for US public companies.
  • Recordo: Task Planner & Notes — v1.0.0
    Brain dump, routines, task planning, focus, and instant thought retrieval
  • STEADYWRK Field Service Dispatch — v0.1.1
    Field service dispatch: instant quotes, tracked work orders, public evals. steadywrk.app
  • ZEN SecDB — v1.0.0
    ZEN SecDB MCP server for CVE intelligence, CVSS/EPSS scoring, advisories, SSVC, and package audits.
  • com.argosvix/server — v1.1.1
    Observability MCP server: query LLM cost/errors/latency & operate alerts/evals from Claude/Cursor

Repository gone or archived (4)

The registry still lists these, but the linked repository returns 404 or the owner archived it. Shown so you do not spend time discovering that yourself.

  • io.github.base76-research-lab/cognos-session-memory — archived by owner, last commit 2026-03-03, v0.1.0
    CognOS trust scoring (C=p·(1-Ue-Ua)) and session trace storage as MCP tools.
  • LimitGuard Trust Intelligence — repository gone, v1.0.1
    Entity verification, sanctions screening, and trust scoring for AI agents.
  • RadiusOS CRM — repository gone, v1.0.0
    34-tool CRM server — contacts, pipeline, quotes, invoices, scheduling, email, and AI scoring.
  • Papertext — repository gone, v0.1.0
    Academic literature search, retrieval, and private library management on top of OpenAlex.

Page 2 of 7

How this page is ordered

Entries are grouped by whether anyone is still working on them, using the date of the most recent push to the repository. They are not ordered by stars, because a star is a bookmark somebody left once and never took back.

Where we have not checked an entry yet, it says so rather than being mixed in with the verified ones.