ZBS Index What actually exists in applied AI, with the source next to it

mcp server

io.github.ahmedEid1/forgejudge

Open eval leaderboard + CI gate for autonomous coding agents (solve, score, trace).

Description as published by the maintainer. Source

  • version 0.1.1
  • active
  • evaluation

active — Most recent push to the repository was 2026-07-25. Dashed tags are derived by ZBS Index from the published description, not stated by the maintainer.

Signals

These are separate measurements of different things. They are deliberately not combined into one score, because a popularity number that mixes website traffic with saves and stars cannot be checked or acted on.

Signal Value What it measures Window Observed Source
GitHub stars 0 Number of GitHub accounts that bookmarked this repository since it was created. It is a bookmark count, not installs, not active users and not quality. cumulative, all time GitHub
Last commit 2026-07-25 Date of the most recent push to any branch. This is the strongest cheap indicator of whether the project is still maintained. point in time GitHub
Open issues 2 Open issues plus open pull requests, as GitHub counts them together. A high number can mean an active project or an abandoned one. as of fetch GitHub
Latest published version 0.1.1 Latest version string the maintainer published to the registry. as of fetch Model Context Protocol
Registry record last updated 2026-05-30 When the registry record was last updated by its maintainer. point in time Model Context Protocol
License MIT Licence GitHub detected in the repository. Detection can be wrong; the LICENSE file is authoritative. as of fetch GitHub
First listed in the MCP Registry 2026-05-30 Date this server was first published to the official MCP Registry. Not a usage or quality measure. point in time Model Context Protocol
repository status active The repository exists on GitHub and is not archived. This says nothing about how recently it was worked on. as of fetch GitHub

Where to get it

Related, by what their authors tagged them

  • CoinRithm Agent Trading — last commit 2026-08-06, shares leaderboard
    Keyless prediction-market data across 12 venues plus paper-trading of crypto spot, futures, and PM.
  • CompletionKit — last commit 2026-08-05, shares llm-evaluation
    Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
  • Citadel — last commit 2026-08-06, shares llm-evaluation
    Encrypted-first embedded database with vector search and agent memory, exposed as MCP tools
  • io.github.Daichi-Kudo/llm-advisor — last commit 2026-06-18, shares swe-bench
    Real-time LLM/VLM benchmarks, pricing, and recommendations. 300+ models, 5 sources.
  • io.github.dcondrey/misterdev — last commit 2026-07-27, shares swe-bench
    Autonomous LLM build orchestrator that plans, edits, and verifies code across languages.
  • io.github.archonics/mcp-audit — last commit 2026-05-14, shares evaluation
    Free context-engineering audits for AI agents. BYOK Anthropic key. Top-3 findings per scan.
  • AgentTrust — last commit 2026-04-27, shares evaluation
    Quality verification for AI agents and MCP servers. 6-axis scoring, adversarial probes.
  • io.github.botzrDev/dreamd — last commit 2026-08-05, shares evaluation
    Local-first, cross-harness memory for AI coding agents.
  • io.github.cyanheads/evals-mcp-server — last commit 2026-07-30, shares evaluation
    Author verifiable eval records through a draft→review→revise→submit loop with enforced graders.
  • ai.testiv/mcp — last commit 2026-08-04, shares ci
    Local-first visual regression for AI agents: verdicts, diff images, explain_snapshot. No API key.

These share tags the maintainers applied themselves, such as leaderboard, llm-evaluation, swe-bench, evaluation. Common tags like "mcp" or "ai" are ignored for this: agreeing with six hundred other projects is not a similarity.

This is not a recommendation and not a test result. It is a map of what the authors said their work is about.

Also from ahmedeid1

  • Atlas Research — last commit 2026-06-01
    Read-only MCP over an agentic SLR workspace with per-claim citation verification
  • Lumen — last commit 2026-06-07
    Self-hostable agentic-AI LMS: catalog, RAG tutor, FSRS reviews, AI authoring, ingest.
  • Thoth — last commit 2026-06-01
    Read-only MCP over an agentic SLR workspace with per-claim citation verification

How the author describes it

Topics the maintainer set on GitHub: ai-agents, autonomous-agents, ci, evaluation, langfuse, leaderboard, llm, llm-evaluation, observability, python, swe-agent, swe-bench.

Bring your own setup

We take apart real AI setups every week and show what broke, what cost too much, and what the trace actually said. If you run agents on real work, that is where the useful conversation is.

Join ZBS AI Practice Lab

Sources

  1. ahmedEid1/forgejudge on GitHub — GitHub, observed , trust tier 3.
  2. Official MCP Registry — Model Context Protocol, observed , trust tier 1.