ZBS Index What actually exists in applied AI, with the source next to it

mcp server

Hlido Agent Reviews

Independent AI-agent reviews: trust checks, evidence scorecards, incident registry, recommendations.

Description as published by the maintainer. Source

  • version 1.0.0
  • active

active — Most recent push to the repository was 2026-07-15.

What this server can do

17 functions, named and described by the server itself. Parameter names are shown because they say more about what a function does than its name usually does.

commerce_check(agent_or_url)
Check whether a Hlido-reviewed agent is ready to be delegated to / transacted with in the agentic-commerce world (MCP/ACP/AP2). Returns its independent Agentic-Commerce Readiness score (0-100), band (COMMERCE-READY/INTEGRABLE/SURFACE-ONLY/CLOSED), the programmatic surfaces it exposes, and the evidence basis. Call this before an orchestrator delegates a paid/identity-bearing task to another agent. Required: agent_or_url.
compare_agents(slugs)
Head-to-head trust comparison of 2-5 Hlido-reviewed agents. Returns each agent's Laddoo score, tier, dimension scores, and key claim verdicts side by side so you can pick the most trustworthy option for a task. Use this once you've shortlisted candidates (via find_trusted, find_similar_agents, or recommend) and need a direct comparison. Required: slugs.
explain(slug, dimension)
Structured natural-language explanation of why a Hlido-reviewed agent has its current score. Pulls claim-by-claim evidence from the published scorecard. Pass an optional dimension (one of: reliability, transparency, integration, security, evidence) to filter; omit for the full picture. Returns each claim with verdict (PASS|FAIL|PARTIAL|UNKNOWN), a quoted evidence snippet, plus a top-line synthesis. Required: slug.
find_similar_agents(top_k, min_score, description)
Semantic search over Hlido's review corpus. Given a task description (e.g. 'I need an agent that can refactor TypeScript and edit multiple files at once'), returns the top-N reviewed agents ranked by embedding similarity, each with their Laddoo score, evidence_tier, and review URL. Use this when you have a task in mind and want Hlido's recommendation — much better than substring matching via find_trusted. Required: description.
find_trusted(need, limit, min_tier)
Discover Hlido-reviewed agents that match a free-text need, ranked by trust. Returns reviewed agents at or above a minimum tier, each with its Laddoo score, tier, and review URL. Use this for keyword/need-based discovery; for semantic task-matching prefer find_similar_agents, and for structured constraint filters (category/score/tier) prefer recommend. Required: need.
get_behavioral_trace(slug, spec_version)
Fetch the behavioral evaluation trace for a Hlido-reviewed agent — per-task pass/fail, adapter used, behavioral tier, and signed trace link. Returns status 'not_yet_bench_tested' if the slug hasn't been evaluated yet, or 'not_testable' if the agent's interface doesn't support automated bench runs. Use this when you need evidence that an agent's coding/task behaviour has been independently verified beyond marketing claims. Required: slug.
get_incidents(slug, limit, category, severity)
Fetch published incidents from Hlido's NTSB-style failure registry — real observed agent failures (availability outages, regressions, hallucinations, safety issues) plus Hlido self-reported process incidents, each with severity, evidence, and vendor-response status. Filter by agent slug, severity, or category. Use this before delegating to an agent to check for known recent failures; an empty list means no published incidents, not a guarantee of reliability.
get_scorecard(slug)
Fetch the full sanitized claim-vs-evidence scorecard for one Hlido-reviewed agent. Returns every claim, verdict, evidence quote, source surface, and (for CLI/API tests) the captured command + exit_code + duration. Schema v1.0. Use this for agent-to-agent pre-flight evaluation. Required: slug.
recommend(constraints)
Constraint-driven recommendation across Hlido's reviewed agents. Pass any combination of: category, min_score, tier, use_case, max_results. Returns ranked candidates each with a why_match line. Use this when you have buyer constraints (budget, category, capability) and want Hlido's filtered shortlist instead of one-by-one trust_check calls. Required: constraints.
report_review_issue(slug, detail, reporter, issue_type)
Report an issue with a Hlido review (stale info, wrong verdict, missing claim, broken link). Use when calling get_scorecard or trust_check returns data you can prove is incorrect. Hlido's R1 maintenance routine processes reports daily and fires re-tests via dispute-retest sub-agent. Required: slug, issue_type, detail.
request_quick_audit(url, why, name, requester)
Request that Hlido audit a NEW AI agent that has no review yet. Use this when trust_check or get_scorecard returns no_review_found and you need a verdict before delegating to the unknown agent. Returns a future scorecard URL + ETA. Free-tier rate-limited (5/day per anonymous, 50/day per identified). The audit produces signed evidence + claim verification within ~24h (sooner if founder triggers manually). Required: url.
scan_mcp(server, requester)
On-demand independent SAFETY scan of an MCP server — call this BEFORE installing or connecting to one. Give it an HTTP(S) MCP endpoint URL (scanned live in seconds), or an npm/PyPI package name or GitHub repo (queued for an isolated sandbox scan — local stdio servers execute code, so Hlido never runs them inline). Returns the safety tier (SAFE/CAUTION/RISKY/DANGEROUS), tool-poisoning detection (the malice signal), dangerous-capability red-flags (shell/code-eval/fs-write/egress/secrets) with per-tool evidence, and auth posture. Tier = blast radius if hijacked, not maintainer trustworthiness. A server Hlido hasn't scanned returns not_scanned — never assumed safe. Register of already-scanned servers: https://hlido.eu/mcp/ Required: server.
submit_agent(url, name, note, email)
Nominate a new AI agent for Hlido to review. Use this when an agent isn't in Hlido's corpus yet (trust_check returned no_review_found) and you want it added. Returns a confirmation with a tracking reference; the review is queued and produces a public scorecard. If you need a verdict right now rather than a queued review, use request_quick_audit (faster, rate-limited) instead. Required: url, name.
subscribe(slug, channel)
Preview — Wave 3 will add persistent webhook + RSS subscriptions. For now this returns the agent's current state plus advisory polling instructions (RSS at /changelog/feed.xml or polling /data/attestations/{slug}.json). Use this to register interest in being notified when a slug's verdict changes. Required: slug.
trust_check(use_case, agent_or_url)
The core Hlido trust query: is a specific AI agent trustworthy? Given one agent (by Hlido slug or product/homepage URL) it returns the independent Laddoo trust score (0-100), tier (VITAL/STEADY/FADING/FLATLINE), a one-line verdict, a claim-verification summary, and any known incidents. Call this FIRST — before delegating to, installing, or relying on another agent — to get a fast trust read. Returns no_review_found if the agent isn't in Hlido's corpus (then call request_quick_audit). For the full claim-by-claim evidence, follow up with get_scorecard. Required: agent_or_url.
verify_claim(agent, claim)
Fact-check one specific marketing or capability claim about an agent against Hlido's independent testing. Returns Hlido's verdict (PASS/FAIL/PARTIAL/UNKNOWN) with a quoted evidence snippet and its source surface — or an honest null when that exact claim wasn't tested (absence of evidence, not proof). Use this to validate a vendor's specific promise before you rely on it. Required: agent, claim.
verify_transparency(url)
Check any AI agent's EU AI Act Article-50 transparency posture — including agents Hlido has NOT reviewed yet. Returns two clearly separated layers: (1) Hlido's independent register verdict when the agent is in our reviewed corpus, and (2) a live public-surface probe of the Article-50 signals (AI-interaction disclosure, machine-readable marking/provenance, deepfake/synthetic labelling, detection tool). Use before adopting or delegating to a tool ahead of the 2026-08-02 transparency obligations. The live probe is a first-pass surface read, NOT a compliance determination and NOT legal advice; an undetected signal means 'not discoverable on the public surface', not 'non-compliant'. Unreviewed agents are automatically queued for a full independent review. Required: url.

Last successful function declaration observed on . Source: https://hlido.eu/mcp. We list what the server declared; we do not call any of these functions.

Endpoint status observed on . Source: https://hlido.eu/mcp.

Signals

These are separate measurements of different things. They are deliberately not combined into one score, because a popularity number that mixes website traffic with saves and stars cannot be checked or acted on.

Signal Value What it measures Window Observed Source
GitHub stars 0 Number of GitHub accounts that bookmarked this repository since it was created. It is a bookmark count, not installs, not active users and not quality. cumulative, all time GitHub
Last commit 2026-07-15 Date of the most recent push to any branch. This is the strongest cheap indicator of whether the project is still maintained. point in time GitHub
Open issues 0 Open issues plus open pull requests, as GitHub counts them together. A high number can mean an active project or an abandoned one. as of fetch GitHub
Latest published version 1.0.0 Latest version string the maintainer published to the registry. as of fetch Model Context Protocol
Registry record last updated 2026-06-12 When the registry record was last updated by its maintainer. point in time Model Context Protocol
License MIT Licence GitHub detected in the repository. Detection can be wrong; the LICENSE file is authoritative. as of fetch GitHub
First listed in the MCP Registry 2026-06-12 Date this server was first published to the official MCP Registry. Not a usage or quality measure. point in time Model Context Protocol
repository status active The repository exists on GitHub and is not archived. This says nothing about how recently it was worked on. as of fetch GitHub
mcp tools declared 17 tools Number of functions the server itself declared when asked to list them. This is what the server offers an agent, not a measure of how well any of them work. as of probe hlido.eu
mcp endpoint status ok The server listed 17 functions when asked. as of probe hlido.eu

Where to get it

Related, by what their authors tagged them

  • io.github.hidai25/evalview-mcp — shares agent-evaluation
    Regression testing for AI agents. Golden baselines, CI/CD, LangGraph, CrewAI, OpenAI, Claude.
  • AgentTrust — Identity & Trust for A2A Agents — last commit 2026-04-09, shares trust
    Identity, trust, and A2A orchestration for autonomous AI agents. Official A2A partner.
  • Fidensa — last commit 2026-04-01, shares trust
    Independent AI certification authority. Trust scores, search, comparison, and verification.
  • com.fronesislabs/dcl-trust-oracle — last commit 2026-08-06, shares trust
    Deterministic AI audit layer for LLM/agent outputs: policy checks, tamper-evident log, x402.
  • dev.forgesworn/bray — last commit 2026-08-06, shares trust
    Trust-aware Nostr for AI agents -- 235 tools covering social, DMs, zaps, trust, and identity
  • info.agentrank/agentrank — last commit 2026-06-30, shares trust
    Is this AI agent or x402 service real and settlement-backed before you pay it? Trust check.
  • io.github.AgentTanuki/agent-guild — last commit 2026-08-05, shares trust
    Free self-serve Agent Passports for AI agents: signed, portable, offline-verifiable credentials.
  • io.github.Aidress-ai/aidress — last commit 2026-08-06, shares trust
    Verify, discover, and rate AI agents before transacting.
  • DOS — the trust substrate for agent fleets — last commit 2026-07-17, shares trust
    Verify what agents actually shipped, arbitrate file collisions, refuse with structured reasons.
  • AgentTrust — last commit 2026-04-27, shares trust
    Quality verification for AI agents and MCP servers. 6-axis scoring, adversarial probes.

These share tags the maintainers applied themselves, such as agent-evaluation, trust. Common tags like "mcp" or "ai" are ignored for this: agreeing with six hundred other projects is not a similarity.

This is not a recommendation and not a test result. It is a map of what the authors said their work is about.

How the author describes it

Topics the maintainer set on GitHub: agent-evaluation, ai-agents, cloudflare-workers, mcp, mcp-server, model-context-protocol, reviews, trust.

This record as data

Every field on this page, with its source and observation date, is in the catalog JSON. Fetch the whole kind at once instead of parsing this HTML.

GET /api/v1/entries/mcp_server.json

Sources

  1. ankitkapur1992-hlido/hlido-mcp on GitHub — GitHub, observed , trust tier 3.
  2. Tools declared by the MCP server at https://hlido.eu/mcp — hlido.eu, observed , trust tier 1.
  3. Official MCP Registry — Model Context Protocol, observed , trust tier 1.