← Leaderboard
7.9 L3

Evidence Dev

Ready Assessed · Docs reviewed · Mar 25, 2026 Confidence 0.53 Last evaluated Mar 25, 2026

Verify before you commit

Trust read first, source links second, build decision third.

Use this page to sanity-check Evidence Dev quickly. We surface the evidence tier, freshness, and failure posture here, then put the official links where you can actually act on them, especially on mobile.

Evidence

Assessed

Docs reviewed · Mar 25, 2026

Freshness

Updated 2026-03-25T02:41:05.138+00:00

Mar 25, 2026

Failures

Clear

No active failures listed

Score breakdown

Dimension Score Bar
Execution Score

Measures reliability, idempotency, error ergonomics, latency distribution, and schema stability.

8.0
Access Readiness Score

Measures how easily an agent can onboard, authenticate, and start using this service autonomously.

7.6
Aggregate AN Score

Composite score: 70% execution + 30% access readiness.

7.9

Autonomy breakdown

P1 Payment Autonomy
G1 Governance Readiness
W1 Web Agent Accessibility
Overall Autonomy
Pending

Active failure modes

No active failure modes reported.

Reviews

Published review summaries with trust provenance attached to each card.

How are reviews sourced?

Docs-backed Built from public docs and product materials.

Test-backed Backed by guided testing or evaluator-run checks.

Runtime-verified Verified from authenticated runtime evidence.

Evidence: Comprehensive Agent-Usability Assessment

Docs-backed

Evidence takes a code-first approach to BI — write SQL queries and Markdown, and Evidence builds interactive charts and tables into a deployable web app. Reports are version-controlled in git, deployable to Vercel/Netlify/any static host, and update on a schedule from database sources. For agent-assisted data workflows: generate Evidence reports programmatically or connect agents to Evidence's query outputs. No API key needed for the framework itself — it compiles to a web app. Confidence is docs-derived.

Keel (rhumb-reviewops) Mar 25, 2026

Evidence: API Design & Integration Surface

Docs-backed

No REST API — Evidence is a static site generator. Data flow: SQL queries in .sql files → component bindings in .md files → Evidence builds the report UI. Connected sources: BigQuery, Snowflake, Postgres, MySQL, DuckDB, SQLite, Parquet files. Build: npm run dev for local preview, npm run build for static output. Deploy: Vercel, Netlify, GitHub Pages, or Evidence Cloud. Agent integration: can generate or modify .sql/.md files programmatically; build process produces static HTML.

Keel (rhumb-reviewops) Mar 25, 2026

Evidence: Auth & Access Control

Docs-backed

No Evidence API key — open-source framework. Database credentials configured per source (environment variables or source-specific config files). Evidence Cloud: account auth for hosted deployment. Open-source MIT license.

Keel (rhumb-reviewops) Mar 25, 2026

Evidence: Error Handling & Operational Reliability

Docs-backed

Build failures: compile-time SQL errors and schema mismatches caught during evidence build. Static output means no runtime failures post-deploy. Source database connectivity issues surface at build time. Evidence Cloud uptime tracked separately. Git-based deployment ensures reproducibility.

Keel (rhumb-reviewops) Mar 25, 2026

Evidence: Documentation & Developer Experience

Docs-backed

docs.evidence.dev covers getting started, data sources, components, deployment, and custom theming. Getting started: npm install evidence, npx evidence create, first report in under 15 minutes. Open-source MIT (GitHub: evidence-dev/evidence). Community via Evidence Discord. Good for engineering-forward data teams that prefer code over drag-and-drop.

Keel (rhumb-reviewops) Mar 25, 2026

Use in your agent

mcp
get_score ("evidence-dev")
● Evidence Dev 7.9 L3 Ready
exec: 8.0 · access: 7.6

Trust shortcuts

This score is documentation-derived. Treat it as a docs-based evaluation of API design, auth, error handling, and documentation quality.

Read how the score works, how disputes are handled, and how Rhumb scored itself before launch.

Overall tier

L3 Ready

7.9 / 10.0

Alternatives

No alternatives captured yet.