← Leaderboard
8.3 L4

Vellum Ai

Native Assessed · Docs reviewed · Mar 30, 2026 Confidence 0.56 Last evaluated Mar 30, 2026

Verify before you commit

Trust read first, source links second, build decision third.

Use this page to sanity-check Vellum Ai quickly. We surface the evidence tier, freshness, and failure posture here, then put the official links where you can actually act on them, especially on mobile.

Evidence

Assessed

Docs reviewed · Mar 30, 2026

Freshness

Updated 2026-03-30T14:41:26.876+00:00

Mar 30, 2026

Failures

Clear

No active failures listed

Score breakdown

Dimension Score Bar
Execution Score

Measures reliability, idempotency, error ergonomics, latency distribution, and schema stability.

8.3
Access Readiness Score

Measures how easily an agent can onboard, authenticate, and start using this service autonomously.

8.2
Aggregate AN Score

Composite score: 70% execution + 30% access readiness.

8.3

Autonomy breakdown

P1 Payment Autonomy
G1 Governance Readiness
W1 Web Agent Accessibility
Overall Autonomy
Pending

Active failure modes

No active failure modes reported.

Reviews

Published review summaries with trust provenance attached to each card.

How are reviews sourced?

Docs-backed Built from public docs and product materials.

Test-backed Backed by guided testing or evaluator-run checks.

Runtime-verified Verified from authenticated runtime evidence.

Vellum: Comprehensive Agent-Usability Assessment

Docs-backed

LLM application development and operations platform for teams building production AI features. Prompt versioning with deployment environments (sandbox/staging/production). Visual Workflow builder for multi-step chains, conditional branching, loops, API calls, and inline code execution. A/B testing across providers and models. Evaluation test suites with custom metrics. Trace monitoring for production debugging. Confidence is docs-derived.

keel-expansion Mar 30, 2026

Vellum: API Design & Integration Surface

Docs-backed

REST/SDK API: execute_prompt(deployment_name, inputs) for prompt execution; execute_workflow(deployment_name, inputs) for workflow execution; streaming variants for token-by-token output; Python SDK (vellum-ai) and TypeScript SDK; Vellum API key as Authorization header; Webhook events for async workflow completion; REST API for managing deployments, documents, and evaluation runs programmatically.

keel-expansion Mar 30, 2026

Vellum: Auth & Access Control

Docs-backed

Vellum API key per workspace (scoped to organization); environments (sandbox/staging/production) for deployment isolation; role-based access control (admin, editor, viewer) within organization; prompt and workflow content stored in Vellum — no customer prompt content used for training; audit log for deployment changes; documents uploaded to Vellum vector store are organization-scoped.

keel-expansion Mar 30, 2026

Vellum: Error Handling & Operational Reliability

Docs-backed

execute_prompt returns structured output with generation, tokens used, and finish reason; workflow execution is synchronous for short chains or async with webhook for long-running workflows; evaluation runs batch-execute against test cases and return metric scores; streaming uses SSE for incremental token delivery; error responses include error_type and message for programmatic handling; rate limits per plan.

keel-expansion Mar 30, 2026

Vellum: Documentation & Developer Experience

Docs-backed

Documentation covers prompt editor and versioning guide, Workflow builder reference (node types, branching, code execution), evaluation test suite setup, Python and TypeScript SDK quickstart, deployment environment guide, and monitoring/trace search. Designed for ML engineers and product teams iterating on production LLM features. Confidence is docs-derived.

keel-expansion Mar 30, 2026

Use in your agent

mcp
get_score ("vellum-ai")
● Vellum Ai 8.3 L4 Native
exec: 8.3 · access: 8.2

Trust shortcuts

This score is documentation-derived. Treat it as a docs-based evaluation of API design, auth, error handling, and documentation quality.

Read how the score works, how disputes are handled, and how Rhumb scored itself before launch.

Overall tier

L4 Native

8.3 / 10.0

Alternatives

No alternatives captured yet.