Evaluate, observe, and actually ship agents that work

Trail traces every step your agents take and scores every run — so you can see what happened, prove it improved, and catch regressions before your users do.

Trace every run today. Add evals, experiments, and human review when you are ready.

Agent runs · overview
Run volume
$4,412 +23.9%
1,284 runs
MonSun
Runs by model
claude gpt-5 gemini llama
Jul 26Aug 6
Slowest runs
runmodellatencyverdict
support-assistantclaude-sonnet-4.54.21 spassed
doc-summarizergpt-53.05 sflagged
agent.plan · step 3gemini-2.5-pro2.88 spassed
search-rerankllama-3.3-70b1.44 sblocked

Trusted by engineering teams running agents in production

Northcurrent
Lumenwave
qillbase
Fieldstone
Verdantly
Harborline
sablewood
Alderpoint
Why Trail

Your agent is in production.
Nobody can prove it works.

Every team demos an agent in a week, then spends nine months afraid to touch it. Not because the model is weak — because there is no loop. The run fans out across a dozen calls and none of it is captured, so debugging starts from a screenshot and “better” means someone liked the demo.

Prompt fixes ship blind, customers find the bugs first, and every release becomes a negotiation. Teams that close the loop ship weekly.

Observability

See every step your agent takes

Trace each run from your app through every tool call and model call and back — latency, tokens, and cost in real time, down to the individual span.

Hierarchical traces, span by span
Filter by session, user, cost, latency, or metadata
Auto-instruments 30+ agent libraries
Runs · last 14 days $1,284.60 +18%
featurecallscost
support-assistant41,208$742.10
doc-summarizer9,655$318.75
search-rerank88,320$223.75
Playground "Summarize this refund policy in two sentences."
gpt-5
1,140 ms · $0.0091
claude-sonnet-4.5
942 ms · $0.0084
gemini-2.5-pro
870 ms · $0.0062
mistral-large
610 ms · $0.0021
command-r+
720 ms · $0.0028
grok-4
1,320 ms · $0.0074
llama-3.3-70b
480 ms · $0.0009
deepseek-v3
1,010 ms · $0.0007
8 providers · 1 prompt best: claude-sonnet-4.5
Evaluation

Score every run, not just the demo

LLM-as-a-judge, heuristic scorers, or human review. Run evaluators on live production traces or inside a controlled experiment, and compare versions side by side.

Judges, heuristics, or human review
Run on live traces or in an experiment
Compare prompts, models, and versions
Experiments

Prove the new version is better

Define test cases, run your agent against them, and gate the release on the result. Set the passing bar once and every experiment is judged against it.

Test cases promoted from real traces
Accuracy, safety, and tone scored per run
Tune the passing bar to fit your product
experiment · support-assistant
0.60
evaluatorscoreverdict
accuracy0.91
citation_match0.82
conciseness0.64
tone0.43
is_json0.28
scored on every run · 71 ms overhead
Also included

The rest of the loop

It reaches every framework

Python and TypeScript SDKs, a CLI for CI pipelines, and an MCP server for IDE agents. Auto-instrumentation over standard OpenTelemetry.

Eval suite Running
citation_matchjust now
tone2 min ago
is_json2 min ago

It brings humans in

Review queues for labelling runs, and golden datasets curated straight from production traces.

One prompt, versioned v7
"Answer only from retrieved context. Cite article ids."
support.reply production pulled at runtime 3 variants

It keeps prompts straight

Version prompts and deploy changes without shipping code — the SDK pulls the live template at runtime.

Secret vault · project-scoped encrypted
OPENAI••••4f9a
ANTHROPIC••••b21c
PINECONE••••7e03

It's yours to keep

Credentials never live in code. Encrypted, project-scoped, rotatable — and on Enterprise, inside your own boundary.

Try Trail
Getting started

Up and running in three steps.

Add the SDK, watch every run stream in, then score them with LLM judges or your own heuristics and gate releases on the result.

01 connect — add the SDK
02 watch — runs stream in
03 evaluate — score and gate
python typescript
$ pip install trail
# app.py
import trail
trail.observe(app)
trail.evaluate("accuracy", dataset="golden-set")
# that's it — every run traced and scored
# across 30+ agent libraries
receiving spans from 3 services
Pricing

Start free. Then pick your plan.

Every plan includes the full platform — observability, evaluation, and experiments. Scale up as your team and your agent traffic grow.

Free For individual developers and first agents.
$0 /mo
Start for free
No credit card required
2 seats
1 project
1M spans / mo
30-day data retention
Community support
Full platform access
Starter For small teams shipping their first agents.
$27 /mo
Contact sales
Talk to our team
Small-team seats
Multiple projects
15M spans / yr
Extended retention
Community support
Everything in Free
Enterprise For large-scale, regulated, or self-hosted deployments.
Custom
Book a demo
Talk to a solutions engineer
Custom seats & volume
On-premise deployment
Custom span volume
PDPL / DPA compliance
Dedicated CSM
Everything in Growth
Included in every plan

The full platform, on every tier.

Full trace waterfalls & span-level detail
OpenTelemetry-native ingestion (OTLP)
Token, cost & latency tracking
LLM analytics dashboards (cost, model, endpoint)
Agent-run analytics (by agent, by tool)
Vector-DB & GPU hardware metrics
Guardrails (prompt-injection & topic detection)
Evaluation datasets
Built-in evals (toxicity, hallucination, bias)
Prompt versioning
OpenGround multi-provider playground
Projects, orgs & role-based access
Project-scoped secret vault
Python & TypeScript SDKs, 40+ integrations
REST API access
Works with
openaianthropicgeminimistralcoheregrokllamadeepseeklangchainllamaindexcrewaidspypineconepgvectorvercel-ai
FAQ

Before you ask

Trail is one platform for running agents in production: it traces every step a run takes, scores runs with LLM judges or your own heuristics, compares versions in experiments, and turns real traces into regression datasets — so you don't stitch five tools together.

Get started

Bring your agents under observation.

Start with full trace visibility today, and add evaluation, experiments, and human review whenever you are ready.