I connected Braintrust to two agents over a three-week test period: a customer support agent handling tier-1 tickets and a document extraction agent parsing invoices.
The support agent processes about 200 conversations per day, the extraction agent runs 50-80 documents daily.
Setup required adding the Braintrust SDK (I used the Python SDK) and wrapping existing LLM calls with trace decorators.
The initial instrumentation took about 45 minutes for the support agent and 20 minutes for the extraction agent, which already had cleaner function boundaries.
Braintrust’s SDK is lightweight compared to LangSmith’s, which wanted me to restructure chain definitions. Here, you wrap what you have and traces start flowing.
The first thing I noticed was trace search speed. Querying across three weeks of support agent logs for a specific failure pattern (“hallucinated refund policy”) returned results in under two seconds across roughly 4,200 traces.
The Brainstore database that powers this is noticeably faster than what I experienced with LangSmith’s trace explorer on similar volumes.
Braintrust was founded in 2020 by Ankur Goyal in San Francisco. The company raised an $80M Series B in February 2026 at a reported $800M valuation.
The customer list includes Notion, Stripe, Vercel, Instacart, Zapier, and Replit. Notion’s published case study claims a tenfold increase in shipping speed for AI issues, going from roughly 3 to 30 fixes per day across 70 engineers.
The platform supports native SDKs for Python, TypeScript, Go, Ruby, and C#. It is framework-agnostic, so you are not locked into LangChain or any specific orchestration layer.
There is also MCP integration for querying logs and updating prompts directly from your IDE, which I did not test but looks useful for teams living in VS Code or Cursor.
Key Features
Production Tracing and Observability
Every agent step gets captured: prompts, responses, tool calls, retrieved context, latency, and token cost.
The trace viewer breaks each step into a timeline you can click through.
I used this to catch a pattern where my support agent was making redundant tool calls on 15% of conversations, adding 2-3 seconds of latency and about $0.004 in wasted tokens per interaction.
That adds up to $24/month at 200 conversations per day. Finding that pattern in raw logs would have taken hours. The trace viewer surfaced it in minutes.
Evaluation Framework with Multiple Scorer Types
Braintrust supports LLM-as-judge scoring, custom code scorers, and human review scoring.
You define what “good” looks like for your use case, build a dataset of inputs and expected outputs, and run experiments that score your agent’s performance.
I set up a factual accuracy scorer for the support agent and a field extraction accuracy scorer for the document agent. The framework is flexible enough that a product manager can define passing criteria in plain language while an engineer wires up the scorer code.
Dataset-from-Traces (Production Failure to Test Case)
This is the feature that made me keep using Braintrust past the initial trial.
When you spot a bad trace in production, one click converts it into a test case in a versioned dataset. I built a regression suite of 47 test cases in the first week just by flagging failures as I reviewed traces.
On every other eval platform I have used, the cold-start problem of “what do I even test against?” stalls adoption for weeks. Braintrust sidesteps it by treating production as the source of truth for test data.
Loop Agent for Automated Prompt Improvement
The Loop agent takes a plain-language description of what you want to improve and proposes updated prompts, new scorers, and expanded datasets.
I asked it to “reduce hallucinated policy citations in refund conversations” and it generated a prompt variant that scored 12% higher on my factual accuracy scorer.
The suggestions are starting points, not production-ready changes, but they cut the manual prompt iteration cycle from hours to minutes.
Discover: Automated Pattern Classification
The Discover feature clusters production traces into topics and surfaces recurring patterns.
It grouped my support agent’s failures into five categories without any manual labeling: policy hallucinations, tone mismatches, incomplete tool call chains, context window overflows, and language detection failures. That classification would have taken a full afternoon of manual review.
AI Proxy with Model Credits
Braintrust includes a proxy layer that routes LLM calls and comes with model credits ($10/month free, $100/month on Pro).
You can use the proxy to compare models side-by-side in the playground without switching API keys. The credits are modest but cover enough experimentation to test prompt variants without dipping into your own API budget.
Competitors Comparison
| Feature | Braintrust | Langfuse | LangSmith | Arize Phoenix | Galileo |
| Free tier | 1 GB data, 14-day retention | Self-hosted (unlimited), cloud free tier | Developer free tier | Open source | Limited free |
| Starting paid price | $249/month (flat) | $29/month (cloud) | $39/user/month | $50/month | $100/month |
| Pricing model | Data + scores | Observations | Per-seat | Spans | Per-seat |
| Self-host option | Enterprise only | Yes (open source) | No | Yes (open source) | No |
| Runtime guardrails | No | No | No | No | Yes |
| SDK languages | Python, TS, Go, Ruby, C# | Python, TS, JS | Python, TS, JS | Python | Python |
| Human review | All plans (1/project free) | Cloud plans | Plus+ | Yes | Yes |
| Dataset from traces | Yes (one-click) | Manual export | Yes | Limited | Limited |


