Braintrust

  • Braintrust — AI-Native Evals & Observability Platform for Production LLM Agents (Brainstore + Loop + MCP)

    SF, founded 2023 by Ankur Goyal (ex-Impira), Series B 800M val, $124M total). AI-native evals + observability for LLM apps/agents. No acquisition, no sunset. Category ai-productivity-data, pipeline_stage=none. File under Dev​ (LLM-eval/observability sublane). This is a re-run of your prior Braintrust turn with verified 2026-09 figures.Official site: https://www.braintrust.dev​Pricing: https://www.braintrust.dev/pricing

    Docs: https://www.braintrust.dev/docs

    Snapshot

    Field Value
    Status Active
    Origin Braintrust Data Inc., SF, 2023, Ankur Goyal
    Lane LLM/agent observability + evaluation + prompt registry + quality gates
    Core parts Tracing, Evals, Brainstore (AI-trace DB, 80× faster), Loop (eval agent), Topics (pattern discovery), Gateway, MCP server
    Customers Notion, Stripe, Vercel, Cloudflare, Replit, Ramp, Dropbox, Zapier, Instacart, Coursera
    Compliance SOC 2 Type II, HIPAA BAA (Ent), GDPR, DPA
    Category ai-productivity-data / Dev (eval/obs sublane)

    Features (table)

    Layer Capability Notes
    Tracing Per-step agent traces (prompt/tool/retrieval/latency/cost), SDK Py/TS/Go/Ruby/C#/Java, OTel-compatible Backed by Brainstore, 80× faster query vs PG
    Evals Dataset + task + scorer model, batch run vs prompts/models, CI gates (GitHub Action posts PR comment + blocks merge) LLM-as-judge / code / human scorers; one-click prod trace → eval case
    Prompt registry Env-labeled versions (prod/staging/dev), playground side-by-side, no-deploy promote Fetch by env name at runtime
    Loop AI agent: draft better prompts, build/refine scorers, synthesize eval datasets from NL Uses model-credit quota (Topics-grade tokens)
    Topics Auto pattern discovery across traces (issues/sentiment/tasks) Metered separately, 0.40 per MTok overage
    Gateway Multi-provider routing, caching, cost tracking, built-in obs Optional proxy
    MCP MCP server → pull traces/evals into Cursor/Claude Code/Cline IDE-native debugging
    Deploy SaaS, BYOC, self-hosted data plane (Ent) Control plane managed
    No No agent runtime/deploy (≠ LangSmith Deploy), no APM for non-LLM services, no OSS community edition (SDKs MIT/Apache, platform closed) Scope = AI apps only

    Pricing (2026-09, USD)

    Tier Platform fee Included usage Overage / notes
    Starter $0 1 GB processed data/mo, 10k scores/mo, $10 Topics/model credits, 14-day retention, unlimited users/projects/datasets/playgrounds Data 2.50/1k, Topics 0.40 out per MTok
    Pro **249 credit till 2026-09-01, then $100/mo credit​ from 2026-09-01) 5 GB data, 50k scores, 30-day retention ($0.50/GB/mo beyond), custom charts, env, RBAC basic, priority email Data 1.50/1k, lower token rates
    Enterprise Custom (annual invoice) Custom limits, SAML/OIDC SSO, SCIM, audit log, BAA, BYOC/self-host, S3 export, Slack channel, SLA Early-stage startups (≤Series A, ≥$100K raised) get 6–12 mo free Pro

    starting_price=$0; first paid $249/mo (no per-seat fee at any tier — bill = platform fee + data + scores + Topics tokens).

    Reviews (2026)

    Source Signal
    Vendor benchmarks / Notion case study Notion: 3→30 AI issues fixed/day after adopting Braintrust evals (10×)
    ToolRadar / Voiceflow / AITrendTool Praised: trace→eval→CI gate loop, Brainstore speed, Loop autonomy, MCP IDE pull, seat-unlimited billing. Criticized: 249 jump steep, Topics/Loop token burn, compliance (SOC2/HIPAA/SSO) Enterprise-gated, closed-source platform (vs Langfuse)
    Aggregate Best-in-class eval-first observability for eng-led LLM teams; less turnkey for non-LangChain PMs than Humanloop/PromptLayer

    Access / Integration (table)

    Method Supported
    SDK Python, TypeScript, Go, Ruby, C#, Java
    Framework OpenAI, Anthropic, LangChain, Vercel AI SDK, OpenTelemetry
    Gateway Braintrust Gateway proxy (multi-provider, cache, cost)
    MCP MCP server → Cursor / Claude Code / Cline
    CI GitHub Actions eval gates (braintrustdata/eval-action)
    Deploy SaaS / BYOC / self-hosted data plane (Ent)
    No No REST “score my text” free endpoint for non-customers, no frontend analytics

    Best For / Not For

    Best For Not For
    Eng-led LLM/agent teams running CI eval gates PM-only prompt-editing shops (→ PromptLayer/Humanloop)
    Notion/Stripe-scale trace volume (breaks PG) LangChain-locked teams wanting agent runtime (→ LangSmith Deploy)
    Framework-agnostic stacks (OpenAI+Anthropic+Vercel) Tiny side-projects unwilling to watch data/score/Topics meter
    Air-gap/BYOC regulated (Ent: HIPAA/SAML/self-host) Teams needing OSS self-host free (→ Langfuse/Phoenix)
    Cursor/Claude Code users wanting MCP trace pull Non-LLM APM needs (→ Datadog)

    Competitors (lane map)

    Lane Tool Braintrust edge / gap
    Eval-first platform LangSmith LangSmith deeper LangChain + agent runtime + seat pricing; Braintrust framework-agnostic + Brainstore + Loop + seat-unlimited
    OSS observability Langfuse, Arize Phoenix OSS/self-host free; Braintrust managed Brainstore + Loop + MCP
    Prompt mgmt PromptLayer, Humanloop Lighter prompt-registry; Braintrust broader eval/obs/trace-to-dataset
    Gateway/obs Portkey, Helicone Routing focus; Braintrust eval/CI gates deeper
    CI prompt test Promptfoo YAML/red-team OSS; Braintrust production-trace feedback loop
    Runtime guardrails Galileo, Guardrails Galileo runtime safety; Braintrust release-control + evals
    APM legacy Datadog LLM Obs, Fiddler General APM; Braintrust AI-trace-native
    In-list overlap Exa, Cody, JB AI, Continue, Aider, Cline, Mintlify, Q Dev(sunset), Phind(shutdown) Dev lane adjacent (eval vs coding/docs)

     

Do Not Sell or Share My Personal Information Cookie Settings