LangSmith

  •  LangChain Inc., SF, founded 2022 by Harrison Chase, Series B 2025 (led IVP, with CapitalG/Sapphire/Sequoia/Benchmark, ~$125M). Agent engineering platform: tracing + evals + prompt hub + managed deployment for LangChain/LangGraph (framework-neutral SDK too). No acquisition, no sunset. Category ai-productivity-data, pipeline_stage=none. File under Dev​ (LLM-eval/observability sublane, sibling to Braintrust).Official site: https://www.langchain.com​ (LangSmith product) / direct: https://www.langchain.com/products/langsmith​Pricing: https://www.langchain.com/pricing

    Docs: https://docs.smith.langchain.com

    LangSmith — LangChain-Native Agent Engineering Platform for Tracing, Evals, Prompt Hub & Managed Deployment

    Snapshot

    Field Value
    Status Active (changelog Aug 2026: Engine >2x issue detection, Tuned Evaluators, BYOC GA on AWS, Managed Deep Agents beta)
    Origin LangChain Inc., SF, 2022, Harrison Chase
    Lane LLM/agent observability + evaluation + prompt registry + managed agent runtime
    Center of gravity Trace-first (production monitoring), eval integrated, LangChain/LangGraph zero-config
    Customers ~1,000 paying; 35% Fortune 500 incl. Workday, Cloudflare, ServiceNow, Klarna, Podium, Notion, Stripe-style
    Compliance SOC 2, HIPAA (Ent), GDPR, DPA, US/EU data residency, SSO/SAML/RBAC (Plus+), self-host/BYOC (Ent)
    Category ai-productivity-data / Dev (eval/obs sublane)

    Features (table)

    Layer Capability Notes
    Tracing Hierarchical traces (prompt/tool/retrieval/LLM/sub-agent), @traceable decorator, OTel-compatible, Py/TS/Go/Java SDK Zero-config auto-instrument for LangChain/LangGraph via env var
    Observability Insights (trace clustering, failure-mode detection), cost tracking, latency/token breakdown, monitoring + alerting Engine clusters traces → prioritized issues + suggested PR fix
    Evals Offline (dataset vs prompt/model versions) + online (live traffic scoring), LLM-as-judge / code / pairwise / human annotation queues Statistical grounding on score deltas (vs Braintrust gap)
    Prompt Hub Versioned prompts, env labels, Playground A/B, AI-assisted prompt rewrite (Polly), Assistants runtime config Runtime param injection without redeploy
    Deployment LangSmith Deployment: managed runtime for stateful LangGraph agents, durable exec, human-in-the-loop pause/resume, 1-click GH deploy, Fleet (no-code agents), Sandboxes (ephemeral code exec), expose agent as MCP server Only platform in this lane with agent runtime
    Engine Autonomous trace monitoring → cluster behaviors → diagnose → recommend prompt/code fix → auto-build offline eval LCU-metered ($1.50/LCU)
    Gateway LLM Gateway (Jul 2026): provider routing, cache, runtime controls Not a general APM
    No No OSS community edition (SDKs open, platform closed), no non-LLM APM, no pure score-metering like Braintrust

    Pricing (2026-09, USD)

    Tier Platform fee Included Overage / notes
    Developer $0 1 seat, 5,000 base traces/mo (14-day retention), community support Hard-capped 5k until CC added; no extended retention; no deployment/Fleet/Engine
    Plus $39/seat/mo Unlimited seats ($39 ea), 10,000 base traces/mo (14-day), 1 dev deployment (unlimited dev runs), Fleet 500/mo, Engine/Sandboxes/Insights, email support Base trace overage 5.00/1k; deployment run 0.0036/min prod; Fleet run 1.50/LCU; Sandboxes vCPU/GiB hourly
    Enterprise Custom (annual) Custom trace volume, SSO/SAML, SCIM, RBAC, audit, BYOC/self-host (AWS/GCP/Azure VPC), BAA, SLA, startup credits

    starting_price=$0; first paid 39 + trace meter + retention choice + optional Engine/Deployment/Fleet/Sandbox usage. LLM token cost billed by your model vendor, not LangSmith.

    Reviews (2026)

    Source Signal
    LangChain vs Braintrust page (vendor) LangSmith wins: LangChain-native zero-config tracing, Engine auto-fix PRs, managed LangGraph deploy, annotation queues with approval gates, statistical eval grounding
    genai.qa / aitoolsatlas / morphllm Praised: deepest LangChain/LangGraph instrumentation, full trace→eval→deploy loop, Fleet/Sandboxes. Criticized: per-seat × trace-meter double bill shocks scaling teams, non-LangChain stacks lose auto-instrument polish, some eval features need extra LLM spend, closed-source (vs Langfuse)
    TopReviewed panel 8.2/10 — “full trace-to-eval-to-deploy loop”, dinged on shallow attribution outside LangChain
    Case studies Klarna 80% faster resolution / 70% automation; Podium 98.6% F1 on agent eval

    Praised: LangGraph deploy + Engine flywheel + annotation workflow + US/EU residency. Complained: 2.50/1k trace compound cost, 5k free traces gone in 1–2 days in prod, lock-in perception, no OSS self-host.

    Access / Integration (table)

    Method Supported
    SDK Python, TypeScript, Go, Java
    Auto-instrument LangChain, LangGraph (env var), OpenAI/Anthropic/Vercel AI SDK via wrappers, OTel
    Deployment LangSmith Deployment (LangGraph agents), Fleet (no-code), Sandboxes, MCP server expose
    CI Custom GitHub Actions (less native than Braintrust’s block-merge action)
    Hosting Cloud (US/EU), Hybrid (SaaS control + self-host data), Fully self-hosted (Ent)
    No No REST “score text” free endpoint, no frontend product analytics

    Best For / Not For

    Best For Not For
    LangChain/LangGraph production agents (zero-config tracing) Framework-agnostic eval-first teams (→ Braintrust/Langfuse)
    Teams wanting managed stateful agent deploy + tracing + evals in one Tiny side-projects (5k trace cap dies in days)
    Enterprises needing SSO/SAML/SCIM/BYOC/self-host + SLA OSS-self-host-only mandates (→ Langfuse/Phoenix)
    Human-in-the-loop agent workflows (pause/resume, annotation queues) PM-only prompt shops unwilling to pay per-seat $39
    Klarna/Podium-scale trace volume with regression gates Non-LLM microservice APM (→ Datadog)

    Competitors (lane map)

    Lane Tool LangSmith edge / gap
    Eval-first platform Braintrust Braintrust scores-meter + Loop + framework-agnostic; LangSmith trace-meter + Engine auto-fix + LangGraph deploy
    OSS observability Langfuse, Arize Phoenix OSS self-host free; LangSmith managed + agent runtime + LangChain-native
    Prompt mgmt Humanloop, PromptLayer Lighter; LangSmith deeper trace→eval→deploy
    Gateway/obs Portkey, Helicone Routing focus; LangSmith eval/deploy broader
    CI prompt test Promptfoo YAML red-team OSS; LangSmith production flywheel
    APM legacy Datadog LLM Obs, Fiddler General APM; LangSmith agent-native + LangGraph runtime
    In-list overlap Exa, Cody, JB AI, Continue, Aider, Cline, Mintlify, Braintrust, Q Dev(sunset), Phind(shutdown) Dev lane adjacent (eval/obs vs coding/docs)
    Tags / Keywords

    langsmith,active-2026-09,llm-observability,agent-tracing,evals,langchain-native,langgraph-deploy,prompt-hub,engine-auto-fix,annotation-queues,fleet,sandboxes,mcp-expose,soc2-hipaa-gdpr,developer-free,plus-39-seat,series-b-125m,harrison-chase,not-shutdown,ai-productivity-data,pipeline_stage=none,competitors-braintrust-langfuse-phoenix-humanloop,dev-eval-sublane

    Want me to fold LangSmith (re-verified) into the locked slug/name/short_desc + structured fields + _event_log(none) batch alongside the eight active Dev/audio/video cards + Phind(shutdown) + Q Developer(sunset) + Braintrust?

Do Not Sell or Share My Personal Information Cookie Settings