The best AI agent observability tools in 2026 are Langfuse for open-source tracing, LangSmith for LangChain and LangGraph teams that want issues found and fixes proposed, Braintrust for evaluation-driven teams, Arize (Phoenix and AX) for OpenTelemetry-first teams, Datadog Agent Observability for companies already on Datadog, Laminar and Latitude for open-source failure detection, Raindrop for user-facing agents, and Opik for the lowest-cost open-source platform. If what you need is not more visibility but a ranked list of what to fix, Papaya works on top of any of them.
We checked every tool below against its own pricing page and documentation on October 1, 2026. This market moved fast this year: several tools were acquired, and most added some form of automatic failure detection.
What to look for in an agent observability tool
Observability for agents is not the same as observability for a single LLM call. A useful tool has to handle long, branching runs with many tool calls, and in 2026 it should do more than store them. These are the questions that separate the options.
- Does it model agent runs well? Nested spans for model calls, tool calls, sub-agents, and retries, with a view that makes a 40-step run readable.
- Does it find problems on its own? Scheduled or continuous detection of recurring failures, not just dashboards you have to remember to check.
- Can it evaluate? LLM-as-a-judge and code evaluators, online scoring of production traffic, datasets built from real traces.
- What does it meter? Traces, spans, gigabytes, events, or seats. One agent run can be one trace but dozens of spans, so the unit matters more than the headline price.
- Open standards and deployment. OpenTelemetry ingestion, open-source licensing, and whether self-hosting is free or Enterprise-only.
The tools at a glance
| Tool | Best for | License | First paid tier | Finds recurring failures automatically |
|---|---|---|---|---|
| Langfuse | Open-source tracing and prompt management | MIT core | $29/month | No built-in equivalent found |
| LangSmith | LangChain and LangGraph teams | Proprietary | $39 per seat/month | Yes: Engine, with proposed PRs |
| Braintrust | Evaluation-driven teams | Proprietary | $249/month | Topics; Patterns in preview |
| Arize Phoenix / AX | OpenTelemetry-first teams | Phoenix: Elastic License 2.0 | AX Pro: $50/month | AX: Signal (scheduled) |
| Datadog Agent Observability | Companies already on Datadog | Proprietary | $160/month (annual) for 100k LLM spans | Topic clustering and Insights |
| Laminar | Long-running and browser agents | Apache-2.0 | $30/month | Yes: Signals you define |
| Latitude | Open-source issue-to-fix loop | MIT | $99/month | Yes: signals, with hand-off to coding agents |
| Raindrop | User-facing agents | Proprietary | $299/month | Yes: Issue Detection on Pro |
| Opik (Comet) | Lowest-cost open-source platform | Apache-2.0 | $19/month | Ollie (interactive, on request) |
| Papaya | Ranked fixes per business use case | Proprietary | $500/month, 3 agents, 50k traces | Yes, per business use case, ranked by impact |
The tools in detail
Papaya
Papaya is built for the step after observability: deciding what to change. It reads traces in Langfuse, LangSmith, Arize Phoenix, and Braintrust formats, or collects them through its own SDK, and maps other formats automatically. It classifies every run by the business use case it served, compares succeeding and failing runs with 200+ research-backed analyses, and ranks findings by their impact on quality, latency, and cost. It also includes live observability dashboards, Slack alerts, an interactive LLM judge, and pull requests for approved fixes. Free for one agent and 1,000 traces a month; Business is $500 a month, never per seat.
Langfuse
The default open-source choice. Tracing, LLM-as-a-judge and code evaluators, datasets, experiments, prompt management with versioning, dashboards, and an OpenTelemetry endpoint, with an MIT-licensed core that is free to self-host. Cloud plans start at $29 a month with unlimited users. ClickHouse acquired Langfuse in January 2026 and has said no licensing changes are planned. It does not yet scan production on its own for recurring failures.
LangSmith
LangChain’s agent engineering platform, with tracing, online evaluations, annotation queues, datasets, and deployment. Its Engine, launched in May 2026, scans tracing projects on a schedule, clusters failures into prioritized issues, and can propose pull requests. Pricing is per seat, base traces are kept for 14 days, and self-hosting is an Enterprise add-on.
Braintrust
Built around one data model for production logs and experiments, which makes it the smoothest path from a bad trace to a regression test. It has online scoring, alerts, an OpenTelemetry endpoint, and Topics, which clusters traces by task, sentiment, and issue. Unlimited users on every plan; the jump from free to the $249 Pro plan is steep.
Arize Phoenix and Arize AX
Phoenix is a free, self-hostable, OpenTelemetry-native platform under the Elastic License 2.0. AX is the commercial product, from $50 a month, and adds Signal, which scans traces on a schedule and turns recurring failures into ranked issues with suggested fixes. Dynatrace completed its acquisition of Arize on October 1, 2026.
Datadog Agent Observability
Formerly LLM Observability. Its advantage is correlation with everything else Datadog already monitors: APM, logs, and infrastructure. It includes evaluations, datasets, topic clustering of production traffic, and automatic Insights on cost and reliability issues, priced per LLM span.
Laminar
An Apache-2.0, OpenTelemetry-native platform aimed at agents with long trajectories, with SQL access to traces and Signals: detectors you describe in plain language that run over traces and cluster their results. Paid plans start at $30 a month, with Signals billed as credits.
Latitude
MIT-licensed and self-hostable. It groups failed evaluations, annotations, and automatic flags into recurring signals with affected-user counts, and can hand an issue to a coding agent. Pro is $99 a month, billed in credits.
Raindrop
Focused on detecting silent failures in user-facing agents, with default and custom signals, issue detection across large volumes of interactions, experiments, and a triage agent available in Slack. The free plan covers 1,000 events a month; Pro is $299 a month plus usage, and full issue detection requires Pro. Raindrop announced a Series A in September 2026.
Opik
Comet’s Apache-2.0 platform: trace trees, datasets, experiments, LLM-as-a-judge metrics, online evaluation rules with alerts, and an agent optimizer. At $19 a month for 100,000 spans, it is the cheapest paid tier here.
The best options for startups
If you are a small team, three things matter more than feature lists: a free tier you can actually use, a reasonable first paid step, and a startup program.
- Langfuse offers 50% off for twelve months to companies that are bootstrapped or have raised up to $5 million and were incorporated in the last five years.
- LangSmith offers up to $10,000 in credits to VC-backed startups.
- Braintrust offers its Pro plan free for six to twelve months to qualifying startups, which removes its biggest drawback for small teams.
- Opik, Langfuse, and Laminar have the lowest first paid tiers, at $19, $29, and $30 a month.
Our honest advice for a startup: pick a tracing tool you will not outgrow in a year, instrument with OpenTelemetry where you can so switching stays cheap, and do not pay for features you do not have the time to use. The scarce resource at a startup is rarely data. It is the hours needed to read it.
Tools that need a caveat in 2026
Helicone was acquired by Mintlify in March 2026 and is in maintenance mode: security fixes and new models keep shipping, but new features do not. Galileo has been acquired by Cisco and is becoming Splunk Agent Observability, so check its standalone plans before committing. W&B Weave is now part of CoreWeave, which has launched Agent Lens for failure clustering across production traces.
Observability tells you what happened. Then what?
Every tool above makes agent behavior visible, and most now help you spot recurring problems. What they still leave to you is the hard part: working out which problems matter, why they happen, and what to change.
Papaya is built for that step. It reads traces in Langfuse, LangSmith, Arize Phoenix, and Braintrust formats, or collects them through its own SDK, and maps other formats automatically. It classifies every run into the business use case it served, compares succeeding and failing runs within each one using 200+ research-backed analyses, and quantifies each finding by its impact on quality, latency, and cost. You get the few fixes worth making first, not another dashboard, and with code access Papaya opens the pull request. We explain the difference in more depth in agent optimization vs. observability.
Try it on your own traces
Papaya’s free plan covers one agent and 1,000 traces a month. See pricing or talk to us.
Frequently asked questions
What is the best AI agent observability tool?
It depends on your stack and team. Langfuse is the leading open-source option, LangSmith suits LangChain and LangGraph teams, Braintrust suits evaluation-driven teams, Arize suits OpenTelemetry-first teams, and Datadog suits companies already using Datadog. Laminar, Latitude, and Raindrop focus on detecting failures automatically. Papaya suits teams that want a ranked list of fixes per business use case rather than more dashboards, and it works on top of traces from the other tools.
What is the best LLM observability tool for startups?
For most startups, Langfuse is the safest default: open source, $29 a month to start, and a 50% startup discount for a year. Opik is the cheapest paid option at $19 a month. Braintrust is a strong choice if you qualify for its startup program, which makes Pro free for six to twelve months.
What is the difference between agent observability and LLM observability?
LLM observability tracks individual model calls: prompts, outputs, tokens, latency, and cost. Agent observability tracks whole runs: sequences of model calls, tool calls, sub-agents, retries, and state changes, and whether the run achieved the user’s goal. Agent failures usually live in that sequence rather than in any single call.
Which agent observability tools are open source?
Langfuse (MIT core), Opik (Apache-2.0), Laminar (Apache-2.0), Latitude (MIT), and MLflow (Apache-2.0) are open source. Arize Phoenix is free to self-host under the Elastic License 2.0, which is source-available rather than OSI open source.
Do agent observability tools detect failures automatically?
Increasingly, yes. LangSmith Engine, Arize Signal, Braintrust Topics, Laminar Signals, Latitude, Raindrop, and Datadog Insights all surface recurring issues to some degree. They differ in whether detection is scheduled or on request, whether you must define the detectors, and whether they propose fixes.
Sources used in this article
- Langfuse pricing, startup program, and Langfuse joins ClickHouse.
- LangSmith pricing and Engine documentation.
- Braintrust pricing and Topics documentation.
- Arize pricing and Dynatrace completes acquisition of Arize.
- Datadog pricing list and Agent Observability documentation.
- Laminar pricing and Latitude pricing.
- Raindrop plans.
- Comet Opik pricing.
- Helicone joins Mintlify, Splunk on the Galileo acquisition, and CoreWeave Agent Lens.
Related
Keep reading
Article
Agent Optimization vs. Observability: Why Watching Isn't Fixing
Your dashboard saw the bad run and did nothing. The gap isn't between having data and not — it's between seeing what happened and knowing what to change.
Comparisons
Langfuse Alternatives (2026): 7 Tools Compared
Why teams leave Langfuse, what each alternative is best at, verified pricing and licensing as of October 2026, and when Langfuse is still the right call.

