← Back to blog
    September 30, 202610 min read

    Langfuse for AI Observability: How to Monitor, Debug, and Optimize LLM Agents in Production

    Langfuse is the open-source LLM observability platform that turns AI agent debugging from archaeology into a trace lookup. What it costs, how it compares to LangSmith and Phoenix, and a rollout plan that gets you from zero to useful traces in an afternoon.

    LangfuseLLM observabilityAI agentsOpenTelemetryLLM tracingAEO

    Langfuse for AI Observability: How to Monitor, Debug, and Optimize LLM Agents in Production

    Langfuse is an open-source LLM engineering platform that records every trace of an AI agent — every prompt, tool call, token count, and dollar — so engineering teams can debug failures, track costs, and score output quality on real production traffic. It is MIT-licensed, free to self-host, and since January 2026 it is part of ClickHouse. If you run agents in production without a tool like it, you cannot fix a prompt chain you cannot see, and you cannot cut a cost you never measured.

    What is Langfuse?

    Langfuse is an observability platform built specifically for applications that use large language models. Where a traditional APM like Datadog tracks HTTP requests and database queries, Langfuse understands LLM-native objects: prompts, completions, tool calls, token usage, and dollar cost per generation. It was created in Berlin, went through Y Combinator's W23 batch, raised a $4M seed round in September 2023 led by Lightspeed Venture Partners, La Famiglia, and Y Combinator, and in January 2026 was acquired by ClickHouse — with the announcement confirming MIT licensing, self-hosting, and the cloud offering continue unchanged.

    The platform covers four things engineering teams need when an LLM app leaves the prototype stage:

    • Tracing: every agent run becomes a trace, rendered as a waterfall of nested spans — the top-level request, each LLM generation with its exact input and output, every tool call, and the latency and token cost of each step.
    • Prompt management: prompts stored as versioned artifacts, fetched at runtime, so you know exactly which prompt version produced which output and can roll out changes without a redeploy.
    • Evaluations: score outputs manually, with an LLM-as-judge, or against saved datasets, and replay production traces as regression tests.
    • Analytics: cost per trace, p50/p95/p99 latency, error rates, and score distributions broken down per user, session, model, or tag.

    Adoption is the strongest signal in the category. Langfuse reports over 50,000 companies using it and positions its tracing infrastructure as powering AI at 19 of the Fortune 50; independent trackers put its GitHub repository above 23,000 stars as of March 2026, the largest reach among open-source LLM observability tools. Canva's AI team publicly credits Langfuse for tracing and debugging its generative design features in production.

    Why do AI agents fail without observability?

    An agent is not a single model call. A research agent searches, evaluates results, and loops back for more searches. A coding agent writes code, runs tests, and iterates on failures. Each hop multiplies the failure surface: a bad retrieval in step two poisons the reasoning in step five, and none of it shows up in server logs.

    This is the pattern that keeps agents stuck in pilot purgatory. The team ships a demo, it works on the happy path, and then in production the agent occasionally burns through 40 tool calls to answer a simple question. Nobody knows which step blew the budget, which prompt version regressed, or which model is quietly costing four times what it did last month. Gartner's widely cited forecast that 40% of agentic AI projects will be canceled by 2027 is, at root, a forecast about teams that could not see what their systems were doing.

    The fix is mechanical: trace every run, attribute cost per step, score outputs on a sample of real traffic, and only then gate prompt changes behind dataset evaluations. Langfuse is one of the cheapest ways to get all four.

    What are the core features of Langfuse?

    Tracing: the debugging waterfall

    Tracing pays for the setup on day one. Every agent run becomes a trace with nested observations: LLM generations with full prompt and completion text, tool executions with arguments and results, and spans for your own logic. Traces group into sessions, attach to users, and accept arbitrary tags and metadata. Because Langfuse accepts OpenTelemetry as a first-class protocol — it acts as a native OTLP backend and handles trace context propagation — a multi-step agent's execution can be tracked across distributed services, not just inside one process.

    The debugging value is concrete: when an agent deletes the wrong file or loops on a failing test, the trace shows the exact sequence — which tool ran, with what arguments, after which model output. Instead of re-running the agent five times with added print statements, you open the trace and look.

    Cost tracking: unit economics, not just token counts

    Langfuse computes token usage and dollar cost per generation, then aggregates per trace, user, session, tag, or model. This is what turns cost from a monthly invoice surprise into unit economics: how much did a specific session cost when the agent went into a fifteen-step loop? Tagging runs with session_id and user_id lets you aggregate spend per tenant or per customer, which is the number that actually matters if you are monitoring margins on an AI product.

    Evaluations: quality you can regression-test

    Traces alone tell you what happened; scores tell you whether it was any good. Langfuse supports manual annotation, LLM-as-judge scoring on production traces, and dataset-based evaluation where saved traces become test cases you replay whenever you change a prompt. A realistic sequence: start with shadow scoring on a sample of production traffic in week one, build a gold dataset from annotated real traces over the first month, and only move to gating deployments on eval scores once the judges agree with your human labels often enough to trust.

    Prompt management: version control for prompts

    Prompts in a config file are a deployment hazard. Langfuse stores prompts as versioned artifacts with deployment labels, fetched at runtime via SDK. You get a rollback path, and — more importantly for debugging — you can correlate output quality regressions with specific prompt versions, because every generation records which version produced it.

    Which frameworks integrate with Langfuse?

    Langfuse does not require rewriting your application. Instrumentation happens through SDK wrappers and OpenTelemetry:

    • OpenAI and Anthropic SDKs: a two-line wrapper that captures every completion with tokens and cost.
    • LangChain, LlamaIndex, and LiteLLM: callback handlers or OpenInference-compatible OTel spans — anything already emitting OTel can pipe straight in.
    • Workflow tools: n8n and Make users can log runs via the REST API or through a LiteLLM proxy sitting between the orchestrator and the model provider.
    • Coding agents: Langfuse ships documented support for tracing coding agents like Claude Code, giving teams cost per developer and an audit trail of what the agent actually did.

    Python and JavaScript SDKs are the primary clients; Go and Java work over OpenTelemetry. Most teams see their first traces the same afternoon they install the SDK.

    How much does Langfuse cost?

    Langfuse meters units — one unit equals one trace, one observation, or one score — with no per-seat fees. Unlimited users ship on Core and above, which is a real difference from LangSmith's $39-per-seat pricing.

    • Self-hosted open source: free, MIT-licensed, all core features with no usage limits. Deploy via Docker Compose, Kubernetes (Helm), or Terraform on AWS, GCP, or Azure. Organization-level RBAC and enterprise SSO are included in the free self-hosted distribution; project-level RBAC and SCIM require the Enterprise tier.
    • Cloud Hobby: free, 50,000 units per month, 2 users.
    • Cloud Core: $29/month, 100,000 units, unlimited users.
    • Cloud Enterprise: published at $2,499/month; self-hosted Enterprise adds management APIs, retention policies, audit logs, and SOC 2 / ISO 27001 reports, bundled with ClickHouse commercial plans.

    A useful worked comparison from Langfuse's public pricing model: at 500k traces with 10 spans and 1 score each, 5 users, 12-month retention, the bill comes to roughly $621/month on Langfuse Pro versus roughly $3,895/month on LangSmith Plus. The gap exists for a structural reason: LangSmith charges per seat plus per-trace usage and retention upgrades, while Langfuse bills only on usage units. Your mileage varies with the shape of your traffic, but the order-of-magnitude gap holds for seat-heavy teams.

    Langfuse vs. LangSmith vs. Phoenix vs. Helicone: which should you pick?

    The standalone AI observability layer consolidated hard in 2026: Langfuse joined ClickHouse in January, Helicone was acquired by Mintlify in March, and Arize agreed to be acquired by Dynatrace in August 2026 for $915 million. What survived as open, self-hostable infrastructure matters more than the logos:

    • LangSmith (LangChain): the tightest integration if you are all-in on LangChain/LangGraph, and the only one that also sells agent deployment infrastructure. But it is proprietary — self-hosting is an Enterprise add-on, and since May 2026 the US cloud runs on SmithDB, LangChain's proprietary storage engine. Costs scale with seats.
    • Arize Phoenix: open-source tracing and evals using OpenInference OTel conventions. Strong evaluation harness, notebook-friendly, a good pick if you want instrumentation-only. Now part of the Dynatrace orbit, which is either reassuring (backing) or concerning (open-source commitment) depending on your read.
    • Helicone: lightweight proxy-based logging, good for cost metering on an API key, thinner on evaluations and prompt management. Its acquisition by Mintlify makes it a question mark for net-new stacks.
    • Comet Opik: open-source tracing and evals from the experiment-tracking incumbent Comet — costs nothing new if you already run Comet.

    The honest decision rule: if your codebase is LangChain and you will never leave, LangSmith's zero-setup tracing is hard to beat. If you value owning your data, need framework-agnostic instrumentation, or your seat count makes per-seat pricing painful, Langfuse is the default. If evaluation-first regression testing is your primary need, look at Braintrust or DeepEval alongside.

    What about data privacy and compliance?

    Full-payload tracing means prompts and completions — which at many companies contain PII, secrets, and customer data — land in your trace store. This is the first question your InfoSec team will ask, and Langfuse answers it architecturally: self-hosting keeps trace data inside your own infrastructure (Docker Compose, Kubernetes, or Terraform on AWS/GCP/Azure), the platform supports server-side data masking to redact sensitive fields before storage, and trace data exports to your own S3 bucket. Cloud tiers ship with SOC 2 Type II and ISO 27001 reports. For HIPAA- or GDPR-constrained deployments, self-hosted with masking is the standard pattern.

    A realistic rollout plan

    • Day 1: self-host with Docker Compose (or start on Cloud Hobby), wrap your OpenAI/Anthropic client with the Langfuse SDK, and confirm traces arrive. Run an auth check — a typo'd key otherwise manifests as traces silently never showing up.
    • Week 1: add cost and latency dashboards; identify the most expensive step in your agent's chain. Start LLM-as-judge shadow scoring on a sample of production traces, without acting on the scores yet.
    • Month 1: build a gold dataset from annotated real traces. Move prompts into Langfuse's registry. Set alerts on cost per trace, p95 latency, and error rate.
    • Ongoing: once judge scores agree with human labels reliably, gate prompt changes behind dataset evaluations before they ship. Watch score drift. Treat the agent like a service, not a script.

    Conclusion

    The gap between an AI demo and an AI product is almost entirely a visibility problem. Langfuse closes it for free at the self-hosted tier, with an MIT license, OpenTelemetry-based instrumentation that fits whatever stack you already run, ClickHouse storage that scales to billions of events, and a pricing model that does not punish you for adding engineers. Debugging agent failures goes from archaeology to a trace lookup; cost goes from a monthly surprise to unit economics; quality goes from vibes to scores with regression tests behind them. You cannot optimize what you cannot see — and with tracing this cheap to add, there is no longer an excuse not to see.

    Frequently asked questions

    Is Langfuse free to use?
    Yes. Langfuse's core platform is MIT-licensed and free to self-host with no usage limits, including enterprise SSO and organization-level RBAC in the free self-hosted distribution (project-level RBAC and SCIM require the Enterprise tier). The managed cloud has a free Hobby tier with 50,000 units per month, and paid plans start at $29/month for the Core plan with 100,000 units and unlimited users.
    Is Langfuse open source?
    Yes. Langfuse is MIT-licensed open source, hosted on GitHub, and self-hosting is a first-class deployment mode running the exact same codebase as Langfuse Cloud. Since January 2026 the company is part of ClickHouse, which confirmed the MIT licensing and self-hosting continue unchanged.
    How does Langfuse compare to LangSmith?
    Langfuse is MIT open source with self-hosting on every tier, framework-agnostic OpenTelemetry instrumentation, ClickHouse storage, and no per-seat fees; LangSmith is proprietary, LangChain-native, meters $39 per seat per month, and restricts self-hosting to its Enterprise plan. At 500k traces with 10 spans each, Langfuse's published worked example comes to roughly $621/month on Pro versus $3,895/month on LangSmith Plus.
    What can you track with Langfuse?
    Langfuse tracks traces (full agent runs rendered as nested waterfalls of LLM generations and tool calls), token usage and dollar cost per step, user and session attribution, latency percentiles, error rates, output quality scores from LLM-as-judge or human annotation, and prompt version history, with everything queryable in dashboards or via SQL over ClickHouse.
    How do you integrate Langfuse into an existing LLM application?
    Install the Python or JavaScript SDK and wrap your OpenAI or Anthropic client, or add a callback handler for LangChain, LlamaIndex, or LiteLLM. Langfuse treats OpenTelemetry as its first-class protocol, so anything emitting OTel spans can send traces directly. Most teams see their first traces the same afternoon, with no rewrite of the application required.
    Can Langfuse be self-hosted?
    Yes. Langfuse self-hosts via Docker Compose, Kubernetes Helm charts, or Terraform on AWS, GCP, and Azure, running the same codebase as Langfuse Cloud. All core features — tracing, evaluations, prompt management, datasets — ship in the MIT distribution without usage limits, and trace data exports to S3 so you are never locked in.