LLM Observability Cost
A real LLM observability cost model built from verified vendor pricing retrieved 2026-08-08, showing why Langfuse, LangSmith, Braintrust and Arize cannot be compared on list price alone and how the OpenTelemetry GenAI conventions' Development status quietly breaks naive cost attribution.
Key takeaways 5
Five things to know before pricing an LLM observability or evaluation platform against your own trace volume.
- Vendor pricing pages cannot be compared on price alone Langfuse, LangSmith, Braintrust, Arize and Confident AI meter LLM observability on three meter families, event-count, data-volume and volume-time. Comparable within a family, a league table across families without your own workload constants is dishonest.
- Build a cost model against your own trace volume The model needs three inputs: the judge model's token rate, the platform's metering unit and overage price and, for Braintrust, its per-score charge. State any token-per-trace figure as an assumption, never a citation.
- The Sonnet 5 pricing cliff ages judge-cost models fast Claude Sonnet 5's introductory pricing reverts from $2/$10 to $3/$15 per million tokens on 31 August 2026, and the newer tokenizer produces about 30% more tokens for the same text than earlier Claude generations.
- A conformant OpenTelemetry trace can carry no cost data Every `gen_ai.*` attribute in the OpenTelemetry GenAI conventions is status Development, and the token-usage attributes are Recommended, not Required, so a fully spec-conformant trace can legally ship with zero token counts.
- Evaluation traces double as compliance evidence Logged inputs, outputs and judge scores are exactly the kind of timestamped record EU AI Act Annex IV point 9 asks for, supporting technical documentation duties without automatically satisfying them on their own.
In short: Langfuse, LangSmith, Braintrust, Arize and Confident AI meter LLM observability using three families of billing meter, not five unrelated ones: event-count meters that tick per trace, span or score, data-volume meters that tick per byte and volume-time meters that tick per byte held over time. Meters inside the same family price out to a genuinely comparable number; meters in different families need a workload constant, such as spans per trace or bytes per trace, that only you can measure. Building a cost model against your own trace volume, not a vendor's tier name, is the only reliable way to know what any of them will actually cost you. A second risk sits underneath the pricing question: every `gen_ai.*` attribute in the OpenTelemetry GenAI semantic conventions is status Development, with only two Stable attributes on the page at all, both borrowed from the general conventions rather than GenAI-specific, and the token-usage attributes are only Recommended, so a standards-conformant trace can arrive with no cost data attached at all. Every vendor and model price quoted in this article was retrieved 2026-08-08 and carries its exact published tier name, because these prices move; the one regulatory citation later in the article carries its own retrieval date of 2026-07-28.
Why LLM observability vendors cannot be compared on price alone
These five tools do not bill on the same kind of meter, and converting one into another is not a rounding exercise, it is a missing conversion factor (Langfuse pricing, retrieved 2026-08-08). The right way to read that gap is by meter family, not by vendor name.
Three meter families, not five incomparable units
- Langfuse bills `units`, defined on its pricing page as the sum of traces, observations and scores ingested per billing period, a pure event-count meter
- LangSmith bills base and extended traces at the tier line, another event-count meter, but underneath that it also prices a normalized LangChain Storage Unit (LSU) at $1.00 each, a data-volume layer with no published traces-per-LSU conversion
- Braintrust runs all three families at once: scores billed per thousand, processed data billed per gigabyte and retention overage billed per gigabyte-month on Pro
- Arize bills spans ingested and gigabytes of storage as included-amount caps rather than priced meters, with no overage price published for either, and gives every tier including the free one unlimited evaluations
- Confident AI bills GB-months of trace spans, a pure volume-time meter where a longer retention window multiplies the bill the way an event-count meter never does
The finding that survives scrutiny is not that these five meters cannot be compared. It is that comparison inside a family is arithmetic, and comparison across families needs a workload constant the buyer must supply. Langfuse and Arize both count events, so the conversion between them is spans (observations) per trace, but that is not the same constant as Langfuse's own units-per-trace ratio: a Langfuse unit is traces plus observations plus scores combined, not spans alone, so the vendor's 6.98 units per trace is not 6.98 spans per trace. Langfuse and Confident AI count different physical things entirely, events versus bytes held over time, so the conversion is bytes per trace, then retention months. Each vendor section below states which constant a given comparison needs and what, if anything, the vendor itself publishes as an anchor for it.
A cost model normalized to your own trace volume
If the units do not convert to each other, the fix is to stop comparing units and start comparing your own bill at your own volume. That takes three inputs, all sourced from public pricing pages, plus one number you have to supply yourself. See how AI development cost breaks down more broadly for the surrounding engineering budget this fits inside.
The three inputs the model needs
First, the judge model's own token rate, from the model vendor's pricing page, not the observability vendor's page. Second, the platform's own metering unit and overage price, read directly off the relevant vendor table below. Third, for Braintrust only, the separate per-score charge, since it is the one tool here that bills scoring as its own line item rather than folding it into storage or trace count.
Worked example at three monthly trace volumes
No source in the pricing pages or the specification states how many tokens one LLM-as-judge call consumes. That number depends on prompt length, rubric and output format. Any article that quotes one as a fact is guessing. For this worked example, assume, purely as a stated planning assumption and never as a cited figure, that one judged trace costs the judge model about 1,500 input tokens and 150 output tokens, counted on the Claude 4.7-generation tokenizer, which is what Claude Sonnet 5 actually uses (see the tokenizer note later in this article). A token estimate calibrated on Claude Sonnet 4.6 or an earlier generation, which uses the older tokenizer, will understate spend by roughly 30% once carried over and applied to a 4.7-generation judge model like Sonnet 5.
At Anthropic's Claude Sonnet 5 introductory rate of $2 per million input tokens and $10 per million output tokens (detailed below under the Sonnet 5 pricing cliff), that assumption prices one judged trace at roughly $0.0045 in judge-model tokens alone:
- 10,000 judged traces a month: about $45 in judge-model tokens
- 100,000 judged traces a month: about $450 in judge-model tokens
- 1,000,000 judged traces a month: about $4,500 in judge-model tokens
That is the judge-model line only. Layer the platform's own meter on top using the vendor pricing tables below. On Braintrust, 100,000 scored traces a month is 90,000 billable scores on Starter after its 10,000 included scores, 90 blocks of 1,000 at $2.50 each, $225.00 in score overage on a $0 platform fee. On Pro the same 100,000 scores are 50,000 billable after 50,000 included, 50 blocks at $1.50 each, $75.00 in overage on top of the $249 platform fee, $324.00 all-in. Starter stays cheaper below roughly 199,000 scores a month; above that, Pro's lower per-block rate overtakes the platform fee (working shown in the Braintrust section below). On Langfuse, the same 100,000 traces convert using the vendor's own published formula, `Units = Count of Traces + Count of Observations + Count of Scores`, and its worked-example ratio of about 6.98 units per trace: roughly 698,000 units a month, stated as an assumption anchored to that one vendor example, not a general average (see the Langfuse table below for the resulting overage cost). One more variable belongs in the model before you trust it for a live system: LangSmith's own documentation states that reference-based evaluators require reference outputs and work only for offline evaluation (LangSmith evaluation concepts, retrieved 2026-08-08), so a reference-based judge cannot run against live traffic and its cost belongs in an offline test-suite budget.
One trace volume, five vendor meters: what actually computes
The clearest test of the three-family taxonomy is to take one stated monthly volume and try to price it against every vendor meter above. Take 100,000 traces a month, the same figure used above for the judge-model token estimate, and run it through each vendor's published unit, showing the assumption behind every computed cell.
| Vendor | Billing unit | Cost at 100,000 traces/month |
|---|---|---|
| Langfuse | `units` (traces + observations + scores) | Computable with a stated assumption: using the vendor's own worked-example ratio of about 6.98 units per trace, 100,000 traces convert to roughly 698,000 units/month. That is 598,000 units past the 100k included on any paid tier, $47.84 in overage at the flat $8/100k rate for the 100k-1M band (the pricing page publishes the rate but not the billing granularity, so if overage is billed in whole 100k blocks the figure is $48.00), so $76.84/month all-in on Core, $246.84/month on Pro, $2,546.84/month on Enterprise; Hobby's 50k-unit cap has no overage available and cannot carry this volume (Langfuse Cloud pricing, retrieved 2026-08-08) |
| LangSmith | `traces` (base, 14-day retention) | Not computable: Developer includes 5k base traces/mo and Plus includes 10k, so 90,000 to 95,000 of the 100,000 traces fall into pay-as-you-go territory priced only as "an additional fee"; the page publishes a $1.00 LSU rate but not the traces-per-LSU conversion needed to turn that into a dollar figure (LangSmith pricing, retrieved 2026-08-08) |
| Braintrust | scores, separate from processed-data GB and model credits | Computable, scores only: if every trace is also scored once, that is $225.00/month on Starter (90,000 billable scores after the 10,000 included, at $2.50/1k) or $324.00/month all-in on Pro ($75.00 overage after the 50,000 included, at $1.50/1k, plus the $249 platform fee); processed-data GB and model-credit charges add on top and depend on trace size and judge-token volume this article does not assume (Braintrust pricing, retrieved 2026-08-08) |
| Arize | spans plus GB storage | Not computable: AX Pro's 50k spans/month included already falls short of 100,000 traces assuming roughly one span per trace, and no overage price is published anywhere on the page for spans or storage past the included amount (Arize pricing, retrieved 2026-08-08) |
| Confident AI | GB-months of trace spans | Computable with a stated assumption: using the vendor's own published conversion, roughly 100K traces per GB, 100,000 traces a month ingests close to 1 GB. The Free tier's 1 GB-month is a volume-time meter, not a one-time cap, so at a steady 100,000 traces a month the retained data accumulates: month one is already close to using up the 1 GB-month allowance, and month two exceeds it, with anything past that edge dropped rather than billed. Starter's 5 GB-months for $200/month absorbs that accumulation with headroom (Confident AI pricing, retrieved 2026-08-08) |
Three of the five rows above now compute to a dollar figure, once the vendor's own published formula or conversion anchor is applied as a stated assumption. The other two, LangSmith and Arize, still cannot be priced at this volume, and for a documented reason: LangSmith publishes an LSU rate but not the traces-per-LSU conversion, and Arize publishes span and storage caps with no overage price on record at all. That split, not a blanket claim that none of this is comparable, is the honest state of a five-vendor comparison: inside a meter family the arithmetic works, across families it needs a constant only the buyer can measure, and two of these five vendors do not give the buyer even that constant to work with.
One more assumption sits underneath every number above, including the table: it prices 100% trace ingestion. Production systems rarely ingest every trace, and the sampling rate you actually run is a direct multiplier on every figure in this section. If a number above looks too high, sampling rate is the first thing to check, before switching vendors.
Vendor pricing tables, retrieved 2026-08-08
All prices below are unauthenticated cloud list prices retrieved 2026-08-08. Negotiated, annual-commit and regional pricing are out of scope.
Langfuse Cloud and self-hosted
Langfuse Cloud's tiers include the same amount of usage on every paid plan. Core to Pro buys longer retention, 90 days to 3 years, alongside throughput and support; Pro to Enterprise buys higher ingestion throughput, custom governance and support, not more retention, since Enterprise retention matches Pro at 3 years (Langfuse Cloud pricing, retrieved 2026-08-08).
| Tier | Price | Included | Overage | Retention |
|---|---|---|---|---|
| Hobby | Free | 50k units / month included | not available | 30 days data access |
| Core | $29/month | 100k units / month included | $8/100k units (lower with volume) | 90 days data access |
| Pro | $199/month | 100k units / month included | $8/100k units (lower with volume) | 3 years data access |
| Enterprise | $2,499/month | 100k units / month included | $8/100k units (lower with volume) | 3 years data access |
Langfuse does define the unit: a billable unit is any tracing data point sent to the platform, meaning the trace itself, every observation inside it (span, event or generation) and every score. The docs page publishes the formula, `Units = Count of Traces + Count of Observations + Count of Scores`, with a worked example of 20,070 traces, 119,500 observations and 561 scores totalling 140,131 units, about 6.98 units per trace on that example. Observations, meaning every span, event and generation inside a trace, drive the bill far more than the trace count itself, so "100k units / month included" is not 100k application interactions included; on the vendor's own example ratio it is closer to 14,000 traces (Langfuse pricing, retrieved 2026-08-08).
The overage ladder is graduated: $8.00 per 100k units from 100k to 1M, $7.00 per 100k units from 1M to 10M, $6.50 per 100k units from 10M to 50M, and $6.00 per 100k units above 50M (Langfuse Cloud pricing, retrieved 2026-08-08).
| Tier | Price | Notable |
|---|---|---|
| Open Source (Free) | Free | MIT License, unlimited units and usage, Enterprise SSO and RBAC included, community support |
| Self-Hosted Enterprise (Custom Pricing) | Custom | management APIs, project-level RBAC, data retention policies, audit logs, SOC 2 Type II and ISO 27001 reports |
The free self-hosted tier already includes Enterprise SSO and RBAC (Langfuse self-host pricing, retrieved 2026-08-08).
LangSmith base traces versus extended traces
| Tier | Price | Included traces |
|---|---|---|
| Developer | $0/seat per month plus pay as you go | Up to 5k base traces/mo, then pay-as-you-go |
| Plus | $39/seat per month plus pay as you go | Up to 10k base traces/mo, then pay-as-you-go |
| Enterprise | Custom pricing plus pay as you go | Custom |
Base traces carry 14-day retention, extended traces carry 400-day retention, and the pricing FAQ says only that you can "upgrade" base traces to extended traces "for an additional fee" with no per-trace dollar figure published anywhere on the page (LangSmith pricing, retrieved 2026-08-08). What the page does publish is a normalized unit layer underneath the trace tiers: LangChain Compute Units (LCU) at $1.50 each for work done and compute, and LangChain Storage Units (LSU) at $1.00 each for traces and storage, with the FAQ defining an LSU as a normalized unit of data stored or managed by the LangSmith platform, used because different LangSmith services meter at different rates. The page does not publish how many base traces consume one LSU, so a per-trace dollar figure still cannot be derived from it, and the on-page calculator itself carries the disclaimer that actual LCU/LSU consumption varies with the work an agent performs.
Braintrust's three-way metering
Braintrust is the only tool here that meters LLM-as-judge scoring as its own billable unit, separate from data volume and separate from the tokens the judge model itself burns (Braintrust pricing, retrieved 2026-08-08). The metered unit names on Braintrust's page are Processed Data and Model Credits, alongside scores; there is no unit called "judge tokens" anywhere on that page.
| Tier | Price | Model credits | Processed data | Scores | Retention |
|---|---|---|---|---|---|
| Starter | $0/month | $10 credits plus token rates | 1 GB processed data, $4/GB overage | 10k scores, $2.50/1k overage | 14-day retention |
| Pro | $249/month | $249 credits plus token rates | 5 GB processed data, $3/GB overage | 50k scores, $1.50/1k overage | 30-day retention plus $0.50/GB/month |
| Enterprise | Custom pricing | - | - | - | Custom data retention and export, RBAC and premium support with on-prem or hosted deployment |
At 100,000 scored traces a month, assuming one score per trace, Starter's post-allowance overage is 90,000 billable scores (100,000 minus the 10,000 included), 90 blocks of 1,000 at $2.50, $225.00 total on a $0 platform fee. Pro's post-allowance overage is 50,000 billable scores (100,000 minus the 50,000 included), 50 blocks at $1.50, $75.00 in overage plus the $249 platform fee, $324.00 all-in; quote that as "$75 of score overage on top of the $249 platform fee", not "$75 total". Setting the two cost lines equal, $0.0025 x (S minus 10,000) = $249 + $0.0015 x (S minus 50,000), gives a crossover at S = 199,000 scores a month, where both plans cost $472.50; below that volume Starter is cheaper all-in, above it Pro is. Neither figure includes processed-data GB, model credits the judge itself burns or Pro's $0.50/GB/mo retention overage.
Arize AX and Phoenix
Arize gives every AX tier, including the free one, unlimited evaluations and charges instead for span ingest and storage (Arize pricing, retrieved 2026-08-08). That is the inverse of Braintrust, which meters scoring directly.
| Tier | Price | Spans | Storage | Retention | Evals |
|---|---|---|---|---|---|
| AX Free | Free | 25k spans per month | 1 GB per month | 15 days | Unlimited |
| AX Pro | $50 per month | 50k spans per month | 10 GB per month | 30 days | Unlimited |
| AX Enterprise | Custom | Custom | Custom | Custom | Unlimited |
No overage price is published on the page for spans or storage past the included amounts, so whether overage is billed, blocked or negotiated per account is not something this article can confirm. Phoenix, Arize's open-source sibling, is described on the same page as "local-first" and free of license cost, which is a licensing fact, not a claim about the infrastructure bill of running it yourself.
Confident AI / DeepEval
| Tier | Price | Trace spans | Overage |
|---|---|---|---|
| Free | Forever $0 | 1 GB-month of trace spans | additional trace spans are dropped |
| Starter | $200/month | 5 GB-months of trace spans | then $1 per GB-month ingested or retained |
| Team | $2,000/month | 75 GB-months of trace spans | then $1 per GB-month ingested or retained |
| Enterprise | Custom pricing | Unlimited GB-months of trace spans | - |
Confident AI's bill accrues per gigabyte held per month (Confident AI pricing, retrieved 2026-08-08). The vendor's own calculator publishes a rough conversion, roughly 100K traces per GB, about 10 KB per trace, the anchor used for the computable Confident AI cell in the table above; treat it as that vendor's own rule of thumb, not a typical value for every workload. Below the Free tier's 1 GB-month, overflow is not billed at an overage price, it is dropped: the page states additional trace spans past that limit are simply discarded, a materially different failure mode from an unpriced overage.
LLM-as-judge token cost and the Sonnet 5 pricing cliff
Every judge-cost model in this article rests on a model vendor's own token price, and that price is not fixed.
Claude Sonnet 5 introductory pricing ends 31 August 2026
Anthropic's own pricing page states it plainly: introductory pricing of $2 per million input tokens and $10 per million output tokens is in effect for Claude Sonnet 5 through 31 August 2026, after which standard pricing of $3 per million input tokens and $15 per million output tokens takes effect (Anthropic model pricing, retrieved 2026-08-08). Note the citation: use `platform.claude.com`, not `anthropic.com/pricing`, which redirects to a consumer plans page carrying none of these per-model figures. That reversion is a 50% jump on both sides of the meter, and any judge-cost model built today on the introductory number and published without the date is stale before most readers finish evaluating the tool it recommends.
The Claude 4.7-generation tokenizer produces about 30% more tokens for the same text
Claude 4.7-generation models and later, including Claude Mythos Preview, use a newer tokenizer that produces approximately 30% more tokens for the same input text than the tokenizer used by Claude Sonnet 4.6 and earlier (Anthropic model pricing, retrieved 2026-08-08). If your judge-cost estimate carries over a token-per-word ratio calibrated on an older Claude generation, it understates spend on any newer model by roughly that margin before the per-token rate is even applied, which is exactly the adjustment the worked example earlier in this article flags for its own 1,500/150 token assumption. Stack the two facts together and a judge-cost figure computed today can be wrong on both axes within the same quarter. See how RAG versus fine-tuning choices affect the token count this section's rate gets multiplied against.
Why a spec-conformant OpenTelemetry trace can carry no cost data
OpenTelemetry does not guarantee the token counts a cost model needs, and the specification itself moved while making that clear. The old `opentelemetry.io` GenAI pages now return a stub notice that the conventions moved to their own repository (OpenTelemetry GenAI semantic conventions repository, retrieved 2026-08-08), so any older blog post or vendor doc citing the original path is pointing at dead documentation.
Every gen_ai.* attribute is status Development, not Stable
The specification's own status badge on the span attribute page reads Development (OTel GenAI spans spec, retrieved 2026-08-08). Among the attributes the spec tabulates, only two carry Stable status, and both are borrowed from the older general conventions rather than being GenAI-specific: `error.type` and `server.address`. Every `gen_ai.*` attribute in that table, including the two Required ones, `gen_ai.operation.name` and `gen_ai.provider.name`, is status Development. In OpenTelemetry's own terms, Development means not stable and subject to change.
Token-usage attributes are Recommended, not Required
`gen_ai.usage.input_tokens` and `gen_ai.usage.output_tokens` both carry Requirement Level Recommended, not Required. A trace can fully satisfy the specification while still shipping with no token counts, because token usage was never made mandatory. A cost dashboard assuming OpenTelemetry instrumentation guarantees cost data is building on an assumption the spec does not make. The standard also defines agent-level metrics answering a related question, how many model calls one agent run costs: `gen_ai.invoke_agent.inference_calls` and `gen_ai.invoke_agent.tool_calls` (OTel GenAI metrics spec, retrieved 2026-08-08), carrying the same Development status. Teams instrumenting production agents rather than single-turn calls should read that gap alongside securing AI agents in production. This finding is confirmed at spec-status level and for the attributes tabulated in the spans and metrics documents, not row by row across every sub-page of the new repository.
Evaluation traces double as EU AI Act technical documentation evidence
The same evaluation traces a team logs for quality monitoring have a second use: they are evidence. Article 11(1) of Regulation (EU) 2024/1689 requires technical documentation to contain, at a minimum, the elements set out in Annex IV, and Annex IV point 9 calls specifically for a description of the system in place to evaluate the AI system's performance in the post-market phase, including the post-market monitoring plan required under Article 72(3) (Regulation (EU) 2024/1689, retrieved 2026-07-28). A logged input, output and judge score is exactly the kind of timestamped record that post-market monitoring evidence is built from. That is not a claim that any platform's export, on its own, satisfies Annex IV in full; it is a claim that traces already generated for cost and quality reasons also feed the evidence Article 11(1) requires. See AI Act technical documentation for what the full duty requires.
Self-hosting changes the calculus, not just the sticker price
Self-hosting is not a uniform lever across these vendors. Langfuse's Open Source tier is free, MIT licensed and includes unlimited units alongside Enterprise SSO and RBAC (Langfuse self-host pricing, retrieved 2026-08-08). Arize's Phoenix is open-source and local-first, with no license cost. LangSmith and Braintrust are the opposite, restricting self-hosting to a custom Enterprise tier with no self-serve option. Confident AI's Enterprise tier mentions custom data residency and HIPAA, but whether DeepEval or the hosted platform can be run fully self-hosted below Enterprise was not confirmed in this article.
None of that removes infrastructure cost. "Local-first" and "free of license" describe what you do not pay a vendor, not what it costs to run storage and compute yourself; no operational-cost figures were retrieved for any of these platforms. What self-hosting buys, beyond avoiding a metered bill, is control over where data physically sits, which matters when the driver is regulatory rather than financial. If data residency is the driver, read that requirement against EU AI Act compliance before assuming self-hosting alone closes the gap.
A vendor selection checklist for the team paying the bill
There is no single winner in this list, because the right vendor depends on which meter matches your own traffic shape, not on which vendor has the lowest sticker price on a tier you may never fit into.
- Identify your dominant cost driver first: trace volume, storage duration, evaluation frequency or judge-scoring volume, then match the vendor whose meter prices that driver, not the cheapest headline number
- If pricing Langfuse, apply the published formula, units = traces + observations + scores, against your own observations-per-trace ratio, not the vendor's single worked example, before budgeting against it
- If your evaluation workload leans heavily on LLM-as-judge scoring, price Braintrust's per-score rate against your actual monthly score count after its included allowance, not before
- Match retention windows to your audit and compliance requirements before price, since a short window can force an expensive upgrade later
- Confirm self-hosting is genuinely available on the tier you can afford; LangSmith and Braintrust both gate it behind custom Enterprise pricing
- If evaluation data also feeds a downstream fine-tuning loop, the LLM fine-tuning guide covers what that data needs to look like before it is useful for training
How Pharos Production supports LLM observability and evaluation builds
Pharos Production helps engineering teams design and operate the observability and evaluation stack this article prices out, from instrumenting production traces through to the judge models and dashboards a platform team runs day to day, in line with the practices covered in the state of production AI engineering. That capability spans three service areas: MLOps for the pipelines carrying traces and metrics at production scale, AI product engineering for the evaluation loops that decide what ships, and LLM integration for wiring judge models and observability platforms into an existing application without duplicating billing meters you already pay for.
Sources
This article is engineering guidance, not financial or legal advice, and every vendor and model price above is a cloud list price retrieved without a login on 2026-08-08; the one regulatory citation, Regulation (EU) 2024/1689, was retrieved 2026-07-28, as noted where it appears in the text above. Negotiated, annual-commit, startup-program and regional pricing are out of scope and were not verified. No statistic here was invented beyond the explicitly stated assumptions in the worked cost-model examples. Primary sources:
- OpenTelemetry GenAI semantic conventions repository
- OTel GenAI spans specification
- OTel GenAI metrics specification
- Langfuse Cloud pricing
- Langfuse billable units docs
- Langfuse self-host pricing
- LangSmith pricing
- LangSmith evaluation concepts
- Braintrust pricing
- Arize pricing
- Confident AI pricing
- Anthropic model pricing
- Regulation (EU) 2024/1689 (EUR-Lex)
FAQ
Quick answers to common questions about custom software development, pricing, process and technology.
Type to filter questions and answers. Use Topic to narrow the list.
Showing all 7
No matches
Try a different keyword, change the topic or clear filters
-
You can't compare them directly because the five vendors meter LLM observability on three different families of billing meter, not one shared unit. Langfuse and LangSmith tick per event, such as a trace, observation or score.
Braintrust and Arize tick per byte ingested. Confident AI ticks per byte held over time. Meters inside the same family price out to a real number. Comparing across families needs a workload constant, such as spans per trace or bytes per trace, that only you can measure from your own traffic, not one any vendor publishes for you.
-
Claude Sonnet 5's introductory pricing of $2 per million input tokens and $10 per million output tokens stays in effect only through 31 August 2026. After that date, Anthropic's standard pricing takes over at $3 per million input tokens and $15 per million output tokens, a 50% increase on both sides of the meter.
Any LLM-as-judge cost model built on the introductory rate needs updating before that reversion, per Anthropic's model pricing page retrieved 2026-08-08.
-
No, OpenTelemetry does not guarantee token cost data in your traces. Every `gen_ai.*` attribute in the GenAI semantic conventions carries status Development, not Stable, and the two token-usage attributes, `gen_ai.usage.input_tokens` and `gen_ai.usage.output_tokens`, are Requirement Level Recommended rather than Required.
A trace can fully satisfy the specification while shipping with no token counts at all, so a cost dashboard built on the assumption that OTel instrumentation guarantees cost data is building on ground the spec never promised.
-
An LLM-as-judge evaluation's cost combines three inputs: the judge model's own token rate from the model vendor's pricing page, the observability platform's metering unit and overage price and, for Braintrust specifically, its separate per-score charge. No source publishes how many tokens one judged trace consumes, so any token-per-trace figure in a cost model must be stated as your own planning assumption, never cited as fact, then multiplied by your actual monthly trace volume.
-
Self-hosting is worth considering only once you weigh it against the infrastructure cost it does not remove. Langfuse's open-source tier is free, MIT licensed and includes unlimited units, and Arize's Phoenix is also free and local-first.
LangSmith and Braintrust restrict self-hosting to a custom Enterprise tier with no self-serve option. Self-hosting removes the vendor's metered bill, not the cost of running storage and compute yourself, so it matters most when data residency, not price, is the driver.
-
Evaluation traces support EU AI Act documentation but do not automatically satisfy it on their own. Article 11(1) of Regulation (EU) 2024/1689 requires technical documentation covering the elements in Annex IV, and Annex IV point 9 calls for a description of the system used to evaluate performance in the post-market phase.
A logged input, output and judge score is exactly the kind of timestamped record that evidence is built from, but no platform's export alone completes the full Annex IV duty.
-
There is no single vendor that wins on cost predictability, because the right choice depends on which meter family matches your own traffic shape, not on which vendor lists the lowest headline price. Identify your dominant cost driver first, trace volume, storage duration, evaluation frequency or judge-scoring volume, then match it against the vendor whose meter actually prices that driver.
Confirm self-hosting is genuinely available on a tier you can afford before assuming it changes the calculus.
I work with startup founders who need a dedicated software development team but don’t want to gamble on hiring, random outsourcing, or opaque delivery.
Most founders face the same problem sooner or later.
Early technical and team decisions lock the product into tech debt, slow delivery, missed milestones and constant re-hiring. By the time this becomes visible, fixing it is already expensive.As a CTO and software architect, I help founders design, build and run dedicated development teams that work as a true extension of the startup. Not as a black-box vendor.
My focus is on complex products where mistakes are costly:
- Web3 and blockchain platforms
- FinTech and regulated products
- High-load startup systems
- MVP → scale transitions
We don’t do body-shopping.
We don’t sell generic outsourcing.Instead, we help founders:
- build the right team structure from day one
- keep technical ownership and transparency
- scale delivery without losing control
- avoid vendor lock-in and hidden risks
Teams are aligned with the product roadmap, business goals and long-term architecture. Not just short-term velocity.