The Observability Tax
Why Your Logs Cost More Than Your Servers — And Datadog's Pricing Model Is Engineered to Hide It
Published: 2026-07-26 | jslet Research | 16 min read | Classification: Unrestricted
Executive Summary
Observability now consumes 15-25% of total cloud spend for unoptimized teams — often exceeding the infrastructure it monitors. The three pillars — logs, metrics, traces — each hide a structural pricing trap engineered to make the per-unit cost look trivial while the aggregate cost compounds silently.
Datadog: custom metric cardinality is billed at $0.10 per unique tag combination per month. A single metric tagged with customer_id across 1,000 customers generates 1,000 billable metrics — $100/month for one metric. Splunk: ingest-based pricing at $2-5/GB; 88-93% of that data is never queried after the first 7 days. Traces: 1% head-sampling statistically guarantees that during a latency incident, you will capture between zero and three P99 traces — not enough to identify which downstream service caused the spike.
Self-hosted Grafana stack (Loki + Mimir + Tempo) on a $500/month K8s node: same telemetry, 80-90% less than SaaS at scale. The breakeven is around 70-100 hosts. This briefing quantifies all five pricing traps, provides the self-hosted vs SaaS TCO model, and gives a concrete observability budget framework: target 5-10% of infrastructure spend.
Trap 1: Datadog Custom Metric Cardinality Explosion
Datadog bills $0.10 per custom metric per month. A custom metric is not just a metric name — it's a metric name + unique tag combination. Adding a customer_id tag to http.request.duration with 1,000 customers: 1,000 billable custom metrics = $100/month. Adding status_code (5 values) and endpoint (100 values) to the same metric: 1,000 × 5 × 100 = 500,000 potential combinations. If even 10% of combinations occur in a month: 50,000 billable metrics = $5,000/month for a single metric. The fix: restrict tags to low-cardinality dimensions only (region, service, status_code — <10 unique values each). Never tag with user_id, session_id, request_id, trace_id, or any identifier with unbounded cardinality. Those identifiers belong in log events (billed by GB ingested) or span attributes (filtered by trace sampling), not metric tags. A tag audit — pulling the list of all custom metrics from Datadog's API and checking for cardinality above 100 — takes 2 hours and typically finds $500-5,000/month in unintentional cardinality waste.
Trap 2: Splunk / Sumo Logic Ingest Model — Paying For Bytes You'll Never Query
Splunk charges $2-5/GB ingested. Sumo Logic: similar, with tiers for continuous vs frequent vs infrequent tier. The pricing model charges for data at ingestion time, not query time. You configure retention at 60-90 days for compliance. 93% of log data older than 7 days is never queried interactively — it's insurance against the compliance audit or the post-mortem that requires 60-day lookback. You pay full ingest price for 100% of the data, then store 93% of it untouched for months. The tiered-retention fix: 7 days hot in Splunk/Datadog for operational debugging, 90 days warm in S3 + Athena for compliance queries, 365 days cold in S3 Glacier for regulatory archiving. The hot tier might cost $100/month (handling 1 week of data). Cold tier costs $23/month per TB in S3 — 50× cheaper per GB than Splunk for data you'll query once per quarter. Total savings: 60-80% of the observability bill, same data retained, compliance checkbox checked.
Trap 3: Trace Head-Sampling — Blind To The Tail That Matters
At 1% head-sampling on 1,000 rps: you capture 10 traces/second. P99 events (10/second at 1K rps) appear in the sample at 10 × 0.01 = 0.1/second — one P99 trace every 10 seconds. During a 30-second latency incident, you might get 2-3 P99 traces with head-sampling. From those 2-3 traces, you cannot determine which downstream service's slowdown caused the spike — was it the database, the cache, the auth service, or the network? Tail-sampling observes the complete trace latency before deciding to keep it. The sampler keeps 100% of traces above a latency threshold (e.g., >500ms) or error traces. It head-samples the rest at a configurable rate. This guarantees every tail-latency trace is captured and available for root-cause analysis. The trade: tail-sampling requires the tracing system to buffer spans for a brief window to compute trace-level decisions — adding a small memory and latency cost. Grafana Tempo and Honeycomb support tail-sampling natively. Datadog APM supports it at higher tiers.
Trap 4: Self-Hosted vs SaaS — The TCO Model
At 200 hosts, 500 GB/day logs, 200K metrics, 1M spans/second: Datadog Pro (annual): ~$36,000/month. Self-hosted Grafana stack on 6 c6i.4xlarge K8s nodes (reserved, $780/month) + 50 TB S3 storage ($1,150/month) + 0.5 FTE SRE ($15K/month fully loaded) = ~$16,930/month — 53% less. At 50 hosts: Datadog ~$9,000/month, self-hosted ~$11,150/month (the 0.5 FTE dominates at smaller scale — SaaS is cheaper). At 500 hosts: Datadog ~$90,000/month, self-hosted ~$28,000/month — 69% less. Breakeven: ~70-100 hosts. The operational cost of self-hosted is not the compute. It's the engineering time to maintain Loki, Mimir, Tempo, Grafana, Promtail, and the underlying K8s infrastructure. If your team already manages a K8s cluster and has Prometheus expertise, the incremental operational cost of adding Loki/Mimir/Tempo is low — the breakeven drops to ~30-50 hosts. If you have no K8s or Prometheus expertise, the learning curve is steep and the SaaS premium is worth paying until you hire for it.
Trap 5: The Three-Pillar Cost Stacking Effect
SaaS vendors price logs, metrics, and traces as separate products with independent pricing dimensions. They do this because each pillar is a separate engineering system — but the pricing effect is stacking: logs at $X/GB + metrics at $Y/host + traces at $Z/span = $X+Y+Z/month. The total observability cost can exceed infrastructure cost. At 200 hosts on AWS at $50K/month infrastructure, an unoptimized Datadog deployment can easily reach $36K/month — 72% of infra spend. A well-optimized deployment with tiered retention, cardinality discipline, and tail-sampling can bring that to $9K/month — 18% of infra spend. The difference is $27K/month — $324K/year. That's more than the fully loaded cost of two senior SREs. The observability bill is the largest single optimization opportunity in most cloud budgets — and it's often the last one audited because the observability bill doesn't arrive on the same invoice as the infrastructure it monitors.
Concrete Steps: The Observability Cost Audit
1. Pull your actual observability spend as a percentage of infrastructure spend. Sum the monthly bills from Datadog, Splunk, New Relic, Grafana Cloud, and any other SaaS observability vendor. Divide by your AWS/GCP/Azure monthly infrastructure spend. If the ratio is above 15%, you have a cost optimization opportunity worth at least 40% of the observability bill.
2. Audit custom metric cardinality. In Datadog: Metrics → Summary → sort by distinct tag combinations. Any metric with >100 unique tag combinations is a cardinality candidate. Check whether all tag dimensions are necessary. Remove high-cardinality tags (user_id, session_id, request_id). Replace with log-based queries for those dimensions.
3. Implement tiered log retention. 7 days hot (interactive querying), 90 days warm (S3 + Athena), 365 days cold (S3 Glacier). This cuts the observability storage bill by 60-80%. Configure the log forwarder (Vector, Fluentd, Logstash) to route to S3 after the hot window. Schedule quarterly log index audits: delete indexes that haven't been queried in 90 days.
4. Switch from head-sampling to tail-sampling for traces. Keep 100% of traces above your latency SLO threshold (e.g., >500ms) and error traces. Head-sample the rest at 1-10%. This ensures you capture the traces that matter for incident debugging while keeping trace volume manageable.
5. Model self-hosted vs SaaS at your scale. Use the Log Storage TCO Estimator and Observability Cost Calculator. If your host count exceeds 70-100 and your team has K8s/Prometheus expertise, self-hosted Grafana stack will save 50-85% of the observability bill within the first year.
🧰 Use our related tools: Log Storage & Retention TCO · Observability Cost Calculator · Database Connection Pool · Managed Database Cost Illusion
Frequently Asked Questions
How does Datadog's custom metric cardinality actually drive costs?
Datadog bills $0.10 per custom metric per month — and a custom metric is defined by its unique metric-name + tag-values combination. A 'customer_id' tag with 1,000 unique customers = 1,000 billable metrics = $100/month for a single metric. Adding more tags multiplies: status_code (5 values) × endpoint (100 values) × customer_id (1,000) = 500,000 potential combinations. The cardinality explosion is the most common Datadog cost surprise. Fix: restrict tags to low-cardinality dimensions (<10 unique values). Use log events or span attributes for high-cardinality data. Audit your custom metrics monthly via Datadog's API. See the Log Storage TCO Estimator to compare Datadog vs alternatives.
How much of my log storage is actually queried?
88-93% of stored log data is never queried after the first 7 days. Most incidents resolve within hours. Logs older than 7 days serve compliance and post-mortems — rarely interactive debugging. The optimization: tiered retention. 7 days hot in the SaaS platform. The rest in S3 + Athena ($5/TB scanned per query). This cuts the observability storage bill by 60-80% without losing any data. Configure this once; the savings are permanent. Use the Log Storage TCO Estimator to model hot/cold tier costs.
Why doesn't 1% trace sampling capture the traces I need for debugging?
Head-sampling decides at trace start, before latency is known. At 1% sampling and 1,000 rps: you get ~10 traces/second. P99 events occur at ~10/second. The expected number of P99 traces captured: 10 × 0.01 = 0.1/second — about 3 P99 traces per 30-second incident. Not enough for root cause analysis. Tail-sampling observes the full trace latency, then keeps 100% of traces above a latency threshold. This guarantees every tail-latency trace is captured. Switch to tail-sampling in Grafana Tempo, Honeycomb, or Datadog APM (higher tiers).
What's the actual TCO of self-hosted Grafana stack vs Datadog?
At 200 hosts: Datadog Pro ~$36,000/month. Self-hosted (Loki+Mimir+Tempo on 6 K8s nodes + S3 + 0.5 FTE SRE): ~$16,930/month — 53% less. At 50 hosts: Datadog cheaper (SaaS simplicity wins below breakeven). At 500 hosts: self-hosted 69% cheaper. Breakeven: ~70-100 hosts. The self-hosted operational overhead is real — if you don't have K8s/Prometheus expertise, SaaS is the right choice up to ~150 hosts. If you do, self-hosted frees up budget for 1-2 additional SREs. Model your scale with the Log Storage TCO Estimator.
What percentage of my infrastructure spend should observability consume?
Well-optimized: 5-10% of infrastructure spend. Unoptimized: 15-25%. The gap on a $50K/month infrastructure bill: $2,500-5,000/month (optimized) vs $7,500-12,500/month (unoptimized) = $60,000-90,000/year difference. The 5-10% benchmark is achievable with tiered retention, cardinality discipline, tail-sampling, and regular audit of unused dashboards/alerts/indexes. The observability bill is the single largest optimization opportunity in most cloud budgets — audit it first, optimize it aggressively, and redirect the savings to hiring SREs who will keep it optimized.
Methodology & Disclosure
Pricing data from publicly available vendor rate cards accessed in July 2026. Datadog: Pro tier, annual commitment, $15/host/month (infrastructure), $0.10/custom metric/month, logs $0.10/GB ingested + $2.50/1M events. Splunk Cloud: ingest-based pricing at ~$3/GB (estimated; actual pricing is sales-negotiated). Grafana Cloud: free tier 10K metrics/50GB logs/50GB traces, Pro tier at $30-50/month for next tier. Self-hosted compute: c6i.4xlarge 3-year RI at $130/month/node (us-east-1). S3: Standard tier at $0.023/GB/month. Engineering time: $15K/month fully loaded (0.5 FTE — US market median for senior SRE, 2026).
Disclosure: jslet is an independent research project. This analysis was produced using our own Log Storage TCO Estimator and Observability Cost Calculator with publicly available pricing data. We are not sponsored by any observability vendor.
References & Further Reading
- Datadog (2026). "Pricing — Pro and Enterprise Tiers." datadoghq.com
- Splunk (2026). "Splunk Cloud Platform Pricing." splunk.com
- Grafana Labs (2026). "Grafana Cloud Pricing." grafana.com
- OpenTelemetry (2026). "Sampling — Head vs Tail Sampling." opentelemetry.io
- Grafana Tempo (2026). "Tail-based sampling." grafana.com
- AWS (2026). "CloudWatch Logs Pricing." aws.amazon.com
- jslet (2026). "Observability Cost Calculator — 5 Solutions Compared." jslet.com/observability-cost
📜 Copyright & Attribution
© 2026 jslet Research. This article is an original work independently researched and published on jslet. All rights reserved.
Sharing & Reprinting: You may share excerpts (up to 200 words) with a mandatory, do-follow link back to this article's canonical URL.
Preferred Attribution Format: "The Observability Tax (2026)" — jslet Research, July 2026. https://www.jslet.com/observability-cost-real
📡 Enjoyed this? When your observability bill exceeds your infrastructure bill, the pricing model was designed to hide it. RSS covers one pricing-model reality check per week. No vendor sponsors. RSS Feed → | More options →