The true cost of self-hosted observability during an incident

Work out what self-hosted observability costs when telemetry backs up, queries compete and alerts still need to arrive, using a practical recovery test.

Backlog-clearance chart: an 18 GB queue drains to zero in 150 minutes at 12 MB/s ingestion or 30 minutes at 20 MB/s, while fresh logs arrive at 10 MB/s.
Illustrative calculation, not a product screenshot or benchmark. A 30-minute outage queues 18 GB of decimal payload at 10 MB/s. Clearance times begin after recovery and assume complete buffering, constant fresh traffic and sustained ingestion, with no retry overhead.

Self-hosted observability costs include the capacity and engineering work needed to keep telemetry useful during a failure. To compare options fairly, first agree how much data must survive, how quickly it must become searchable and which alerts must arrive. Then price a deployment that meets those requirements.

A stack running Prometheus, Loki and Grafana might handle routine traffic comfortably. During an outage, the same deployment might need to store queued telemetry, catch up with fresh traffic and serve several engineers investigating at once. Its monthly infrastructure bill tells you what it consumed. Recovery tests tell you whether the capacity you paid for is enough.

The familiar costs are already well covered. Grafana's self-hosting guide describes upgrades, security, high availability and engineering time, although its scope is Grafana itself. Bleemeo's guide overview covers on-call work, replication, retention and opportunity cost. Parseable's cost analysis examines storage, compute, data movement and switching. These vendors sell observability services, so check their commercial comparisons against your workload and current quotes.

The connection less developed in the cost articles reviewed here is how a specific recovery requirement changes the bill. "Reliability" needs a measurable definition: which records survive, when they become usable and what resources achieve that result. That gives a platform team something concrete to test and price.

How much capacity does recovery need?

Take a log pipeline whose backend stops accepting data for 30 minutes. For this worked example, assume the following:

  • Logs arrive at a constant 10 MB/s, measured as payload bytes at the buffer boundary.
  • The collection path stays operational and retains every accepted payload during the outage.
  • Fresh logs continue arriving at 10 MB/s after the backend recovers.
Illustrative calculation, checked 2 October 2026. Assumed constant rates and complete buffering. MB and GB are decimal payload units, not provisioned disk capacity. No product or customer benchmark is represented.
QuantityCalculationResult
Payload queued during the outage10 MB/s × 1,800 seconds18,000 MB = 18 GB
Time to clear the backlog at 12 MB/s18,000 MB ÷ (12 − 10) MB/s9,000 seconds = 150 minutes
Capacity to clear it within 30 minutes10 MB/s + 18,000 MB ÷ 1,800 seconds20 MB/s

At 12 MB/s, fresh traffic uses 10 MB/s and leaves only 2 MB/s to drain the queue. The backend is reachable again after 30 minutes, but clearing the accumulated payload takes another two and a half hours. The age of individual records and the freshness of new data depend on scheduling, so measure backlog clearance and live-data freshness separately.

To clear that backlog within 30 minutes of recovery, the pipeline needs at least 20 MB/s of successful ingestion: 10 MB/s for new arrivals and 10 MB/s for queued data. This is a lower bound under constant rates. Retries, changing record sizes, backend limits and competing work affect the throughput a deployment sustains. If effective processing capacity stays at or below the arrival rate, the backlog never clears under these assumptions.

That recovery target has a cost. You might keep spare capacity running or scale up when the backend returns. Include the expense either way, and test how long additional capacity takes to become available.

Measure the full path before adding resources. OpenTelemetry's scaling guidance explains that more collectors do not resolve a saturated database or network. Adding workers might put more pressure on the overloaded destination.

What does telemetry buffering guarantee?

A useful buffering requirement names the outage duration and the failures it must survive. For example: "Retain the selected log stream through a 30-minute backend outage and a collector restart, then make the retained records searchable within the recovery target." Test that statement on the actual exporter, version, storage and backend configuration.

Version-specific behaviour matters. An OpenTelemetry Collector issue reported against version 0.92.0 described a persistent queue blocking start-up. A maintainer later pointed to a fix for version 0.100.0. That historical report gives a reason to repeat recovery tests after upgrades. It does not establish that current releases have the same defect.

Prometheus has its own limits. Its remote-write tuning guide warns that an extended destination outage leaves unsent samples at risk of loss after write-ahead-log compaction, describing a two-hour window for the documented path. Check the implementation and configuration in use before applying that duration to a deployment.

Size the buffer for the incident you intend to survive. If payload volume rises during the outage, the 18 GB estimate above is insufficient. Use observed peaks or stated stress assumptions, and decide which data receives priority when storage runs out.

Will queries and alerts work during recovery?

Engineers need to investigate while the pipeline catches up. Loki's query-fairness documentation describes contention between users sharing a tenant's query resources and mechanisms for allocating work across them. You still need to measure whether your deployment meets its latency target.

Build a query set from previous investigations, including the time ranges, label selections and searches engineers used. Run those queries with the expected number of concurrent investigators while replay continues. Measure latency, errors and result completeness separately. A fast query over an incomplete incident window has passed only the latency test.

Check alert delivery separately, from rule evaluation through to notification receipt. If this workload needs more query workers, cache capacity or isolation, include those resources in the estimate. A passing test with existing capacity is equally useful: it gives you evidence against buying resources you do not need.

Which failures does observability share with production?

Draw the dependencies between the application and its observability path: compute, networking, storage, identity, DNS and deployment tooling. For each failure scenario, establish which parts of collection, querying and notification remain accessible.

In its account of moving to managed observability, Clouds of Europe describes tuning and upgrade work, monitoring affected by management-cluster problems, and the control it gave up in the move. This is one organisation's experience. It supports checking shared dependencies, but does not supply a general staffing benchmark. A separate namespace alone is not evidence of independence, and the need for a second region depends on the failure you must survive.

Providers face the same design questions. In its November 2023 control-plane and analytics post-mortem, Cloudflare explained that logging services sat outside its high-availability cluster because delayed analytics had been considered acceptable. The report stated that undelivered Logpush data from the affected period would not be recovered. Its network and security services continued operating during the incident.

That incident illustrates the difference between restoring a service and recovering its missing data. It does not establish a failure rate for either hosting model. Price the separation your self-hosted design needs. For a managed service, establish where the provider's responsibility begins and which collection, network and access dependencies remain yours.

How do you compare the full cost fairly?

Start with operational targets the team agrees to meet. The worksheet below translates them into evidence and cost categories. The targets are yours to choose, rather than industry standards.

Proposed evaluation worksheet, 2 October 2026. Define targets for your workload before pricing either option.
RequirementEvidence to collectCost to include
Preserve selected telemetry through the chosen outageRecover unique test records and identify gaps and duplicatesBuffer storage, write throughput, persistence and any replication
Clear the backlog within the recovery targetMeasure backlog size and drain rate while fresh traffic continuesExtra ingestion capacity, transfer and scaling operations
Keep investigations usableMeasure query latency, errors and completeness during replayQuery resources, caches and workload isolation
Deliver alerts during the scenarioTime the test condition through to notification receiptEvaluation capacity, notification path and independent checks
Survive the chosen infrastructure failureTest with the relevant dependency unavailableRedundancy, separate dependencies and tested access paths
Repeat recovery after changesKeep versioned configuration and upgrade/recovery-test resultsTest environment and operating time

Build a monthly total from four categories, counting each cost once:

Cost-model structure, 2 October 2026. No vendor prices or staffing benchmarks are assumed.
CategoryWhat to countHow to calculate it
Recurring servicesRoutine infrastructure, spare capacity, subscriptions and supportMonthly bills or quotes for the required workload
Operating workMaintenance, upgrades, recovery exercises and relevant on-call workMeasured hours × fully loaded hourly cost
Setup and migrationOne-off work needed to put the option into serviceAllocate over a stated evaluation period
Additional test and recovery usageConsumption not already included in recurring servicesExpected usage × applicable rates

Allocate shared platform costs consistently. Where operating hours are uncertain, show a range and how it changes the result. Use a cash view for changes to invoices, hiring and overtime, alongside an economic view that also values existing staff time.

For managed services, include the collection and configuration work the customer retains, the support level required and charges for the test workload. Compare prices once both options meet the agreed requirements, or display unmet requirements beside the price.

Keep speculative incident losses outside the base total. A recovery test measures technical behaviour. It does not establish how many minutes of application downtime observability will prevent. If incident history supports a financial sensitivity analysis, count only additional impact plausibly attributable to impaired observability.

Can you reduce volume and keep the same service?

Reduce unnecessary telemetry before sizing either option. SigNoz's cost-control documentation describes analysing high-volume sources, applying reductions and validating the result. This work belongs before a purchasing decision when avoidable volume dominates the workload.

Record the filtering, sampling, retention and metric resolution used in the comparison, along with the incident questions the remaining data must answer. Shorter retention or more aggressive sampling changes the service being priced. Apply the same policy to both options.

Selective preservation is a reasonable design choice. You might prioritise particular error logs, metrics and traces while accepting loss elsewhere. State that policy and test it. Requiring every record to survive indefinitely adds cost without establishing business value.

When is self-hosting the better choice?

Self-hosting deserves to win when its measured operating work, infrastructure and recovery requirements cost less than suitable alternatives. Required control or customisation might also justify paying more.

An experienced team with reusable automation and a predictable workload should cost those advantages directly. Assigning an arbitrary fraction of an engineer to every deployment would conceal them. Give spare infrastructure an explicit cost allocation and name the person responsible for keeping that capacity available.

A small internal service with an accepted data-loss window needs a different design from a critical customer-facing service. Applying identical resilience requirements to both would distort the decision. Managed operation becomes more attractive when it meets the required service level for less, including the work the customer retains.

What should you test before approving the budget?

Use a representative test environment and the failure scenario you agreed to price. Record the resources and engineering hours consumed as you work through these steps:

  1. Interrupt the destination while the collection path continues receiving uniquely identified records.
  2. Restart a collector if restart survival is part of the requirement.
  3. Restore delivery and run the investigation queries while queued data drains.
  4. After the recovery deadline, reconcile the records and check notification delivery.

Use the results to identify the cheapest acceptable way to meet the requirements. A deployment tested only under routine traffic still has unverified incident performance. Include the work needed to establish that performance in its evaluation cost.

Does 18 GB of queued payload require an 18 GB disk?
Provisioned storage also needs space for queue encoding, metadata and operational margin. Compression and replication change physical storage requirements, so measure them in the intended deployment.
Do fewer maintenance hours mean a lower payroll bill?
Fewer hours release engineering capacity. Payroll falls only if the choice changes spending on staffing or overtime, which is why the cost model separates cash expenditure from the value of existing staff time.
Is backend availability enough to prove the data has recovered?
Check backlog clearance, live-data freshness and query completeness separately. A reachable backend might still be processing queued records, and the age of those records depends on scheduling.

Sources checked on 2 October 2026. Vendor articles describe existing cost coverage. Technical claims use project documentation and attributed incident reports. The recovery calculation is illustrative.