Metrics costs: keep the labels your investigations need

Compare metrics costs without losing useful detail. Test labels, sample resolution and query results before changing collection or moving to another backend.

Illustrative metrics comparison: whole-service error rate 1%, instance A 5% and instance B 0%, showing why the instance label matters.
Illustrative example from this article, not a product screenshot or benchmark. The same workload has a 1% service-wide error proportion, 5% for instance A and 0% for instance B.

A useful metrics saving leaves your team able to investigate the failures it needs to resolve. Before changing collection or buying a different backend, specify the labels, sample resolution, retention and queries you need. Compare the cost of meeting those requirements, then test whether the proposed change preserves them.

At xScaler, we start with the workload behind the bill. That might lead to a backend change, a collection change or both. The distinction matters: a lower price for the same data and a lower bill from collecting less data are different results.

Existing guidance explains how to reduce volume. Grafana's cardinality guide covers finding expensive metrics and reviewing labels. Google Cloud's cost controls cover collection frequency and local aggregation. Grafana's Adaptive Metrics documentation describes recommendations based on dashboards, alerts, recording rules and queries.

Those are useful ways to find candidates. This article focuses on how to decide which changes you can accept, and how to demonstrate that the remaining data still answers your operational questions.

What workload are you paying to retain?

Prometheus identifies a time series by its metric name and label set. Changing a label value creates a new series. Labels therefore need to be assessed in the combinations that actually occur, as described in the Prometheus data model.

A label with 100 possible values multiplies an existing series count by 100 only when every original series occurs with all 100 values. Adding one fixed region value to each existing series does not automatically multiply the simultaneous count. During a transition, old and new label sets may both appear in retained data.

Sample frequency is a separate dimension. For 100,000 continuously active scalar series, assuming one sample per series at every collection interval:

Illustrative arithmetic, checked 2 October 2026. Assumes a full day of successful collection, with no additional series or duplicate samples. These are sample counts, not storage sizes or vendor prices.
Collection intervalCalculation per daySamples per day
15 seconds100,000 × 86,400 ÷ 15576 million
30 seconds100,000 × 86,400 ÷ 30288 million
60 seconds100,000 × 86,400 ÷ 60144 million

Moving from 15 seconds to 60 seconds reduces observations by 75% while leaving the assumed series count unchanged. The financial result depends on the billing model. Google Cloud documents per-sample charging for its managed Prometheus service, but that does not make a sample reduction an identical percentage reduction in every provider's total bill.

Retention adds another requirement: how long must those observations remain available, and at what resolution? Keeping every label while replacing original samples with coarser aggregates changes the evidence available later. Record any accepted reduction in resolution explicitly.

Which labels are worth keeping?

Start with questions your team must answer: which deployment introduced the failure, whether the problem is confined to a region, or which instance needs investigation. Attach each question to the metric and label that makes it answerable.

Proposed decision worksheet, 2 October 2026. These are examples, not a requirement to add every label to every metric.
Operational questionCandidate detail to preserveValidation task
Did failures begin with a deployment?A version or deployment dimension on the relevant metricCompare behaviour across a deployment window
Is the problem confined to one region?Region on the service metricIsolate the affected region in a query
Is one instance failing more often?Instance on request and error countersCalculate the error proportion for each instance

Review dashboards, alerts, recording rules and query history, then ask the service owner about investigations absent from that history. A label unused during a quiet week might still be part of an infrequent recovery procedure. This is a reason to review the decision, not a reason to retain every dimension indefinitely.

Grafana already supports Adaptive Metrics exemptions for metrics or labels that should be excluded from recommendations, including those needed for particular investigations. Preserving useful detail is not unique to xScaler. Our recommendation is to make that requirement explicit in the cost comparison.

Unbounded labels deserve separate attention. Prometheus's instrumentation guidance warns about the resource cost of large label sets. A unique request identifier might belong in a trace or log, with its own retention and cost, rather than on every metrics observation. Keep useful dimensions and measure their cost. High cardinality still consumes resources when the labels are justified.

What can a service-level average conceal?

Suppose two instances serve the same application. In this hypothetical workload, their average request rates over the same window are:

Illustrative scenario, checked 2 October 2026. Error proportion is 5xx responses divided by all requests. These are assumed values, not customer results or measured PromQL output.
ScopeRequests per second5xx responses per secondError proportion
Instance A200105%
Instance B80000%
Whole service1,000101%

The service's 1% result is correct. Instance A's 5% result answers a different question. If you keep only the service aggregate and have no other source of instance detail, you cannot use that aggregate to identify A as the affected instance.

A proposed cost reduction should therefore name the questions it preserves. In this example, comparing only the whole-service dashboard would miss the loss of instance-level diagnosis. An alert with a hypothetical threshold above 2% would also behave differently depending on whether it evaluates each instance or the service total.

For a counter with service, instance and status labels, the following query expresses the per-instance error proportion as a percentage:

promql
100 *
sum by (service, instance) (
  rate(http_requests_total{status=~"5.."}[5m])
)
/
sum by (service, instance) (
  rate(http_requests_total[5m])
)

Adapt the label names to your instrumentation. This example assumes matching denominator series and positive request rates. Missing numerator series and zero traffic need an explicit display and alert policy. Treat them as test cases before adopting the expression in production.

The Prometheus rate() documentation recommends calculating the rate before aggregation so counter resets remain detectable. The query follows that order. The illustrative table is arithmetic, not evidence that this query has been executed against a deployment.

Where should aggregation happen?

Query-time aggregation lets you choose the dimensions of a result while retaining the source series. For example:

promql
sum by (service) (
  rate(http_requests_total[5m])
)

This returns a service-level request rate. If instance and status labels remain stored, another query can still use them within the available retention and query limits. The aggregation reduces the result's dimensions, not the number of source series already stored. This is a PromQL capability, not an xScaler-specific compression mechanism.

Prometheus recording rules precompute expressions and store the results as new time series. They help with repeated queries, but keeping both the recorded result and its source data adds stored output. To reduce exported volume, you must also decide what source data stops being exported or retained. Google's cost-control guide describes an approach that sends aggregates while keeping raw data locally.

Choose the policy according to the question. Service aggregates may be sufficient for a long-term trend while per-instance data remains necessary for a shorter investigation window. If that is your policy, price and test both retention periods. Do not describe the result as preserving every original observation.

How does xScaler approach the backend cost?

xScaler covers collection, managed storage and query, exploration, dashboards and notifications across applications, infrastructure and AI, with central management of telemetry collectors. Metrics are one part of that platform.

Our metrics backend uses Grafana Mimir, with blocks in object storage and caches on the read path. Prometheus-compatible collection sends data through remote write, and PromQL provides the query interface. The upstream Mimir store-gateway documentation explains its block-querying and caching mechanisms.

We operate and tune that foundation, including reviewing infrastructure use and the cost of running the service. The financial effect of storage, caching and query configuration depends on the workload. An architecture description alone does not establish a percentage saving, a latency guarantee or unlimited cardinality.

For a proposed xScaler configuration, confirm retention, sample-resolution requirements and applicable limits. Compare those terms with your existing service before deciding whether changing collection is necessary. If the existing system already meets your needs for less, keeping it is a valid result.

What would prove that a saving preserves useful detail?

Agree the comparison before changing either system. Include these requirements in the workload profile:

Proposed comparison worksheet, 2 October 2026. No vendor pricing or universal acceptance threshold is assumed.
AreaHold constant or discloseEvidence to compare
CollectionMetric names, label sets, sample intervals, expected growth and series turnoverCoverage and ingestion results through deployments and busy periods
StorageRetention and resolution at each age of dataQueries over recent and older windows
Queries and alertsExpressions, time ranges, concurrency and evaluation settingsValues, missing results, latency and notification behaviour
ServiceAvailability needs, support, regional placement and operating responsibilitiesAgreed service scope and the work each party retains
Commercial termsBilling units, included usage, overages, discounts, currency and extra chargesMatched quotes or bills with assumptions attached

For self-hosting, include engineering work for upgrades, capacity planning and incident response alongside infrastructure spend. Separate cash expenditure from the value of existing staff time. A collection reduction should appear as a changed workload, even if both options become cheaper after it.

Once the commercial comparison looks useful, run a parallel evaluation:

  1. Send the same representative input to both paths, recording intentional filters and temporary duplicate-ingestion charges.
  2. Compare metric names, required label combinations, missing series, ingestion failures and sample coverage across matching time windows.
  3. Run the agreed queries during ordinary and busy periods. Include the per-instance failure example, restarts and zero-traffic cases where relevant.
  4. Check alert evaluation and notification behaviour before retiring the previous path. Record differences and agree which are acceptable.

Series totals alone are insufficient: the same count could contain different labels or missing required series offset by unrelated additions. Conversely, matching query results on one dashboard does not establish that raw samples were preserved.

Set tolerances before the test. Record query step, evaluation time, ingestion delay and cache conditions when comparing results or latency. Investigate each difference before attributing it to data loss. If older data has not been transferred or retained for the required period, mark historical parity as untested.

How do you connect Prometheus and Grafana?

An existing compatible Prometheus setup can send metrics to xScaler through remote write. Add the following minimal block to prometheus.yml, using your assigned metrics host, bearer token and tenant ID:

yaml · prometheus.yml
remote_write:
  - url: https://<your-metrics-host>/api/v1/push
    authorization:
      credentials: <token>
    headers:
      X-Scope-OrgID: <tenant-id>

Follow the remote-write configuration guide for queue settings and request limits. This fragment shows connection fields, not a complete production configuration.

For your own Grafana, add a Prometheus data source with https://<your-metrics-host> as the server URL. Set the custom headers Authorization: Bearer <token> and X-Scope-OrgID: <tenant-id>. Use the host root without /prometheus, as described in the Grafana connection guide.

Can I recover a removed label from an aggregate later?
If only the aggregate remains, it does not contain the discarded per-label breakdown. Recovery requires another retained source with that detail. Check how long such a source remains available before approving the change.
Does a 75% sample reduction mean a 75% saving?
No. The example reduces sample count by 75%, but the total bill depends on billing units, included usage and other charges. It also changes sample resolution, so it is not an unchanged-workload comparison.
Does adding a recording rule reduce storage?
The rule creates stored result series. Retaining those alongside all source series adds output. Reducing storage or exported volume requires a separate decision about which source data to retain or send.

External sources checked 2 October 2026. Examples and worksheets are illustrative, not benchmarks. The xScaler setup guides are linked in the connection section.