Metrics costs: keep the labels your investigations need
Compare metrics costs without losing useful detail. Test labels, sample resolution and query results before changing collection or moving to another backend.

A useful metrics saving leaves your team able to investigate the failures it needs to resolve. Before changing collection or buying a different backend, specify the labels, sample resolution, retention and queries you need. Compare the cost of meeting those requirements, then test whether the proposed change preserves them.
At xScaler, we start with the workload behind the bill. That might lead to a backend change, a collection change or both. The distinction matters: a lower price for the same data and a lower bill from collecting less data are different results.
Existing guidance explains how to reduce volume. Grafana's cardinality guide covers finding expensive metrics and reviewing labels. Google Cloud's cost controls cover collection frequency and local aggregation. Grafana's Adaptive Metrics documentation describes recommendations based on dashboards, alerts, recording rules and queries.
Those are useful ways to find candidates. This article focuses on how to decide which changes you can accept, and how to demonstrate that the remaining data still answers your operational questions.
What workload are you paying to retain?
Prometheus identifies a time series by its metric name and label set. Changing a label value creates a new series. Labels therefore need to be assessed in the combinations that actually occur, as described in the Prometheus data model.
A label with 100 possible values multiplies an existing series count by 100 only when every original series occurs with all 100 values. Adding one fixed region value to each existing series does not automatically multiply the simultaneous count. During a transition, old and new label sets may both appear in retained data.
Sample frequency is a separate dimension. For 100,000 continuously active scalar series, assuming one sample per series at every collection interval:
| Collection interval | Calculation per day | Samples per day |
|---|---|---|
| 15 seconds | 100,000 × 86,400 ÷ 15 | 576 million |
| 30 seconds | 100,000 × 86,400 ÷ 30 | 288 million |
| 60 seconds | 100,000 × 86,400 ÷ 60 | 144 million |
Moving from 15 seconds to 60 seconds reduces observations by 75% while leaving the assumed series count unchanged. The financial result depends on the billing model. Google Cloud documents per-sample charging for its managed Prometheus service, but that does not make a sample reduction an identical percentage reduction in every provider's total bill.
Retention adds another requirement: how long must those observations remain available, and at what resolution? Keeping every label while replacing original samples with coarser aggregates changes the evidence available later. Record any accepted reduction in resolution explicitly.
Which labels are worth keeping?
Start with questions your team must answer: which deployment introduced the failure, whether the problem is confined to a region, or which instance needs investigation. Attach each question to the metric and label that makes it answerable.
| Operational question | Candidate detail to preserve | Validation task |
|---|---|---|
| Did failures begin with a deployment? | A version or deployment dimension on the relevant metric | Compare behaviour across a deployment window |
| Is the problem confined to one region? | Region on the service metric | Isolate the affected region in a query |
| Is one instance failing more often? | Instance on request and error counters | Calculate the error proportion for each instance |
Review dashboards, alerts, recording rules and query history, then ask the service owner about investigations absent from that history. A label unused during a quiet week might still be part of an infrequent recovery procedure. This is a reason to review the decision, not a reason to retain every dimension indefinitely.
Grafana already supports Adaptive Metrics exemptions for metrics or labels that should be excluded from recommendations, including those needed for particular investigations. Preserving useful detail is not unique to xScaler. Our recommendation is to make that requirement explicit in the cost comparison.
Unbounded labels deserve separate attention. Prometheus's instrumentation guidance warns about the resource cost of large label sets. A unique request identifier might belong in a trace or log, with its own retention and cost, rather than on every metrics observation. Keep useful dimensions and measure their cost. High cardinality still consumes resources when the labels are justified.
What can a service-level average conceal?
Suppose two instances serve the same application. In this hypothetical workload, their average request rates over the same window are:
| Scope | Requests per second | 5xx responses per second | Error proportion |
|---|---|---|---|
| Instance A | 200 | 10 | 5% |
| Instance B | 800 | 0 | 0% |
| Whole service | 1,000 | 10 | 1% |
The service's 1% result is correct. Instance A's 5% result answers a different question. If you keep only the service aggregate and have no other source of instance detail, you cannot use that aggregate to identify A as the affected instance.
A proposed cost reduction should therefore name the questions it preserves. In this example, comparing only the whole-service dashboard would miss the loss of instance-level diagnosis. An alert with a hypothetical threshold above 2% would also behave differently depending on whether it evaluates each instance or the service total.
For a counter with service, instance and status labels, the following query expresses the per-instance error proportion as a percentage:
100 *
sum by (service, instance) (
rate(http_requests_total{status=~"5.."}[5m])
)
/
sum by (service, instance) (
rate(http_requests_total[5m])
)Adapt the label names to your instrumentation. This example assumes matching denominator series and positive request rates. Missing numerator series and zero traffic need an explicit display and alert policy. Treat them as test cases before adopting the expression in production.
The Prometheus rate() documentation recommends calculating the rate before aggregation so counter resets remain detectable. The query follows that order. The illustrative table is arithmetic, not evidence that this query has been executed against a deployment.
Where should aggregation happen?
Query-time aggregation lets you choose the dimensions of a result while retaining the source series. For example:
sum by (service) (
rate(http_requests_total[5m])
)This returns a service-level request rate. If instance and status labels remain stored, another query can still use them within the available retention and query limits. The aggregation reduces the result's dimensions, not the number of source series already stored. This is a PromQL capability, not an xScaler-specific compression mechanism.
Prometheus recording rules precompute expressions and store the results as new time series. They help with repeated queries, but keeping both the recorded result and its source data adds stored output. To reduce exported volume, you must also decide what source data stops being exported or retained. Google's cost-control guide describes an approach that sends aggregates while keeping raw data locally.
Choose the policy according to the question. Service aggregates may be sufficient for a long-term trend while per-instance data remains necessary for a shorter investigation window. If that is your policy, price and test both retention periods. Do not describe the result as preserving every original observation.
How does xScaler approach the backend cost?
xScaler covers collection, managed storage and query, exploration, dashboards and notifications across applications, infrastructure and AI, with central management of telemetry collectors. Metrics are one part of that platform.
Our metrics backend uses Grafana Mimir, with blocks in object storage and caches on the read path. Prometheus-compatible collection sends data through remote write, and PromQL provides the query interface. The upstream Mimir store-gateway documentation explains its block-querying and caching mechanisms.
We operate and tune that foundation, including reviewing infrastructure use and the cost of running the service. The financial effect of storage, caching and query configuration depends on the workload. An architecture description alone does not establish a percentage saving, a latency guarantee or unlimited cardinality.
For a proposed xScaler configuration, confirm retention, sample-resolution requirements and applicable limits. Compare those terms with your existing service before deciding whether changing collection is necessary. If the existing system already meets your needs for less, keeping it is a valid result.
What would prove that a saving preserves useful detail?
Agree the comparison before changing either system. Include these requirements in the workload profile:
| Area | Hold constant or disclose | Evidence to compare |
|---|---|---|
| Collection | Metric names, label sets, sample intervals, expected growth and series turnover | Coverage and ingestion results through deployments and busy periods |
| Storage | Retention and resolution at each age of data | Queries over recent and older windows |
| Queries and alerts | Expressions, time ranges, concurrency and evaluation settings | Values, missing results, latency and notification behaviour |
| Service | Availability needs, support, regional placement and operating responsibilities | Agreed service scope and the work each party retains |
| Commercial terms | Billing units, included usage, overages, discounts, currency and extra charges | Matched quotes or bills with assumptions attached |
For self-hosting, include engineering work for upgrades, capacity planning and incident response alongside infrastructure spend. Separate cash expenditure from the value of existing staff time. A collection reduction should appear as a changed workload, even if both options become cheaper after it.
Once the commercial comparison looks useful, run a parallel evaluation:
- Send the same representative input to both paths, recording intentional filters and temporary duplicate-ingestion charges.
- Compare metric names, required label combinations, missing series, ingestion failures and sample coverage across matching time windows.
- Run the agreed queries during ordinary and busy periods. Include the per-instance failure example, restarts and zero-traffic cases where relevant.
- Check alert evaluation and notification behaviour before retiring the previous path. Record differences and agree which are acceptable.
Series totals alone are insufficient: the same count could contain different labels or missing required series offset by unrelated additions. Conversely, matching query results on one dashboard does not establish that raw samples were preserved.
Set tolerances before the test. Record query step, evaluation time, ingestion delay and cache conditions when comparing results or latency. Investigate each difference before attributing it to data loss. If older data has not been transferred or retained for the required period, mark historical parity as untested.
How do you connect Prometheus and Grafana?
An existing compatible Prometheus setup can send metrics to xScaler through remote write. Add the following minimal block to prometheus.yml, using your assigned metrics host, bearer token and tenant ID:
remote_write:
- url: https://<your-metrics-host>/api/v1/push
authorization:
credentials: <token>
headers:
X-Scope-OrgID: <tenant-id>Follow the remote-write configuration guide for queue settings and request limits. This fragment shows connection fields, not a complete production configuration.
For your own Grafana, add a Prometheus data source with https://<your-metrics-host> as the server URL. Set the custom headers Authorization: Bearer <token> and X-Scope-OrgID: <tenant-id>. Use the host root without /prometheus, as described in the Grafana connection guide.
- Can I recover a removed label from an aggregate later?
- If only the aggregate remains, it does not contain the discarded per-label breakdown. Recovery requires another retained source with that detail. Check how long such a source remains available before approving the change.
- Does a 75% sample reduction mean a 75% saving?
- No. The example reduces sample count by 75%, but the total bill depends on billing units, included usage and other charges. It also changes sample resolution, so it is not an unchanged-workload comparison.
- Does adding a recording rule reduce storage?
- The rule creates stored result series. Retaining those alongside all source series adds output. Reducing storage or exported volume requires a separate decision about which source data to retain or send.
External sources checked 2 October 2026. Examples and worksheets are illustrative, not benchmarks. The xScaler setup guides are linked in the connection section.
- Grafana: Managing high-cardinality metrics 2 Oct 2026
- Google Cloud: Managed Prometheus cost controls 2 Oct 2026
- Grafana: Adaptive Metrics 2 Oct 2026
- Prometheus: Data model 2 Oct 2026
- Grafana: Adaptive Metrics exemptions 2 Oct 2026
- Prometheus: Instrumentation guidance 2 Oct 2026
- Prometheus: rate function 2 Oct 2026
- Prometheus: Recording rules 2 Oct 2026
- Prometheus: Relabel configuration 2 Oct 2026
- Grafana Mimir: Store-gateway 2 Oct 2026