OpenTelemetry metrics best practices: 2027

Choose OpenTelemetry instruments, attributes, histogram resolution and collection settings that preserve the answers your dashboards and alerts need.

Illustrative service trace waterfall above a blue latency histogram
Editorial illustration, not output from the examples or a production dashboard.

OpenTelemetry metrics best practices start with the decisions a metric must support. Choose the right instrument, follow semantic conventions, preserve producer identity, budget useful dimensions and set histogram resolution deliberately. Then configure collection for the destination and verify the dashboard or alert against known workload behaviour.

For each production metric, record its owner, measurement boundary, unit, required breakdowns and consumer. A checkout latency SLO needs a defined request population and threshold; a worker-capacity decision needs current backlog and throughput. These questions lead to different instruments and collection choices.

Which instrument matches the measurement?

Use a Counter for non-negative increments, an UpDownCounter for additive changes in both directions, a Gauge for a current value that should not be summed across producers, and a Histogram for a distribution of individual observations. The Metrics API specification also distinguishes recording an event as it happens from observing an existing value during collection. Synchronous and asynchronous here describe how measurements are obtained, not whether your application uses async code.

Instrument choices follow the OTel Metrics API. Observable instruments use collection callbacks; language-specific names and support must be checked against the deployed SDK.
Measurement availableInstrumentCheck before using it
One completed job at a timeCounterAdd 1 per completion; do not repeatedly add a lifetime total
A library's monotonic lifetime byte countObservable CounterReport the absolute reading; verify reset handling downstream
A worker starts or finishes a jobUpDownCounterAdd +1 and -1 at matching lifecycle points, including errors
Current queued jobs read from a queueObservable UpDownCounterReport the absolute value; sum only disjoint queues
Current temperature from a sensorGauge or Observable GaugeRetain sensor identity; adding temperatures is not meaningful
Each completed request's durationHistogramRecord each duration in the declared unit, including relevant failure paths

The distinction between an increment and an absolute reading prevents a common counting error. If a library reports totals of 100 and then 130, calling Counter.add with those values records 230. An observable counter reports 100 and then 130 as the source's cumulative readings. Keep callbacks bounded and avoid slow network requests inside them: a collection timeout can make the metric unavailable when the dependency is already unhealthy.

For active work, use the same attribute set on an UpDownCounter's increment and decrement, and cover errors and cancellations. For a non-additive current reading, choose a synchronous Gauge when code records changes, or an Observable Gauge when collection can read an existing value. Do not replace a latency distribution with an average-only gauge when operators need percentiles or an SLO threshold fraction. Check each instrument with a small, known sequence and include failure paths so its population matches the dashboard.

How should names, units and conventions stay consistent?

Use the applicable semantic conventions for standard HTTP, database and runtime instrumentation. For custom business metrics, use a stable domain-specific name, an explicit unit and a description of the measurement boundary. Put dimensions in attributes rather than generating a metric name for every customer or route. Reuse existing automatic instrumentation where it measures the event you need; adding a second instrument for the same event can create competing totals.

The HTTP server duration convention defines http.server.request.duration in seconds. http.route, when available, is the matched route template, such as /orders/{orderId}. Do not substitute a raw path containing an order ID. If the framework cannot provide the route template, follow the convention's omission rules rather than inventing an unbounded fallback.

Treat a convention change as a query migration. The experimental-to-stable HTTP migration changed http.server.duration in milliseconds to http.server.request.duration in seconds. Renaming alone leaves the values wrong by a factor of 1,000. Where an instrumentation library supports http/dup, old and new representations may coexist: choose one authoritative population for each calculation so the same requests are not counted twice.

Pin instrumentation-library versions separately from the SDK, and record the conventions and backend translation settings in use. Inspect an actual exported point and its stored series name before updating selectors. A known 250 ms operation must remain 0.25 seconds in the new representation. Honeycomb's migration account describes the query gaps that arise during mixed-convention deployments; OpenTelemetry Weaver can complement runtime checks with schema and policy validation.

Which attributes identify the producer?

Separate the entity producing telemetry from the library measuring it and the event being measured. The OTel data model carries resource, instrumentation scope and measurement attributes separately. Mixing those roles makes fleet aggregation and troubleshooting harder.

Model attributes according to their meaning. These locations do not imply that every backend stores or promotes them as the same labels.
Attribute locationWhat belongs thereExample
ResourceIdentity and deployment context of the producerservice.name, service.namespace, service.instance.id, deployment.environment.name
Instrumentation scopeThe library or module that creates the instrumentsScope name and version
MeasurementBounded properties needed to group or filter observationsRoute template, HTTP method, outcome

Configure a meaningful service.name and distinguish instances using an ID that remains consistent for the producing instance. Use service.namespace where names can collide, and record the deployment environment and version when needed to separate releases or environments. Follow the service resource conventions and your platform's resource detection; do not generate a new instance ID for every request.

Keep independent writers distinguishable until an aggregation step combines them correctly. Removing an instance attribute from two cumulative counters can make independent producers appear to write one stream; attribute deletion does not add their values together. The single-writer principle applies even when a dashboard ultimately needs only a service-wide total.

Promote resource attributes to backend labels selectively. Fleet queries need enough identity to remain correct, but copying every deployment attribute onto every series can increase series count and churn. Prometheus's OTLP guide describes direct promotion and metadata joins. Check two instances, two environments and a restart: confirm separate producers at ingestion and the intended service aggregate at query time.

How do Views control cardinality without removing useful answers?

Start with the breakdowns required for decisions. Route and outcome may be necessary for an error budget; a request ID usually belongs in a trace or log rather than a metric attribute. An intentional tenant dimension can be justified by per-tenant SLOs, but it needs an explicit active-set budget. Dropping every expensive dimension can remove the reason the metric exists.

Estimate combinations before choosing a limit. An illustrative metric with 100 routes, five methods and two outcomes has up to 100 × 5 × 2 = 1,000 combinations per producer. Adding 10,000 possible customer IDs raises the full Cartesian upper bound to 10 million. Not every combination will occur. Measure the active set and its growth; replicas, resource changes, histogram representation and retention affect backend cost separately. A per-stream SDK limit is not a fleet-wide series budget.

Use SDK Views to select the measurement attributes and aggregation a stream retains. For a custom Python instrument, View(instrument_name="demo.checkout.duration", attribute_keys={"route", "outcome"}) keeps those two measurement keys. When omitted keys would otherwise distinguish observations, the histogram aggregates them under the retained set. This does not remove resource attributes or guarantee redaction from exemplars. Keep sensitive or unnecessary identifiers out of instrumentation in the first place.

In the pinned Python SDK, matching Views create independent streams. Put attribute filtering and the chosen aggregation on each intended View; separate filter and histogram Views do not form a sequential transformation. Adding a filtered View alongside an existing unrestricted one may retain the expensive stream too. Reusable libraries should expose instruments without taking ownership of the application's provider. Check your SDK's feature support rather than copying configuration between languages.

Use a cardinality limit as a memory safeguard, with explicit overflow monitoring. In SDKs implementing the specified overflow behaviour, measurements beyond the limit can be aggregated into otel.metric.overflow=true, losing their original measurement-attribute breakdown. The official cardinality guide explains why additive totals can remain correct while route or outcome queries undercount. Resource and scope attributes are separate; moving customer IDs into the resource is not a safe workaround.

Exercise the View with two observations that differ only in a removed attribute, then confirm one aggregated point and the correct count or sum. Also test a realistic active set and, where supported, deliberate overflow. An empty overflow query alone does not prove safety: support and backend label translation vary. The useful-labels guide provides a worksheet for deciding which dimensions are worth retaining.

Which histogram resolution does the SLO require?

Choose histogram aggregation according to the questions it must answer. For a fixed SLO such as requests completed in at most 300 ms, an explicit boundary at 0.30 seconds preserves the qualifying count among recorded observations. For exploratory latency analysis and changing thresholds, exponential histograms can cover a wide range without a hand-maintained list of boundaries. Confirm support and acceptable error across the SDK, Collector and backend before choosing either representation.

A coarse histogram cannot always answer a threshold question after collection. Consider these constructed workloads, each containing five requests. Both have the same coarse histogram, count, sum, minimum and maximum, yet different results at the 300 ms threshold.

Synthetic observations, not customer traffic. Coarse bounds are [0.25, 0.50, 1.0]; refined bounds are [0.25, 0.30, 0.50, 1.0]. Arrays include the final bucket above 1.0. Floating-point sums use a tolerance.
Measurement or resultWorkload AWorkload B
Durations, seconds0.10, 0.29, 0.39, 0.41, 0.900.10, 0.30, 0.30, 0.49, 0.90
Count / sum5 / 2.09 s5 / 2.09 s
Minimum / maximum0.10 s / 0.90 s0.10 s / 0.90 s
Coarse bucket counts[1, 3, 1, 0][1, 3, 1, 0]
Counts with 0.30 added[1, 1, 2, 1, 0][1, 2, 1, 1, 0]
True fraction at <= 0.30 s2 / 5 = 40%3 / 5 = 60%

The OTel histogram model uses inclusive upper bounds and per-bucket counts. Prometheus classic le buckets instead contain cumulative counts up to each bound. The equal coarse aggregate fields cannot distinguish these workloads; adding 0.30 preserves the answer by making the first two OTLP buckets count the qualifying requests. Sampled exemplars are not a complete replacement for the lost distribution.

Reproduce the bucket choice with the Python SDK

Use Python 3.10 or later in an isolated environment. SDK 1.45.0 is pinned for reproducibility, not prescribed as the version every production service should use. Save the two Python blocks below, in order, as metrics_contract.py.

sh
python3 -m venv .venv
. .venv/bin/activate
python -m pip install 'opentelemetry-sdk==1.45.0'

python · metrics_contract.py
from math import isclose
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import InMemoryMetricReader
from opentelemetry.sdk.metrics.view import View, ExplicitBucketHistogramAggregation
from opentelemetry.sdk.resources import Resource

def replay(values):
    reader = InMemoryMetricReader()
    layouts = {
        "coarse": (0.25, 0.50, 1.0),
        "threshold": (0.25, 0.30, 0.50, 1.0),
    }
    views = [View(instrument_name="demo.checkout.duration", name=name,
                  aggregation=ExplicitBucketHistogramAggregation(boundaries=bounds))
             for name, bounds in layouts.items()]
    provider = MeterProvider(metric_readers=[reader], views=views,
        resource=Resource.create({"service.name": "metrics-contract"}))
    duration = provider.get_meter("histogram-example").create_histogram(
        "demo.checkout.duration", unit="s")
    for seconds in values:
        duration.record(seconds)
    data = reader.get_metrics_data()
    provider.shutdown()
    metrics = data.resource_metrics[0].scope_metrics[0].metrics
    return {metric.name: metric.data.data_points[0] for metric in metrics}

Each replay creates a private provider and reader. Both Views receive the same observations and deliberately emit different comparison streams. This is useful for the example; retaining both in production would add storage and processing work. Append the assertions, which check the aggregate values as well as the threshold result.

python · metrics_contract.py
cases = {
    "A": ([0.10, 0.29, 0.39, 0.41, 0.90], 2),
    "B": ([0.10, 0.30, 0.30, 0.49, 0.90], 3),
}
for name, (values, expected_good) in cases.items():
    points = replay(values)
    for point in points.values():
        assert (point.count, point.min, point.max) == (5, 0.10, 0.90)
        assert isclose(point.sum, 2.09, rel_tol=1e-12)
    coarse, threshold = points["coarse"], points["threshold"]
    assert tuple(coarse.bucket_counts) == (1, 3, 1, 0)
    assert tuple(threshold.explicit_bounds) == (0.25, 0.30, 0.50, 1.0)
    good = sum(threshold.bucket_counts[:2])
    assert good == expected_good == sum(value <= 0.30 for value in values)
    print(f"{name}: coarse={tuple(coarse.bucket_counts)}; "
          f"<=300 ms={good / threshold.count:.0%}")

Run the combined file. It prints the common coarse bucket counts and then 40% for A and 60% for B. The result demonstrates why checking only the total count and sum is insufficient when choosing bucket boundaries.

sh
python metrics_contract.py

Keep histogram layouts compatible across the population a query aggregates. If only upgraded replicas expose a new Prometheus classic le="0.3" bucket, its numerator covers less traffic than a fleet-wide count denominator. Prometheus's histogram guidance documents this incomplete-result risk. Scope a new calculation to a complete cohort during the transition; inspect the actual translated bucket labels before enabling its alert.

More buckets cost memory and payload, and fixed boundaries still do not make arbitrary percentiles exact. Prometheus recommends native histograms where possible, but OTel exponential histograms and a backend's native representation require a compatible translation path. Test the question at the retained resolution. Never average per-instance p99 values to obtain a fleet p99: aggregate compatible histogram distributions first, then calculate the fleet quantile.

How should collection, export and temporality be configured?

Configure the application's provider, resource and readers at startup, then reuse instruments. If automatic instrumentation already owns that setup, extend it rather than attempting a second global installation. Record synchronous measurements on their event paths; collection and export happen separately. With a periodic MetricReader, choose an export interval based on alert freshness, payload and CPU budget. Exporting less often does not automatically sample fewer Counter or Histogram recordings, but observable callbacks run when collection occurs and can miss changes between observations.

Allow bounded flush and shutdown work during application termination, using the SDK's documented APIs. A reader with a long interval can otherwise miss a short-lived process's final export. Termination hooks cannot protect against abrupt process death, and export success alone does not prove that the final query is complete.

Architecture choices, not a requirement to deploy every tier. Validate the selected exporters and processors in the Collector distribution and version you run.
Collection pathUse it whenOperational cost
SDK directly to a backendThe destination supports the SDK's representation and configuration needs are simpleEach application manages remote endpoint, credentials and export behaviour
SDK to a local CollectorLocal resource enrichment or a standard host-level collection path is usefulAn agent or sidecar needs resources, updates and failure monitoring
Agent or SDK to a gateway CollectorCentral processing, destination policy or shared export controls are requiredGateways need capacity planning; stateful processing also needs correct stream routing

Start with direct export when it meets the requirements. Add Collector tiers for specific work: local collection or enrichment, central policy, or independently managed export capacity. Each tier adds a dependency to secure, size and monitor. Budget total freshness through collection, batch waiting, export, backend ingestion and alert evaluation; shortening only the SDK interval may leave the dominant delay elsewhere.

Choose temporality with the destination for sums and histograms; OTLP Gauges do not have cumulative or delta temporality. Cumulative points cover measurements since a start time; delta points cover successive intervals. Do not add successive cumulative totals as if they were deltas. Nor should a delta stream be interpreted as a cumulative counter without conversion. Verify the exported representation and actual receiver support.

Cumulative output can retain totals across a missed export if a later point from the uninterrupted producer arrives, although it cannot recover the timing within that gap or an unexported tail lost when the process dies. Delta output can reduce long-lived attribute-set state in applications, but losing an interval leaves a gap the next interval does not fill. Choose where that state and recovery responsibility should live; neither representation is universally cheaper or safer.

Avoid unnecessary conversions. Where delta-to-cumulative conversion is needed, consecutive points for a stream must reach the component holding its accumulated state. Check routing, state expiry and restart behaviour. Current Prometheus OTLP documentation describes experimental delta-to-cumulative ingestion with a feature flag; this does not establish support in every Prometheus-compatible backend or remote-write path.

Use the receiver's documented OTLP protocol, endpoint and authentication requirements. gRPC and HTTP/protobuf settings are not interchangeable, and endpoint path handling depends on whether configuration is generic or signal specific. Use TLS for remote transport, protect receiver access and keep credentials out of examples and source control. Check against the OTLP exporter specification and the implementation actually deployed.

Configure sending queues and retries deliberately on network exporters that support them. Size for expected arrival rate, recoverable outage duration and resource limits; record whether capacity is measured in requests, items or bytes. An in-memory queue does not survive a Collector crash. Persistent storage can preserve queued data across restarts, but disk failure, full queues and retry expiry still create loss boundaries. It does not persist upstream batch buffers or converter state. The Collector resilience guide and exporter-helper configuration explain these trade-offs.

Place the memory limiter first among processors so pressure can be returned towards receivers. Leave headroom below the container limit; periodic checks are not a guarantee against every out-of-memory failure. If using the batch processor, place it after the limiter and intentional filtering. Batching reduces request overhead but adds memory and waiting time. Account for exporter queue batching too, rather than adding buffers without measuring their combined delay.

Test a receiver outage and recovery under continuing load. Confirm that backlog drains, failures are observable and no unintended writer collisions appear. Persistent-queue recovery and converter recovery need separate checks because they preserve different state. Use the remote-write guide for the separate Prometheus sender's queue and replay behaviour.

How do you keep metrics useful in production?

Keep a small acceptance workload with every metric used for paging, SLOs or scaling. Record its expected population, window and numerical tolerance before changing the pipeline. Compare the generated workload with emitted points, the backend query and the alert state. Agreement between old and new pipelines is useful but insufficient if both lose the distinction an operator needs.

Recommended deployment checks. The local histogram example exercises one aggregation property; these wider operational scenarios must be run against the actual deployment.
Change or failure to exerciseEvidence to retainDo not accept
Instrument or unit changeKnown increments, absolute readings and physical durationsDouble-counted totals or milliseconds interpreted as seconds
New attribute or ViewRequired breakdowns, active combinations and overflow stateA cheaper metric that no longer supports the decision
Two producers and a restartDistinct writer identities, start timestamps and fleet resultA reset presented as demand or colliding streams
Histogram or convention migrationOld-only, new-only and mixed-cohort query resultsA numerator representing fewer requests than its denominator
Receiver or converter outageQueue behaviour, recovery, state loss and missing intervalsAn HTTP success or fresh sender timestamp treated as complete delivery
Alert and missing-data behaviourKnown healthy, failing and absent intervalsMissing telemetry treated as proof of a healthy service

Monitor the telemetry path itself. The Collector's internal telemetry catalogue includes otelcol_exporter_queue_size and otelcol_exporter_queue_capacity; compare matching exporter instances and check the configured sizing units. Also track rejected metric points and otelcol_exporter_send_failed_metric_points, which is not automatically a permanent-loss total because retries can later succeed. Metric availability, exposed names and suffixes vary by component and version, so inspect the deployed output before writing selectors.

Watch Collector process resource use and backend ingestion health alongside those counters. Do not treat accepted points minus sent points as a general loss formula: filtering, fan-out and buffering can make the counts incomparable. Keep an independent availability check where a failing pipeline could also hide its own telemetry, and define an explicit alert outcome for missing data.

Give high-impact metrics an owner and review their consumers when instrumentation or traffic changes. Retire unused streams only after checking dashboards, recording rules, alerts and external users. Roll out convention and histogram changes to a cohort first, preserve old queries until migration is complete, and keep a rollback plan for both instrumentation and alert definitions. Stable Prometheus instrumentation does not need to migrate merely because the calendar says 2027.

When comparing xScaler with another platform, hold required dimensions, histogram resolution, retention, collection interval and query behaviour constant. Measure the cost of that workload. Removing a tenant breakdown or coarsening latency buckets can reduce telemetry volume, but it changes the comparison if operators still need those answers.

What are the main OpenTelemetry metrics best practices?
Match instruments to measurement semantics; follow names and units from applicable conventions; preserve resource identity; retain bounded, useful attributes through Views; choose histogram resolution for the query; configure compatible collection and temporality; and test the resulting dashboards and alerts with known workloads.
Should every measurement use an asynchronous instrument?
No. Record event increments and durations with synchronous instruments when the event happens. Use observable instruments to read an existing total or current state during collection. An observable counter reports an absolute monotonic value; a synchronous counter records increments.
Does lowering the export frequency solve high cardinality?
No. Export frequency is not an attribute budget. Its effect on aggregation state depends on temporality and SDK behaviour, while the backend still sees distinct exported identities over time. Bound dimensions deliberately, use supported Views and limits, and monitor overflow and fleet series growth.
Should every application send metrics through a Collector?
No. Direct export can be simpler when the SDK and destination meet the requirements. Add local or gateway Collectors for specific collection, processing or operational needs, and account for the extra failure domains and any stateful routing.
What makes these guidelines relevant to 2027?
They address ongoing instrumentation and operating decisions using sources checked on 2 October 2026. The year is a planning horizon, not a prediction of new releases. Recheck SDK, convention, Collector and backend compatibility before a 2027 rollout.