Migrating from Datadog without downtime
Migrate from Datadog with evidence for operational decisions: verify telemetry, query meaning, paging handover and rollback before retiring the old path.

Migrate from Datadog one operational decision at a time. Keep the incumbent collection and paging paths working while you verify the replacement's data, queries and incident workflow. Transfer paging explicitly, rehearse rollback with an active alert, then retire dependencies separately. Parallel ingestion alone cannot establish equivalent monitoring or guarantee application uptime.
What does a migration without downtime require?
An application can keep serving requests while its monitoring stops detecting failures. A migration therefore needs separate acceptance criteria for application availability and observability coverage. Watch the application's latency, errors and resource use during collector or instrumentation changes, and independently test whether responders can detect and investigate a failure.
Existing guides already recommend parallel operation. SignalCost covers shadow operation and rollback, while Chronosphere prioritises critical use cases and monitor backtesting. The method here makes the acceptance evidence explicit: each operational decision gets a record of its dependencies, observed behaviour, current paging authority and reversal procedure.
Start with an inventory of what Datadog does for the team: metrics, log searches, traces, monitors, SLOs, scheduled jobs, cloud integrations and notification routes. Record the producers, SDKs, agents, processing rules, retention and access requirements. Include RUM, security, profiling or other workflows if you use them. A working Prometheus query does not replace those dependencies.
Choose the replacement scope before deploying it. This guide uses Prometheus-compatible metrics and a Grafana investigation workflow as examples; name the actual rule engine and notification service in your plan. Prometheus-managed rules and Grafana-managed rules do not have identical no-data behaviour. Obtain verified ingestion endpoints, authentication details and supported formats from your destination, rather than deriving them from a vendor's name.
What evidence belongs in the cutover record?
Assign an accountable owner to each critical decision. For a checkout incident, that might mean detecting a service-wide error increase, finding the affected deployment and opening a related trace or log record. Preserve the old monitor and dashboard identifiers so discrepancies remain traceable. Save target queries and configuration versions beside the evidence. Agree acceptable numerical differences and detection delays before comparing results; there is no useful universal percentage tolerance.
| Record field | Illustrative checkout decision | Evidence needed |
|---|---|---|
| Meaning and scope | Completed server responses; HTTP 5xx / all responses; one service and environment | Producer, units, filters, grouping and window documented |
| Decision behaviour | Service-wide ratio above an agreed threshold; separate coverage alarm | Normal, breach, recovery, no-traffic and missing-source cases |
| Investigation | Find the deployment, request log and retained trace | Recorded identifiers and working links under responder permissions |
| Paging authority | Datadog until an approved handover; target writes shadow evidence | Threshold and telemetry-error notifications reach the intended destination |
| Reversal and retirement | Restore the old route while old collection remains live | Active-alert rollback receipt; remaining dependencies signed off |
Use additional records for decisions with different failure modes. A missed overnight job needs a deadline and explicit absence detection. A latency investigation needs a trace search that finds the relevant service and preserves useful context. Neither is covered by passing the checkout error-rate test. If a target deliberately behaves differently, record the reason and obtain the owner's approval instead of calling the difference parity.
Separate four approvals: telemetry is available; the operational decision is reliable; the target may page production; the old dependency may be removed. Each approval needs its own reversal procedure. Reverting a notification route should not require changing application instrumentation, and retaining an old dashboard should not be mistaken for retaining live detection.
How should you establish parallel collection?
Choose a small service cohort with a known owner and a reversible deployment. Keep the current Datadog path intact initially. Separate changes to instrumentation from changes to transport and storage where the existing setup permits it; otherwise a disagreement has several possible causes at once.
| Existing producer | Migration approach | Check before expanding |
|---|---|---|
| OpenTelemetry | Evaluate supported Collector fan-out to both destinations | Per-destination identity, transformations, sampling and failure behaviour |
| Prometheus exposition | Scrape with a scoped collector and use the destination's documented ingestion path | Scrape load, duplicate collection, labels and counter semantics |
| Datadog-specific instrumentation or integrations | Retain the incumbent path; assess a supported receiver or staged replacement | Exact payload support and every dependent product workflow |
| Logs and traces | Verify the source, processing and destination path independently | Redaction, retention, sampling, timestamps and correlation fields |
The OpenTelemetry Collector architecture supports fan-out, but sharing a process or receiver also shares failure risks. The documentation warns that a blocking processor on one pipeline can block other pipelines attached to the same receiver. Test destination unavailability on an isolated cohort and observe the incumbent path as well as the new one.
Queues and retries need their own acceptance checks. Collector resiliency guidance documents loss through queue overflow, retry expiry, crashes without persistence and storage failures. Record queue capacity and units, retention or retry limits, disk requirements and how backlog drains after recovery. For a Prometheus sender, use the remote-write operational guide to assess lag and recovery headroom.
Datadog supports OpenTelemetry, so adopting it does not require immediately leaving Datadog. Check the current compatibility matrix for the exact route. It warns against mixed Datadog and OpenTelemetry SDK instrumentation in most cases, with language-specific exceptions, and against a separate Agent and Collector on the same host. Do not treat installing another agent as a universally safe dual-write recipe.
For logs, compare a controlled set of event identifiers, timestamps, parsed fields and redaction results. For traces, follow a known cross-service request and check parent-child relationships, errors and log correlation. If sampling differs, equal span counts are not a sensible acceptance target. Record where request metrics are generated relative to sampling; arrival of some traces does not establish a complete request denominator. Datadog's Collector setup guide explicitly separates span-metric generation from trace sampling.
How do you preserve query and alert meaning?
For each important dashboard and monitor, write down the measurement before translating its syntax: producer, unit, population, grouping, time window, rollup, interpolation and missing-data behaviour. Datadog's metric type modifiers change count/rate representation; .as_rate() is not a text substitution for PromQL rate() on cumulative counters. OpenTelemetry metrics guidance covers instrument and identity choices.
Datadog's as_count() monitor documentation shows that evaluation order can change the answer: its count path aggregates over time before division, while the classic path performs arithmetic before time aggregation. Preserve the intended operation, not just the threshold. Datadog's rollup guidance also warns that rollup buckets and monitor evaluation windows need not align.
Two correct calculations, different paging decisions
Consider a synthetic service with two instances. In every one-minute interval, instance A completes 10 requests, including one HTTP 500; instance B completes 990 requests with no failures. Assume every request appears once, all counters are collected and both instances belong to the same service and environment. This example concerns aggregation across instances, distinct from Datadog's time-aggregation example.
| Population | Requests / minute | 5xx / minute | Error fraction |
|---|---|---|---|
| Instance A | 10 | 1 | 10% |
| Instance B | 990 | 0 | 0% |
| All requests | 1,000 | 1 | 0.1% |
At an illustrative 1% service-wide threshold, an equal-weighted average would trigger while the request-weighted fraction would not. The 10% on instance A may still justify a separate instance-level alert. Keep that decision explicit instead of accidentally using an average as a substitute for either policy.
Save this standard-library Python example as migration_ratio.py and run python3 migration_ratio.py. It prints 5.0% for equal instance weights, 0.1% for all requests and 0.0% after dropping A from the observed population. The last result represents lost coverage, not recovery.
from fractions import Fraction
# Synthetic requests and HTTP 5xx responses per minute.
instances = {
"A": (10, 1),
"B": (990, 0),
}
mean = sum(Fraction(e, n) for n, e in instances.values()) / 2
weighted = Fraction(
sum(e for n, e in instances.values()),
sum(n for n, e in instances.values()),
)
assert mean == Fraction(5, 100)
assert weighted == Fraction(1, 1000)
remaining = Fraction(instances["B"][1], instances["B"][0])
print(f"Equal instance weights: {float(mean):.1%}")
print(f"All requests: {float(weighted):.1%}")
print(f"A missing: {float(remaining):.1%}")Test the rule against missing telemetry too
For a local rehearsal, the following rule expects the synthetic cumulative counter migration_http_requests_total, labelled by service, instance and status. Use only the isolated test environment. It assumes a zero-valued 5xx series exists even on the healthy instance and that the service selector identifies a single environment. Replace the metric contract and grouping deliberately before adapting it to real traffic.
groups:
- name: migration-rehearsal
interval: 1m
rules:
- record: migration:checkout_error_ratio:rate5m
expr: |
sum by (service) (
rate(migration_http_requests_total{service="checkout",status=~"5.."}[5m])
)
/
sum by (service) (
rate(migration_http_requests_total{service="checkout"}[5m])
)
- alert: CheckoutErrorRate
expr: migration:checkout_error_ratio:rate5m > 0.01
for: 2m
labels:
severity: pageValidate the file with promtool check rules migration-rules.yml. Then save the fixture below alongside it as migration-test.yml and run promtool test rules migration-test.yml. The counter histories add 1 and 9 per minute to A's 500 and 200 series, and 0 and 990 to B's. The assertion checks the request-weighted fraction and absence of a service-wide alert at ten minutes. See the promtool test format for extending the rehearsal.
rule_files: [migration-rules.yml]
evaluation_interval: 1m
tests:
- interval: 1m
input_series:
- series: 'migration_http_requests_total{service="checkout",instance="a",status="500"}'
values: '0+1x10'
- series: 'migration_http_requests_total{service="checkout",instance="a",status="200"}'
values: '0+9x10'
- series: 'migration_http_requests_total{service="checkout",instance="b",status="500"}'
values: '0+0x10'
- series: 'migration_http_requests_total{service="checkout",instance="b",status="200"}'
values: '0+990x10'
promql_expr_test:
- expr: migration:checkout_error_ratio:rate5m
eval_time: 10m
exp_samples:
- labels: 'migration:checkout_error_ratio:rate5m{service="checkout"}'
value: 0.001
alert_rule_test:
- eval_time: 10m
alertname: CheckoutErrorRate
exp_alerts: []The baseline should report SUCCESS. For the partial-loss case, remove A's two input series and change the expected fraction to 0; keep the expected alert list empty. The service rule now passes while A is unobserved, which demonstrates why it needs an independent coverage check. The rule applies rate() before summing, as Prometheus recommends, to preserve counter-reset detection.
| Synthetic case | Observed or expected result | Acceptance consequence |
|---|---|---|
| Both instances present | Observed: weighted 0.1%; equal instance weights 5% | Confirm which population the operational policy concerns |
| A absent for a full query window; B still collected | Observed: weighted ratio becomes 0% | Reject coverage even though the error rule is quiet |
| No series in the query window | Observed: empty vector, not a numeric zero | Require explicit missing-data detection |
| Both instances scraped but counters stay constant | Rule test: 0 / 0 is not a usable error fraction; no threshold alert | Distinguish genuine no traffic from collection failure |
| B now has 20 errors and 970 successes per minute | Rule test: 2.1%; pending before firing after the configured hold | Check threshold timing separately from notification delivery |
Keep coverage checks independent of the error fraction. Compare expected producers from a deployment inventory against observed producers, with an explicit allowance for startup and scale-down. A service-wide absent_over_time() cannot detect A disappearing while B remains. A failed scrape can be detected with up, but a target removed from discovery needs an independent expectation that it should still exist. Avoid converting all absent results to zero: that would erase the distinction the rehearsal is meant to expose.
Timing also needs translation. In Prometheus alerting rules, for: 2m requires the expression to remain active at evaluations for that duration. It is not a two-minute data window, and it does not by itself specify notification delivery time. Preserve the original Datadog monitor's evaluation delay, new-group handling and recovery policy where required, or record the accepted change. Check downstream notification grouping and timing separately.
What should the shadow rehearsal prove?
Choose the observation period around the service's operating cycle and the decisions being migrated. Include a deployment, expected low-traffic periods and scheduled workloads where relevant. Quiet dashboards for a fixed number of days cannot establish how a rare failure will behave. Use isolated or approved fault injection for cases ordinary traffic will not exercise.
- Exercise a threshold breach and recovery. Record the source event, query result, pending/firing/resolved times and affected group in both systems.
- Stop one expected telemetry producer, then test complete absence and a query error separately. Confirm which coverage or pipeline alarm is responsible.
- Restart a producer and introduce a new instance. Check counter resets, identity, discovery delay and whether the new group receives the intended treatment.
- Make the target destination unavailable in the isolated cohort. Check backlog, drops, recovery and whether the incumbent collection path remains usable.
- Follow a known request through the target dashboard, log search and retained trace using ordinary responder access. Record any loss of investigative context.
- Send controlled notifications through the shadow route. Verify receipt, grouping, escalation and recovery messages without accidentally paging the production responder.
Record every disagreement with its population, time range and configuration revision. Classify it as a collection gap, transformation difference, query difference, rule-state difference or routing difference, then assign an owner. Compare against independently known events where possible: two systems agreeing is weak evidence if both receive the same incomplete input.
For Grafana-managed rules, inspect No Data and Error handling. DatasourceNoData and DatasourceError are separate alerts with different labels; policies or silences matching the original alert may not match them. Prove that these notifications also stay on the intended shadow route. Apply the documented behaviour of your deployed version rather than assuming this applies to Prometheus-managed rules.
How do you transfer paging and roll back?
Keep one designated production paging authority for each decision during the shadow period. Both systems can evaluate; the target's notifications go to a controlled test destination. Record the scope and owner of any temporary overlap during handover so duplicate pages are recognised rather than treated as unrelated incidents.
- Before handover, confirm both collection paths are fresh, the target's incident cases passed and the on-call team can use its runbooks. Record the exact monitors and groups being transferred.
- Enable the target's production route for that scope under an attended change. Prove notification receipt, then suppress the corresponding incumbent route. Accept a controlled overlap rather than an unobserved gap; reconcile any duplicate incidents.
- Keep the old ingestion and evaluation paths live for the agreed fallback period. Record who may reverse the handover and what evidence triggers reversal.
- Rehearse reversal while a condition is already active. Restore incumbent delivery, verify the responder receives the current condition and reconcile open incidents before suppressing target delivery.
This distinction comes from Datadog's downtime documentation, where Downtime means notification suppression, not application unavailability. A saved rollback configuration is insufficient if it restores a green integration indicator but delivers no active alert. Datadog's own on-call migration guidance also calls for testing receipt and escalation; the same discipline is useful when moving away from it.
Define concrete rollback triggers before the change: a missed critical condition, unresolved loss of an expected producer, broken investigation access or failed notification delivery. Keep the rollback procedure separate for instrumentation, transport and paging. If both paths share the failed collector, changing the paging route alone will not restore the missing input.
When can you retire Datadog, or choose to stay?
Retire only the dependencies whose owners have accepted the target behaviour and fallback arrangements. Check for agents still collecting another signal, SDK features used elsewhere, cloud integrations, scheduled reports, SLO windows, historical investigations and incident links that still depend on Datadog. Revoke an integration credential only after its remaining consumers have been identified and removed or migrated.
Treat historical telemetry separately from new ingestion. State the earliest available target timestamp and test older incident and SLO queries before removing access to the old system. Confirm export, retention and access options for the data you actually need. Copying dashboard JSON does not transfer its historical data, and stopping ingestion does not settle subscription commitments.
Budget for the overlap using your measured workload: old and new ingestion, retained storage, collector resources, network transfer, query load and engineering time. Compare the recurring cost only after matching retention, useful dimensions, investigation needs and support. Use the Datadog alternatives guide to assess the destination; a claimed percentage saving cannot replace that workload comparison.
xScaler is a cost-effective observability platform covering collection, storage, exploration and notification across applications, infrastructure and AI, with telemetry collector fleet management. Open standards and BYO Grafana support a staged move. The acceptance record still needs to establish compatibility for your workloads; this guide promises neither an automatic converter nor a particular saving.
Keep Datadog for a workload when its required RUM, security, profiling or other product workflow lacks an accepted replacement, or when the team cannot safely operate the overlap. Standardising instrumentation while retaining Datadog can be a useful intermediate outcome. A narrower migration with working incident coverage is preferable to removing the last dependable detection path to meet a renewal date.
- How long should both platforms run?
- Long enough to cover the operating cycles and failure cases for the decisions being moved. Set a review date and cost budget, but make approval depend on evidence such as deployment, scheduled-job, low-traffic and incident rehearsals. There is no universal two-week acceptance period.
- Can Datadog dashboard exports be imported directly into Grafana?
- Do not assume format conversion preserves meaning. Check the destination tool's supported objects, then validate queries, variables, units, grouping and incident links. An exported dashboard is not an export of the telemetry behind it.
- Do I need to replace every Datadog SDK first?
- No blanket ordering is safe. Inventory the features it supplies and choose a supported route. Datadog supports OpenTelemetry, but its compatibility matrix makes feature support depend on instrumentation and collection choices. Avoid unsupported mixed SDK instrumentation.
- Does a passing Prometheus rehearsal prove Datadog parity?
- No. The synthetic example proves specific aggregation and missing-input behaviours in Prometheus. Your Datadog source queries, collector transformations, rule states and actual notification delivery need separate comparison against known events.
- Can I stop paying for Datadog as soon as the target pages?
- Technical handover and commercial termination are separate decisions. Check remaining telemetry and product dependencies, historical access, the agreed rollback period and your contract before removing agents, revoking credentials or changing the subscription.
Technical sources and complementary migration coverage, checked 2 October 2026. Vendor recommendations and operator accounts are attributed, not treated as proof of an xScaler migration. Living documentation must be checked against deployed versions.
- Datadog: OpenTelemetry feature compatibility 2 Oct 2026
- Datadog: OpenTelemetry Collector setup and span-metric generation 2 Oct 2026
- OpenTelemetry: Collector architecture and fan-out 2 Oct 2026
- OpenTelemetry: Collector resilience and loss boundaries 2 Oct 2026
- Datadog: as_count() in monitor evaluations 2 Oct 2026
- Datadog: rollups and monitor windows 2 Oct 2026
- Prometheus: rate(), absent() and absent_over_time() 2 Oct 2026
- Prometheus: alerting rules and pending state 2 Oct 2026
- Prometheus: rule and expression testing 2 Oct 2026
- Prometheus 3.5.0: pinned example release 2 Oct 2026
- Grafana: No Data and Error alert routing 2 Oct 2026
- Datadog: metric type modifiers 2 Oct 2026
- Datadog: monitor timing, recovery and grouping 2 Oct 2026
- Datadog: downtime cancellation and expiry notifications 2 Oct 2026
- Datadog: on-call migration validation 2 Oct 2026
- SignalCost: low-risk observability migration 2 Oct 2026
- Chronosphere: Datadog migration considerations 2 Oct 2026