Chapter 21: Observability and Operations
How do I know what's happening in production and respond to issues?
Production Workers change how you investigate failures: there is no host to SSH into, persistent local log directory to inspect or process to attach an interactive debugger to. Invocations may run in different locations and isolates may disappear. Capture the evidence needed for diagnosis through logs, traces and durable application records.
An incident is harder to diagnose when the useful evidence lived only in an isolate that has disappeared. Decide what must survive the request before choosing where to send it.
Design the evidence before the incident
Workers removes access to the host you might otherwise inspect during an incident. Isolates can serve several requests and be reused, but you cannot depend on their memory surviving or attach an interactive debugger in production. Workers Logs and tracing provide managed evidence; your application still has to supply the context that explains its decisions.
Start by identifying the questions an incident will raise. Which deployment handled the request? Which tenant and operation were affected? Did time go into CPU work, a database call or an external dependency? Capture stable identifiers and safe context before the incident, and retain them for the period your response process needs.
A request can cross several Workers and storage services. Use the platform's tracing where it follows those calls, then add correlation across the boundaries it does not cover. The aim is to reconstruct a failing operation without relying on access to the machine that ran it.
What you can and cannot observe
Honesty about limitations prevents wasted debugging time. Some capabilities you take for granted elsewhere don't exist on the edge.
You cannot attach a debugger to a production Worker. Isolates are reused across requests but can be evicted without notice, and the platform does not expose them for interactive production debugging.
Production Workers do not expose the same CPU profiling workflow as local development. Reproduce expensive execution paths locally and profile them with DevTools; use production metrics and traces to decide which paths need investigation.
You cannot inspect memory state after a request completes. No heap dump, no core dump, no memory snapshot. If you needed to see what was in memory, you needed to log it during the request. What you can see is how much memory isolates use: the dashboard's memory usage chart reports per-invocation usage in P50 to P999 percentiles with deployment markers, enough to catch a leak or answer right-sizing questions without revealing what the heap contains.
You cannot access the underlying system. No system metrics, no network statistics, no disk I/O measurements. The abstraction that enables global deployment hides infrastructure details completely.
Local reproduction is useful when production telemetry cannot explain a failure. Capture the deployment, operation, relevant identifiers and a redacted description of the input. Record full bodies or headers only when their diagnostic value justifies the data exposure and retention, and exclude credentials and sensitive fields.
Workers Logs: the default starting point
Start with Workers Logs for managed retention and queries. Add live inspection or exports when a specific investigation, retention or alerting requirement calls for them.
Workers Logs ingests and indexes logs for requests selected by its sampling configuration. Retention is three days on Free and seven days on Paid. Enable it in configuration if it is not already enabled for the Worker:
{
"observability": {
"enabled": true,
"logs": {
"invocation_logs": true,
"head_sampling_rate": 1
}
}
}
Workers Logs makes console output, invocation metadata and errors queryable without operating a separate logging service. Their availability still depends on sampling, retention and platform limits.
On Paid, twenty million log events per month are included, with additional events charged at $0.60 per million. Ten million requests producing two events each fit that allowance. Count invocation events and application logs together when forecasting volume.
Head-based sampling for cost control
High-traffic Workers can exceed their included logging allowance. Head sampling retains a proportion of requests and their associated logs:
{
"observability": {
"enabled": true,
"logs": {
"invocation_logs": true,
"head_sampling_rate": 0.1 // Log 10% of requests
}
}
}
At a rate of 0.1, approximately one request in ten is selected. Assuming two events per request, a billion requests produce about 200 million retained events. After the paid allowance, that is approximately $108 at $0.60 per million. Sampling reduces investigation detail as well as cost; a rare failure may not be retained.
Head sampling selects requests before their outcomes are known, and excludes all logs from requests it does not select. Logging errors unconditionally in application code cannot override that decision. If the application must retain every error log, keep the head sampling rate at one and reduce successful-request logging in code, or choose an export pipeline with suitable filtering and delivery guarantees.
The query builder
Raw log storage is useful, but queryable log storage is powerful. The Query Builder provides a SQL-like interface for investigating your logs without external tooling.
The Query Builder operates through the Cloudflare dashboard at Workers & Pages → Observability → Investigate. Select fields to display, filters to apply, groupings for aggregation, and time ranges to search. The interface translates your selections into queries against the Workers Observability dataset.
Typical investigations:
Highest error rate endpoints? Filter to 5xx status codes, group by request path, visualise as bar chart.
Latency distribution for a specific Worker? Select wall time, filter to Worker name, visualise as histogram.
What happened to a specific user? Filter by user ID (assuming you log it), sort by timestamp, display as event list.
The Query Builder also exposes a REST API for programmatic access: listing available keys, running queries, listing unique values for a specific key. Build custom dashboards, integrate with existing monitoring systems, or automate investigation workflows.
When Workers Logs isn't enough
Workers Logs solves the common case: operational visibility with minimal configuration and reasonable cost. Some requirements exceed what it provides.
Retention beyond seven days requires exporting to another system. For 90-day compliance retention or historical analysis spanning months, Logpush to R2 remains appropriate. Workers Logs provides the operational window; R2 provides the archive.
Alerting on log content needs an additional processing path. Workers Logs stores and indexes but does not watch and notify. Use a log platform with alerting or a Tail Worker: application code that receives another Worker’s execution logs and exceptions after an invocation and can evaluate alert conditions.
Cross-system correlation with non-Cloudflare logs requires a unified platform. If your architecture includes Lambda functions, Kubernetes pods, and Workers, correlating traces across all three requires exporting to a system that ingests from all sources.
The decision framework: Workers Logs handles most needs by default. Do you have specific requirements (long retention, real-time alerting, cross-system correlation) that justify additional infrastructure?
For most applications, Workers Logs is the answer. Enable it, query it when needed, invest your engineering time elsewhere.
Live inspection and persistent exports
Workers Logs provides the retained investigation window. Live streams help during an active investigation; exports support longer retention and external tooling. Choose the additional path from the evidence you need to inspect or preserve.
Real-time streams
Use wrangler tail to inspect console output, exceptions and request metadata during a deployment or active investigation. It complements retained Workers Logs: a live stream is useful for watching what happens next, while stored logs explain what happened before you connected.
Filter to the affected Worker, operation or request where possible. At high volume, sampling keeps a stream manageable but can omit the particular failure under investigation. A running terminal is not a retention strategy.
Persistent exports with Logpush
Workers Logs handles most operational logging needs with zero configuration beyond enablement. Logpush remains valuable for retention beyond seven days, integration with existing log platforms, or log data in specific formats for compliance.
Choose a log destination by the investigations it must support, the retention period and the cost of querying it. Cheap object storage and an indexed log platform provide different services.
At ten million requests daily and one kilobyte per log, the raw volume is about 300 GB each month. Object storage can hold that cheaply, but interactive search also needs ingestion, indexing or a query engine, and operational ownership. Compare the complete investigation path rather than a storage bill with an indexed-platform bill.
Choosing a Log Destination
| If your situation is... | Choose... | Because... |
|---|---|---|
| Already paying for Splunk/Datadog with budget headroom | Log platform | Query capability worth the cost |
| Long retention with infrequent queries | R2 + query infrastructure | Model storage, query and operating cost together |
| Need real-time alerting on log content | Log platform | R2 requires building alerting yourself |
| Compliance requires long retention | R2 for archive, platform for recent window | Balance cost with query capability |
| Small scale, exploring | Workers Logs | Queryable without building an archive pipeline |
R2 is just storage; you need to build or buy query infrastructure. This works well if you already have a data pipeline (Spark, Athena, BigQuery) that can query files in object storage. It's a poor choice for querying logs interactively during an incident.
A middle path keeps a recent queryable window in a log platform and an archive in R2. Choose and monitor the export delivery path as well as the retention period; storing an archive does not by itself prove that every expected event arrived.
For SQL queries over archived logs, Pipelines can write records into Iceberg tables in R2 for R2 SQL to query. This avoids operating a separate query cluster, while adding ingestion, table layout and query-cost decisions. Chapter 14 covers that pipeline. Choose it for the retention and investigation workload rather than assuming every archive needs an analytics stack.
For the zone's http_requests dataset, Logpush can merge eligible subrequests into a Subrequests array on the parent record. That simplifies investigations of a Worker that fans out to several backends, but the record can be incomplete: at most 50 subrequests completing within five minutes qualify, and others remain separate. Preserve RayID and ParentRayID correlation in downstream queries, and confirm the feature is enabled for the zone.
The container logs view can correlate related Worker and Durable Object logs. Use that shared context when investigating a request that crosses these boundaries, rather than treating the container as the only possible source of delay.
Analytics Engine: high-cardinality custom metrics
The dashboard's aggregated metrics and Workers Logs answer operational questions: error rates, latency distributions, request volumes. Analytics Engine answers business questions: revenue by customer, usage by feature, performance by geographic region. Dimensions that would explode traditional time-series databases.
Traditional metrics systems like Prometheus or InfluxDB struggle with high-cardinality labels. A metric labelled by user ID across millions of users creates millions of time series, each consuming memory and storage. The database slows, queries timeout, and you're forced to aggregate away the detail you wanted.
Analytics Engine handles exactly this pattern. Write millions of distinct data points with arbitrary cardinality, then query them with SQL. Cloudflare uses Analytics Engine internally to power the per-product metrics in the dashboard for D1, R2, and other services. The same infrastructure handles your custom metrics.
Writing data points
Configure an Analytics Engine dataset in your Wrangler configuration and write data points from your Worker:
{
"analytics_engine_datasets": [
{
"binding": "METRICS",
"dataset": "application_metrics"
}
]
}
// Record a user action with rich dimensions
env.METRICS.writeDataPoint({
blobs: [userId, action, region, planType], // String dimensions
doubles: [latencyMs, requestSize], // Numeric values
indexes: [tenantId] // Sampling key
});
The blobs array holds string dimensions: user IDs, action names, countries, feature flags. The doubles array holds numeric values: latencies, counts, sizes, amounts. The indexes field determines sampling behaviour at extremely high volume.
Writes do not await delivery; the platform handles batching and delivery asynchronously. Invalid API inputs can still fail synchronously. Use these points for telemetry, not as a durable ledger of financial or business events.
Querying with SQL
The data point above records latency in double1 and plan type in blob4. A query should preserve that schema and account for sampling:
SELECT
blob4 AS plan_type,
SUM(_sample_interval) AS estimated_actions,
SUM(double1 * _sample_interval) / SUM(_sample_interval) AS mean_latency_ms
FROM application_metrics
WHERE timestamp > NOW() - INTERVAL '7' DAY
GROUP BY plan_type
Analytics Engine may sample during ingestion and querying. Each returned row's _sample_interval is its weight: use it inside aggregations. Indexing by tenant helps separate sampling decisions for high-volume tenants; it does not mean all of a tenant's points are included or excluded together.
Use cases that fit
Analytics Engine excels at specific patterns:
Usage-based billing: aggregate usage per customer with the sampling weights applied. Match the estimate and retention to the billing contract; use an event ledger when charges need exact event-level evidence. Chapter 25 covers that distinction.
Feature analytics: which features customers use, how frequently, with what parameters. High cardinality is the point.
Performance monitoring by customer: which customers experience slow responses and why. Correlate latency with customer characteristics to identify patterns.
Business metrics dashboards: expose custom analytics to your own customers. Query Analytics Engine from a Worker, format the results, return through your API.
The pattern that doesn't fit is real-time alerting. Analytics Engine is an analytics store, not a monitoring system. For alerts on specific conditions, use Workers Logs with external alerting or implement real-time checks in Tail Workers.
How much observability do you need?
Not every application needs the same observability investment. A side project and a revenue-critical API have different requirements.
Minimum Viable Observability
For low-stakes applications, internal tools, or early exploration:
- Workers Logs enabled, with retention and sampling checked against the investigation window
wrangler tailfor active debugging- Built-in dashboard metrics
- No custom instrumentation
Cost: often within the paid plan's twenty million included log events each month. Free and paid Workers plans have different allowances and retention, so check both against the incident response window. This baseline suits systems where investigation can be slow.
Production-Grade Observability
For revenue-generating systems where incidents have real cost:
- Workers Logs for operational investigation and querying
- Logpush to R2 or a log platform for retention beyond seven days
- Structured logging with request ID propagation
- Custom metrics via Analytics Engine for business-critical paths
- Alerting on error rates and latency by endpoint
Cost: meaningful but proportional. Appropriate when fast detection and diagnosis matter.
Enterprise Observability
For systems where Cloudflare is core infrastructure:
- All of the above
- Native tracing across Workers and Durable Objects, with custom instrumentation for uncovered paths
- Regional alerting (not just global aggregates)
- Integration with existing APM and incident management
- Dedicated observability budget and ownership
Cost: significant investment. Appropriate when operational excellence is a competitive advantage.
Invest in observability proportional to the cost of incidents. A side project can tolerate hours of debugging. A payment system cannot.
Logging strategy
Every log line costs money to store and attention to analyse. At edge scale, the question isn't what to log but what's worth logging.
What earns its place
Some events always warrant logging:
Request boundaries provide the skeleton for understanding system behaviour. Log when requests arrive and complete, with enough context to correlate them.
Errors and exceptions are obviously essential. You can't debug what you don't know occurred.
External service calls are where edge applications spend most of their time and encounter most of their failures. When your Worker calls D1, R2, an external API, or a Durable Object, log the call, its duration, and its outcome.
Authentication decisions matter for security auditing. Who accessed what, and were they authorised?
Business-critical actions (purchases, account changes, data modifications) provide audit trails that outlive debugging needs.
Passwords, API tokens, session secrets, credit card numbers, government identifiers: never log these regardless of how useful they'd be. The debugging benefit never outweighs the security risk.
Structure for correlation
Use a stable operation or request identifier across logs. Include the authenticated tenant when relevant, the deployment version, operation name, outcome and useful timing. Pass the identifier across HTTP or RPC boundaries and retain a business identifier when work continues through queues or workflows.
A correlation identifier connects records; it does not authenticate the caller. Keep user-supplied values separate from trusted identity and avoid copying arbitrary headers into every downstream log.
The debuggable log line
An error needs enough context to narrow the cause, not an unrestricted copy of the request. Record the failing operation, error type and safe stack information, relevant identifiers and the input shape or validation result. For example, knowing that user.id was absent may be enough; retaining the user's complete profile is unnecessary.
Use that evidence to construct a local reproduction against the same deployed version. Log additional payload details only where they answer a specific diagnostic question and the data-handling policy permits it. Production logs and traces should support investigation without becoming another uncontrolled store of customer data.
Tracing
Workers provide automatic tracing for supported binding operations and JavaScript RPC across Workers and Durable Objects. Custom instrumentation connects paths that the platform cannot follow, particularly calls into external systems.
Automatic tracing
Automatic tracing captures timing and metadata for supported Cloudflare operations, including D1 queries, KV reads, R2 operations, and JavaScript RPC calls. Enable it alongside Workers Logs in your Wrangler configuration:
{
"observability": {
"enabled": true,
"traces": {
"enabled": true,
"head_sampling_rate": 0.01 // Sample 1% of traces
}
}
}
Traces appear in the Workers Observability dashboard alongside your logs, showing which operations consumed time and where errors occurred within a request. This answers questions that logs alone cannot: which database query was slow, which external call failed, where a request spent its time.
Traces can also be exported using the OpenTelemetry Protocol (OTLP) to destinations such as Honeycomb, Sentry and Grafana. Configure a tracing destination in the Cloudflare dashboard. This replaces the manual waitUntil(fetch(...)) pattern for sending trace data to external collectors.
Automatic tracing is in open beta. The core capability is stable and recommended by Cloudflare's own best-practices guide, but expect the configuration surface to evolve. A future compatibility date will enable tracing by default when observability.enabled is set.
Tracing across service boundaries
Automatic tracing follows JavaScript RPC sessions across Workers and into Durable Objects, including individual method calls, nested calls, and callbacks. A session span groups calls that share the same session, while the callee spans show where execution continues. Start with this view when investigating a slow service chain; add custom instrumentation for gaps you can identify.
Traditional distributed tracing assumes services run in known locations. Edge tracing adds geographic distribution that traditional tools weren't designed for. Your Worker runs wherever the user is closest. If it calls a Durable Object, that object lives in a specific location, potentially different from where the Worker executed. If the Durable Object calls another Worker via service binding, that Worker runs co-located with the object, not with the user. A single request might traverse multiple continents without explicit routing.
What to trace
Service binding calls, Durable Object invocations, and external API calls warrant tracing. These are the boundaries where requests cross locations or leave the Cloudflare network.
Internal computation within a single Worker usually doesn't need span-level tracing. Curious where time goes within a Worker? Profile locally. Production tracing should focus on distributed aspects that can't be reproduced locally.
Propagating context
For an external service, agree on a trace-context format with the receiving system and propagate it with the request. Keep application correlation IDs alongside native traces when work continues through queues, callbacks, or long-running business processes. An order ID can connect events separated by hours even when no single request trace spans the whole operation.
Custom headers carrying an ID do not automatically join native traces. Verify that both systems extract and record the context you send, and use your observability platform's supported propagation format when you need one distributed trace.
Integration with external tracing
For automatic traces, configure an OTLP-compatible export destination in the Cloudflare dashboard. Traces flow to Honeycomb, Grafana, Sentry, or any collector supporting the OpenTelemetry protocol without code changes.
For uncovered boundaries, use an instrumentation library and exporter compatible with the receiving system. Arbitrary JSON posted to a traces URL is not necessarily valid OTLP. Keep export work outside the critical response path where possible, and monitor failures and dropped telemetry.
Automatic tracing closes much of the gap with hyperscaler APM for operations within Cloudflare's ecosystem. JavaScript RPC chains have native visibility, while external systems and unsupported paths still need deliberate context propagation. This does not provide production CPU flame graphs or complete application-level attribution.
When to invest in custom tracing: For simpler architectures (a Worker calling D1 and maybe one external API), automatic tracing combined with structured logging covers most debugging needs. Invest in custom tracing when native spans leave a specific operation unexplained or when edge traces must join an existing distributed tracing system. A chain of Workers alone is not sufficient reason to build your own tracing layer.
Detecting named failure modes
The failure patterns named earlier in the book suggest specific evidence to collect.
Timing assumption violations (Chapter 5): high wall time with low CPU can indicate serial I/O or a slow dependency. Compare spans and call timings with the local assumption. Record a safe query identifier and rows processed, rather than treating a truncated SQL string as redaction.
Placement latency mismatch (Chapter 7): compare Durable Object call latency by the calling Worker's colo and object identifier. That can reveal a geographic pattern, but it does not directly report where the object ran. Investigate placement alongside dependency load and query time before attributing every slow call to distance.
Poison message loops (Chapter 9): track repeated operation identifiers, attempt counts, error classes and dead-letter outcomes. Avoid logging raw message bodies by default. Alert when retries are no longer making progress and retain enough information to replay safely after fixing the cause.
Stale read after write (Chapter 15): correlate the written version with the version subsequently read. Timestamps alone may not identify which value was returned. If the application needs current state, reconsider the store or read path rather than treating stale reads as an intermittent network fault.
These signals complement general error and latency alerts. They narrow the investigation; they do not prove a cause without supporting evidence.
Durable Object observability
Durable Objects have different observability characteristics than stateless Workers. They persist between requests, maintain state, and have lifecycles extending beyond individual invocations.
State visibility
A Durable Object's in-memory state exists only while the object is active. Once it hibernates or is evicted, the in-memory state is gone. Persistent state in SQLite storage survives. Data Studio, in beta, exposes SQL-backed data in the dashboard, but a current database view cannot reconstruct how the object reached that state or reveal its live JavaScript memory.
Record meaningful state transitions with the object identifier, previous and new state, triggering operation and outcome. Log the transition after its durable commit, and distinguish an attempted change from a committed one. This history explains how the current database state was reached without repeatedly logging the entire object.
Connection monitoring
Record connection opens, closes and close reasons, then compare them with active connection counts and error rates. Keep connection metadata per WebSocket, including the start time when duration matters.
Hibernation discards ordinary in-memory fields. Use reconstructible connection metadata, such as WebSocket attachments, for information needed after wake-up. One object-level timestamp cannot describe every connection in a room.
Alarm observability
Record the intended alarm time when scheduling it, then compare that stored value with the handler's start time. Calling getAlarm() inside the running handler normally returns null, so it cannot supply the timestamp of the alarm being handled.
Log the handler's retry metadata, outcome and next scheduled action. An uncaught exception triggers automatic retries, but those retries are bounded. Alert on exhausted recovery and make a deliberate decision about rescheduling unfinished work; do not log willRetry: true for every failure indefinitely.
Alerting for geographic distribution
Alert thresholds on the edge must account for distribution. A 2% error rate might mean a global problem or a submarine cable cut affecting one region. Your alerts should know the difference.
Regional vs global
The most useful distinction is between regional and global incidents. Regional: one geography affected, perhaps a network issue, a regional service degradation, or a deployment problem in specific locations. Global: everywhere simultaneously, typically a code bug, a configuration change, or a dependency failure.
Configure both: "Error rate > 5% globally" catches universal failures; "Error rate > 20% in any single region" catches localised issues that might not register in global aggregates. The second alert is crucial. Singapore at 50% errors but only 5% of traffic means global error rate might be 2.5%, below your threshold. Your Singapore users are suffering while your dashboard shows green.
Baseline calibration
Establish baselines for the actual request paths. Geographic distribution, data placement, cache warmth, model inference and backend behaviour can all widen latency distributions even when Worker startup is small.
Establish baselines per metric and per region before setting alert thresholds. P99 latency of 500ms might be concerning for requests served from the same continent as your database, but normal from the opposite hemisphere.
Threshold selection
Overly sensitive alerts create alert fatigue. Fire three times daily and usually false? You'll start ignoring it and miss the real incident. Overly lenient alerts miss real problems. Fire only when 50% of requests fail? You'll have angry users before you know.
The right threshold is the lowest value that doesn't produce false positives during normal operation. Find it empirically: collect baseline data for a week, calculate normal variance, set thresholds above that variance. Error rate normally between 0.1% and 0.3%? Alerting at 0.5% catches real problems without noise. Alert at 0.2%? Constant pages for normal fluctuation.
Alert routing
Alerts should reach people who can understand the problem and take action. The team that wrote the code should receive alerts about that code's behaviour.
When developers carry operational responsibility for their own systems, they build systems that are easier to operate. When operations is someone else's problem, developers optimise for shipping features, not operational clarity. Structure alerting so consequences flow to decision-makers. New feature causes latency spikes? The team that shipped it should know first.
This isn't about blame; it's about learning. The fastest path to reliable systems is feeling the results of your choices directly.
Incident response
When alerts fire, speed matters. The difference between five-minute and five-hour incidents is often whether you had a plan before the alert fired.
Restore service before investigating deeply
When an incident follows a deployment, check whether routing back to a retained version can restore service. Confirm that the previous code remains compatible with current data and in-flight work; a code rollback cannot reverse mutations or external effects. Chapter 22 covers that preparation.
Mitigation without deployment
Feature flags can disable a failing path or reduce dependency load, provided the fallback has been designed and exercised. Choose the flag store according to how quickly the change must take effect.
KV suits flags that tolerate cached values and delayed propagation. A change can take 60 seconds or longer to become visible elsewhere, so it is not a guaranteed immediate kill switch. For urgent global intervention, use a control path whose propagation and failure behaviour you have measured. Document what happens when the flag store itself is unavailable.
Post-incident learning
After incidents resolve, understand what happened. Not to assign blame, but to improve systems.
What happened? Timeline of events, from first symptom to resolution. Be specific about times and actions.
Why didn't we detect it faster? Alerts missing or misconfigured? Failure mode not matching any alert? Should we add specific detection?
Why didn't we resolve it faster? Rollback not an option? Missing feature flags? On-call engineer lacking access or context?
What systemic change prevents recurrence? Not "person X will be more careful" but concrete changes: new alerts, updated runbooks, additional tests, architectural changes.
The output is action items with owners and deadlines. Post-incident review without action items is storytelling, not improvement.
Carrying an existing observability practice across
Workers Logs supplies managed retention and querying, while native traces cover supported bindings and JavaScript RPC calls. Teams moving from CloudWatch or Application Insights should inventory the signals their current incident process uses, then map each one to native telemetry or a deliberate export.
Check coverage rather than assuming that a product called APM follows every operation. Verify propagation into external services, retention, sampling, query access and alert ownership. Keep an existing cross-platform observability system when it provides a coherent investigation path across Workers and the dependencies that remain elsewhere.
Local profiling remains useful for CPU attribution; production logs and traces answer a different question about which operations and dependencies failed. A small representative incident exercise reveals these gaps more clearly than a feature comparison table.
Maintaining observability
Telemetry can drift from the application as log schemas, dependencies and traffic change. Review it as part of feature delivery, and assign ownership for the collection and alerting path itself.
Review alert effectiveness quarterly. Which alerts fired? Which were actionable? Which were noise? Alerts that never fire might be misconfigured. Alerts that fire frequently and get ignored are worse than useless; they train your team to dismiss alerts.
Audit log schemas when code changes. New features should log appropriately. Deprecated features should stop cluttering logs. Engineers shipping features should update observability as part of the work.
Test that alerts actually fire. An alert that has never triggered might be broken. Periodically verify your alerting pipeline end-to-end: generate a synthetic error, confirm the alert fires, confirm the notification reaches the right person.
Schedule this work alongside application maintenance. Choose a cadence that matches the rate of change and review the collection path after significant releases or incidents.
What comes next
The same evidence that explains an incident should help assess a deployment. Chapter 22 connects observability to permissions, data boundaries and controlled releases, so you can establish both what changed and who had the authority to change it.