Skip to content

Monitoring

Each deployment renders its own CloudWatch dashboard, named after the stack; the DashboardUrl stack output links to it. It is built in cdk/stack.py from the same metric objects the alarms use, so there is nothing to import or configure — deploy and it exists.

Alarms

Five alarms page (via ALARM_EMAIL, when set):

Alarm Fires when It usually means
DlqMessagesAlarm anything lands in the DLQ granules rejected 20 times — a UR/time collision or a persistent parse failure; the dashboard's Rejected granules table shows which (by url for validation rejections, by SQS message id for granules that raised mid-processing)
ConsumerErrorsAlarm the SQS consumer throws check the consumer's log group
PollerErrorsAlarm the CMR poller fails its run and both async retries CMR unreachable, or watermark state unreadable
ResortErrorsAlarm the re-sort job fails its run and both async retries (a single failure that a retry heals, e.g. a lost promote CAS, does not page) the fold failed before promoting; the ledger keeps growing until fixed
AxisEndLagAlarm the store's newest time slot is > 24 h old, or the AxisEndLag metric goes missing, for 24 consecutive hours the top-line staleness check. Treating missing data as breaching is deliberate: a dead poller or a re-sort killed by its timeout emits no error metric at all — the freshness metric going quiet is the only signal. The 24-hour evaluation window exists because TEMPO is daylight-only: the emitters legitimately go quiet overnight, and a single quiet hour must not page. Only created when forward processing is enabled. A fresh deployment may hold ALARM for up to its first day: the alarm's evaluation window predates the first emission, and those missing hours count as breaching until the first committed batch or promoted re-sort.

The two scheduled-job alarms count AsyncEventsDropped, not Errors: EventBridge invokes them asynchronously and Lambda retries a failed run twice, so a single failure is usually healed a minute later. To see how often that happens, the Poller and Re-sort failures widgets plot both series; failed attempts minus failed-after-retries is the healed share.

Throttles on the consumer are expected (its reserved concurrency is 1; SQS redelivers) and are displayed on the dashboard but never alarmed.

Custom metrics

The handlers emit CloudWatch metrics in the TempoPipeline namespace, dimensioned by Collection and Stage (so queries never need physical resource names):

Metric Emitted by How to read it
AxisEndLag (seconds) consumer after each commit; re-sort after each promote store freshness; production lag is normally a few hours — mostly upstream (scan → CMR publication), see runbook-production-lag for the attribution
GranulesRouted (dimension Route) consumer, per consumed batch APPENDED = growth, UNCHANGED = redeliveries of unchanged sources skipped without a write (the steady band; its absence with a flowing queue is the anomaly), OVERWRITTEN = genuine republications (rare), PENDING = out-of-order arrivals headed for the re-sort (routinely a large share), REJECTED = collisions headed for the DLQ (counted on first delivery only; redeliveries are not re-counted)
ProductionLag / CmrLag (seconds) poller, per fresh arrival (scan within 24 h, first seen this poll) upstream share of the lag: scan start -> ProductionDateTime, and ProductionDateTime -> CMR revision-date; stacked with VirtualizationLag on the Lag attribution widget
VirtualizationLag (seconds) consumer, per APPENDED granule whose message carries the poller's published the pipeline's share: CMR publication -> store commit; healthy is under POLL_SCHEDULE_MINUTES plus a few minutes
PendingLedgerDepth consumer and re-sort nonzero is healthy; trending up across days means the re-sort is not keeping pace
FoldedGranules re-sort (0 when it ran with an empty ledger) pinned at RESORT_MAX_FOLD every run means falling behind
PromoteFailures re-sort, when its promote raises occasional ones are the single-writer design working (a concurrent commit won the CAS); sustained ones mean writers are fighting — or S3 trouble, the counter does not distinguish
CommitFailures consumer, when its batch commit raises shown on the Consumer duration widget; the whole batch redelivers
PartitionsDone / PartitionsTotal backfill reduce / partition steps backfill progress; the dashboard plots their running sum against the carried-forward total
CompletenessDelta verify_store.py --completeness (CodeBuild) granules CMR lists that the store lacks, plus store entries CMR dropped; one point per verify run, so the series is sparse

Emission is best-effort: a metric failure never fails a batch or a re-sort run. If a widget shows no data, first check the corresponding job has actually run (e.g. CompletenessDelta appears only after a verify run).

During a backfill

The dashboard's backfill section (rendered when BACKFILL_ENABLED) shows Step Functions executions, the cumulative partitions-done/total graph, and worker errors — watch it during the initial fill. Afterward, the two numbers worth a daily glance are the Store lag (scan -> store) and Pending ledger depth tiles; the Lag attribution widget next to them says how much of the lag is upstream.