Incident lab · ABR ladder drift on mid-stream CDN failover
Synthetic postmortem · apac/teal weekday primetime multi-CDN failover

ABR package-list drift
on a mid-stream CDN failover.

A working postmortem on the apac/teal weekday primetime broadcast that, mid-stream, pushed a regional slice of viewers from CDN-A to CDN-B through a mid-stream failover after CDN-A's regional POP saturated — CDN-A kept publishing the live-edge tail #EXT-X-MEDIA-SEQUENCE:N while CDN-B anchored a fresh replay-origin timeline #EXT-X-MEDIA-SEQUENCE:N+k. The two CDNs' ladders diverged by k=6 segments (~36 s of ABR package-list drift); the cohort's cdn.pop.ladder_segment_index_delta reads fail, the cohort's ladder-rendition re-aggregate probe fails across 240p / 540p / 720p / 1080p, while cache.availability stayed pass and the multi-CDN health probe multicdn.winner_RTT stayed flat on cdn-a / cdn-b / cdn-c. The Streamwake agentic ops layer classified it as cdn_abr_ladder_drift · dominant at 81% confidence with the cdn_cache_segment_miss lane ruled out by name on the cache leg + segment-leg + license-rollout posture, and remediated with an authorization-tiered policy: Tier 0 autonomous under gate (reissue ABR ladder via replay-origin reanchor), Tier 1 surfaced to humans (CDN-B operator side ladder resync), Tier 2 surfaced for operator-team approval forward (codify the new probe into the cohort fan-in) — recovery verified cohort-side + ladder-side across T+30 s → T+15 m, NOT cache-side.

Protocol: HLS · CMAF · mid-stream failover · cdn-A/sin02 · cdn-B/hkg01 · multi-CDN cdn-a|cdn-b|cdn-c
Window: apac/teal weekday primetime — mid-stream k=6 ladder drift (~36 s PDAT skew)
Streamwake probes: cdn.pop.ladder_segment_index_delta · cdn.pop.cross_cdn_alignment_probe · cohort.ladder_rendition_reaggregate_probe · cdn.pop.affected_renditions · cache.availability · cache.segment_leg_cache_hit · cdn.license_rollout_posture_check · multicdn.winner_RTT · cohort.cohort_join_stall_ratio · cohort.cohort_rebuffer_ratio · cohort.cohort_startup_time_p95_ms · edge.egress_kbps.

Book a technical demo for ABR ladder drift on mid-stream CDN failover

Lead magnet
ABR ladder drift on mid-stream CDN failover

Read the postmortem — then bring your own incident to Streamwake.

Two ways to engage on this exact failure pattern: book a 30-minute technical demo where we walk through the probe cascade on your source, or hand us an archived incident and watch the agent diagnose it end-to-end.

Both routes land on the scoping intake form — no SDR gate.

Viewer impact

What the cohort saw

The first things to read on any real mid-stream CDN failover package-list drift incident are the cohort-level numbers on the affected apac/teal cohort — and crucially, the contrast against the cache leg that stayed green across both CDNs. Three numbers did the heavy lifting here: the ladder-segment-index delta, the cohort join-stall ratio, and the contrast between the cross-CDN ladder alignment probe (fail) and the cache.availability / segment-leg / license-rollout posture probes (all pass on the affected edge POPs across both CDNs). The figures below are simulated telemetry — the disclosure above applies to every figure on this page.

Co-affected (apac/teal cohort)
Sessions whose ABR ladder crossed the cross-CDN alignment probe — CDN-A continuing the live-edge tail while CDN-B anchored a fresh replay-origin timeline.
~14% of cohort

Roughly 14% of the apac/teal weekday primetime cohort — session-level ABR ladder crossed on the failover slice from cdn-A/sin02 to cdn-B/hkg01, where cdn.pop.ladder_segment_index_delta lifted from 0 baseline to k=6 (~36 s of timeline divergence) on both CDNs and cohort.ladder_rendition_reaggregate_probe failed across renditions 240p / 540p / 720p / 1080p. The cache leg was green — cache.availability stayed X-Cache: HIT and cache.segment_leg_cache_hit stayed pass on both CDNs through the entire window — but the cohort's rebuffer ratio lifted to 0.082 behind the mid-stream ladder PDAT drift.

Duration
Window from first cross-CDN ladder alignment probe variance to last cohort probe returning to baseline.
~41 minutes

~41 minutes between the first cross-CDN ladder alignment probe variance at T+0 m and the apac/teal cohort's cohort_join_stall_ratio + cohort_rebuffer_ratio + ladder-rendition re-aggregate probe + cohort startup-time settling within tolerance at T+41 m. The window goes longer than the natural ladder-side reanchor settle because the cdn-B operator-side ladder resync (Tier 1) and the operator-team sign-off on the new probe (Tier 2) ship after the close-out probes land — so the postmortem window is two-tier: T+0 m → T+18 m on the cohort-side + ladder-side close-out, and shaped forward by the cdn-B resync + Tier 2 sign-off that bound forward to the next weekday primetime broadcast.

Surface area
Where the symptom landed — cross-CDN manifest timeline drift on the mid-stream ladder mismatch.
cross-CDN ABR · manifest PDAT drift · k=6 segments

cdn.pop.ladder_segment_index_delta: k=6 (~36.04 s PDAT drift) on cdn-A/sin02 vs cdn-B/hkg01 (apac/teal cohort) — CDN-A continuing the live-edge tail at #EXT-X-MEDIA-SEQUENCE:84217 /#EXT-X-PROGRAM-DATE-TIME:2026-08-20T18:42:11.000Z while CDN-B serves the fresh anchor at #EXT-X-MEDIA-SEQUENCE:84223 / #EXT-X-PROGRAM-DATE-TIME:2026-08-20T18:42:47.040Z. The discriminator is that the cache leg stayed green — cache.availability = X-Cache: HIT, cache.segment_leg_cache_hit = pass, cdn.license_rollout_posture_check = pass, and the multi-CDN health probe multicdn.winner_RTT = 40 / 41 / 42 ms (cdn-a / cdn-b / cdn-c flat at baseline). The failure is on the ladder manifest timeline alignment, NOT a cache-miss posture on either CDN leg.

Classification

How Streamwake classified this incident

Three ranked hypotheses: the top one filing the timeline as cdn_abr_ladder_drift · dominant, the second explicitly tagged cdn_cache_segment_miss · ruled out by so the recovery message lands (the cache leg stayed clean on both CDNs through the entire window), and the third filed as an alternate: cdn_pop_route_bouncing_under_failover with a low confidence that captures the rtt-side stampede shape before the multi-CDN health probe removes it.

Top hypothesis (failure lane)
What the agent named first — the failure category the timeline is filed under.
cdn_abr_ladder_drift · dominant · 0.81

cdn_abr_ladder_drift — a mid-stream CDN failover from CDN-A to CDN-B on a regional slice of the apac/teal weekday primetime cohort pushed the cohort's breadcrumb CDN-A to keep publishing the live-edge tail #EXT-X-MEDIA-SEQUENCE:N while the failover slice on CDN-B got a freshly-anchored timeline #EXT-X-MEDIA-SEQUENCE:N+6; the two CDNs' ladders diverged by k=6 segments, #EXT-X-PROGRAM-DATE-TIME drifted past the cohort's expected join window on the affected slice, and the cohort's ladder-rendition re-aggregate probe failed on 240p / 540p / 720p / 1080p. Four signals line up: ladder_segment_index_delta k=6 vs 0 baseline (cross-CDN), cross_cdn_alignment_probe reads fail, cohort.ladder_rendition_reaggregate_probe fails across monitored renditions, cohort_join_stall_ratio 0.124 vs 0.018 baseline.

Secondary signal (cause lane)
The cdn_cache_segment_miss lane is ruled out by name — discards the cache-miss / segment-leg hit-rate shape that would have gated any CDN-side cache runbook lane on this incident.
cdn_cache_segment_miss · ruled out by · 0.24

cdn_cache_segment_miss was the secondary signal ranked at 24% — but it's tagged ruled out by so the recovery message lands. cache.availability reads X-Cache: HIT on both cdn-A/sin02 and cdn-B/hkg01, cache.segment_leg_cache_hit reads pass on the affected cohort on both CDNs, and cdn.license_rollout_posture_check reads pass on the affected ladder — the failure is on the ladder alignment, NOT on the cache-locale. The dismissal rule was "rank the cause on the cache.availability + segment-leg cache-hit + license-rollout posture pattern, not on the player-visible symptom alone"; the ladder-alignment nature of the failure is the decisive signal pattern.

Severity, region, status
Severity is computed from the co-affected cohort share; region is the geo of the failing probes.
sev3
  • Region: apac/teal weekday primetime broadcast (cdn-A/sin02 → cdn-B/hkg01 failover slice)
  • Status: resolved (window closed; Tier 1 cdn-B resync configured; Tier 2 new probe staged for operator-team sign-off forward)
  • Opened: 2026-08-20 18:42 UTC
  • Spread: contained to the apac/teal weekday primetime broadcast — na-east and eu-west cohorts on the same broadcast unaffected; adjacent multi-CDN cohorts on the same vendor unaffected across all three CDNs; the multi-CDN health probe settled at zero egress change on the apac/teal edge legs.
Confidence
Top hypothesis share of the three ranked hypotheses; remaining mass is split between the ruled-out cdn_cache_segment_miss and the alternate cdn_pop_route_bouncing_under_failover.
81 / 100

Above the 80% threshold the agent treats as a confident top-hypothesis filing. cdn_cache_segment_miss · 0.24 was cleared explicitly because cache.availability, cache.segment_leg_cache_hit, and cdn.license_rollout_posture_check together prove the cache leg was healthy at zero egress change across both CDNs — the discriminator for the ladder-alignment-not-cache-locale shape of the failure. The multicdn.winner_RTT probe reading pass on cdn-a|cdn-b|cdn-c (40/41/42ms flat at baseline) clears cdn_pop_route_bouncing_under_failover as the alternate.

Authorization-tiered remediation policy

Who clears the gate

The governed-action policy on this incident is split across three authorization tiers — the agent may act on its own under an explicit gate (Tier 0), surfaces the work to the cdn-B operator team (Tier 1), or hands the new probe codification to the operator-team sign-off forward (Tier 2). Each tier has an explicit gate; each gate names the threshold before the action lands, and the lane the action belongs to is what makes the split distinct from a generic postmortem.

Tier 0 · autonomous under gate
agent-emitted
Tier 0 acts on its own
Authorization tier 0 — agent may clear the gate and act autonomously. The operator clears via configuration system; the agent emits on its own.

reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill autonomous under gate: top_confidence >= 0.80 AND cdn.pop.ladder_segment_index_delta >= 3 AND multicdn.winner_RTT.pass == true on the affected cohort. The agent clears the gate and reissues the ABR ladder via replay-origin reanchor at T_mid plus a CDN-side segment_index backfill so the cohort resolves on the refreshed ladder — the cache leg is NOT touched, the multi-CDN health probe proves the CDN leg is healthy at zero egress change.

Tier 1 · surfaced to humans
surfaced
Tier 1 needs CDN-side coordination
Authorization tier 1 — surfaces to the cdn-B operator team. Touches another provider / requires a peering or vendor-side action the agent cannot perform.

surface_cdn_failover_evidence_to_cdn_b_for_ladder_resync_via_origin_clip surfaced to humans: cdn_abr_ladder_drift >= 3 segments AND top_confidence >= 0.75. The operator at the cdn-B side configures the ladder anchor / segment_index backfill before cdn-B's manifest tail fully publishes — a misanchor that surfaces on a future primetime isn't masked by a "we reconciled" close-out.

Tier 2 · operator-team sign-off forward
approval forward
Tier 2 codifies the new probe
Authorization tier 2 — surfaces for operator-team sign-off forward. Configuration-owner sign-off on the new probe going into the cohort probe fan-in for future mid-stream failover windows.

surface_cdn.pop.ladder_segment_index_delta_probe_to_cohort_fanin surfaced for operator-team approval forward: incident_resolved AND cohort_side close-out signals cleared. This is the learn-loop act — once the lane is closed, the operator team signs off on pinning the new probe (cdn.pop.ladder_segment_index_delta) into the cohort probe fan-in for future mid-stream failover windows; cohort.ladder_rendition_reaggregate_probe is pinned for the next four weekday primetime windows so the false-positive rate can be measured before promotion.

Chronology

Incident timeline

Eleven events: detection on the apac/teal cohort, classification across three ranked hypotheses (with the cdn_cache_segment_miss lane tagged ruled out by so the recovery message lands), five acts the agent took under the authorization-tiered governed-action gates (Tier 0 autonomous under gate + a status-change handoff at T+5 m + a T+18 m re-probe + a T+25 m status change), four acts it surfaced to humans (Tier 1 cdn-B surface, Tier 2 operator-team sign-off, cdn-B operator ack, reliability-team assignment), and the resolution. The right-hand "act" + "tier" tags are what makes this postmortem distinct from a generic write-up — they pin the split between autonomous agentic ops, the cdn-B side coordination, and the operator-team sign-off forward. All times below are simulated telemetry — the disclosure at the top of this page applies to every minute offset on the timeline.

Today

11 events
  • T+0m
    Detection
    by cohort agent · apac/teal · cdn-A/sin02 + cdn-B/hkg01
    act · autonomous

    Cross-CDN ladder alignment probe fan-in reads fail at k=6

    cdn.pop.ladder_segment_index_delta crossed 0 → 6 within a 30 s window on cdn-A/sin02 AND cdn-B/hkg01; cross_cdn_alignment_probe reads fail (CDN-A live-edge tail at #EXT-X-MEDIA-SEQUENCE:84217 vs CDN-B replay-origin anchor at #EXT-X-MEDIA-SEQUENCE:84223 — ~36.04 s PDAT drift); cohort.ladder_rendition_reaggregate_probe reads fail across 240p / 540p / 720p / 1080p; cohort.cohort_join_stall_ratio lifted to 0.124 (vs 0.018 baseline) and cohort.cohort_rebuffer_ratio lifted to 0.082 (vs 0.012 baseline). cache.availability stays pass on both CDNs; segment-leg cache-hit stays pass; multicdn.winner_RTT stays flat on cdn-a / cdn-b / cdn-c.

    Aug 20, 06:42:11 PM
  • T+1m
    Classification
    by Streamwake reliability agent
    act · autonomous

    Ranked: cdn_abr_ladder_drift · dominant (0.81) · cdn_cache_segment_miss · ruled out by (0.24) · cdn_pop_route_bouncing_under_failover · alternate (0.21)

    Top hypothesis reads 81% confidence. cdn_cache_segment_miss is ruled out by name on cache.availability + segment-leg cache-hit + license-rollout posture — the cache leg was green the entire window on both CDNs and the failure is on the ladder manifest timeline alignment. cdn_pop_route_bouncing_under_failover is the alternate (the multi-CDN health probe reads pass on cdn-a|cdn-b|cdn-c — flat at baseline; the discriminator is on-segment-index delta, not on rtt_ms signature).

    Aug 20, 06:43:11 PM
  • T+2m
    Automated action · autonomous
    by Streamwake reliability agent
    act · autonomous
    Tier 0

    Governed · Tier 0 autonomous under gate: reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill

    Top_confidence 0.81 >= 0.80 gate cleared; cdn.pop.ladder_segment_index_delta k=6 >= 3 gate cleared; multicdn.winner_RTT.pass == true gate cleared on cdn-a | cdn-b | cdn-c. Reissues the ABR ladder via replay-origin reanchor at T_mid plus a CDN-side segment_index backfill so the cohort resolves on the refreshed ladder — the cache leg is NOT touched (it was green the entire window), the multi-CDN health probe proves the CDN leg is healthy at zero egress change.

    Aug 20, 06:44:11 PM
  • T+3m
    Surfaced to human
    by agent → cdn-B operator team
    act · surfaced to humans
    Tier 1

    Governed · Tier 1 surfaced to humans: surface_cdn_failover_evidence_to_cdn_b_for_ladder_resync_via_origin_clip

    CDN-side coordination required. Cdn_abr_ladder_drift >= 3 segments (k=6) AND top_confidence 0.81 >= 0.75 gate cleared. Surfaces the mid-stream failover evidence (CDN-B manifest timeline PDAT drift, k=6 segment skew) to the cdn-B operator team so they configure the ladder anchor / segment_index backfill before cdn-B's manifest tail fully publishes — a misanchor that surfaces on a future primetime isn't masked by a "we reconciled" close-out.

    Aug 20, 06:45:11 PM
  • T+5m
    Status change
    by Streamwake reliability agent
    act · autonomous

    T+5 m handoff: ladder probe re-reads dropping; join-stall settle milestone

    cohort.cohort_join_stall_ratio settled to within 0.05 tolerance by T+5 m (per the staged T+5 m join-stall settle gate — NOT a single-probe reanchor); ladder_segment_index_delta re-reads dropping from k=6 toward k=2 within the first same-length ladder window — ladder-side settle tracking ahead of cohort-side settle, as expected when the reanchor lands first.

    Aug 20, 06:47:11 PM
  • T+7m
    Surfaced to human
    by agent → reliability team
    act · surfaced to humans
    Tier 2

    Governed · Tier 2 surfaced for operator-team sign-off forward: surface_cdn.pop.ladder_segment_index_delta_probe_to_cohort_fanin

    Operator-team sign-off forward required. This is the learn-loop act — once the lane is closed, the operator team signs off on pinning the new probe (cdn.pop.ladder_segment_index_delta) into the cohort probe fan-in for future mid-stream failover windows; cohort.ladder_rendition_reaggregate_probe is pinned for the next four weekday primetime windows so the false-positive rate can be measured before promotion.

    Aug 20, 06:49:11 PM
  • T+10m
    Surfaced to human
    by cdn-B operator team
    act · surfaced to humans

    cdn-B operator acked; ladder anchor / segment_index backfill configured

    Acknowledged within 7 m of the Tier 1 surface; ladder anchor / segment_index backfill configured on the cdn-B side so the next apac/teal weekday primetime broadcast window opens with a reissued ABR ladder — the operator confirmed manifest timeline PDAT alignment to the live-edge tail on cdn-A live-edge tail cadence.

    Aug 20, 06:52:11 PM
  • T+15m
    Surfaced to human
    by reliability team
    act · surfaced to humans

    Postmortem write-up assigned (this page)

    Reliability team assigned the public postmortem; this page is the resulting write-up, with the three ranked hypotheses, the authorization-tiered remediation policy (Tier 0 / Tier 1 / Tier 2), the staged T+30 s → T+15 m viewer-level recovery verification window, and the learn-loop note pinning the new probe into the cohort fan-in.

    Aug 20, 06:57:11 PM
  • T+18m
    Automated action · autonomous
    by Streamwake reliability agent
    act · autonomous

    T+18 m re-probe: first ladder-segment-index delta re-read

    cdn.pop.ladder_segment_index_delta dropped from k=6 → k=2 within the first same-length ladder window (per the T+30 s staged gate — ladder-side settle ahead of cohort-side settle, as expected after the reanchor); cohort.cohort_rebuffer_ratio within 0.05 tolerance (per the T+2 m staged gate); cache.availability / segment-leg cache-hit stayed pass throughout — those probes are NOT a close-out signal (the cache leg was green the entire window).

    Aug 20, 07:00:11 PM
  • T+25m
    Status change
    by Streamwake reliability agent
    act · autonomous

    T+25 m status change: full cohort re-anchor cadence holds

    cohort.cohort_join_stall_ratio remains within 0.05 tolerance (per the T+5 m staged gate); cohort.ladder_rendition_reaggregate_probe reads pass on every monitored rendition (240p / 540p / 720p / 1080p); cohort.cohort_startup_time_p95_ms reads 1920 ms (within startup-time baseline, ahead of the next apac/teal weekday primetime broadcast window opens); cdn.pop.ladder_segment_index_delta sits at k=0 within the same window.

    Aug 20, 07:07:11 PM
  • T+41m
    Resolution
    by Operator + agent
    act · surfaced to humans

    T+41 m incident resolved; Tier 0 ladder reissue landed autonomously; Tier 1 cdn-B resync configured; Tier 2 new probe staged for operator-team sign-off forward

    Cohort-side + ladder-side close-out signals cleared on the apac/teal cohort for the current window (cohort_join_stall_ratio + cohort_rebuffer_ratio + cohort_startup_time_p95_ms + cohort.ladder_rendition_reaggregate_probe + cdn.pop.ladder_segment_index_delta within tolerance) — NOT on cache.availability / segment-leg cache-hit / multicdn.winner_RTT returning to pass (the cache leg was green the entire window, the multi-CDN health probe was healthy the entire window). The audit step codifies: "verify ladder-segment-index delta + cohort ladder-rendition re-aggregate settle within tolerance over the next same-length cohort window" — NOT cache-side close-out, NOT CDN-cache re-anchor.

    Aug 20, 07:23:11 PM
Autonomous acts the agent did
Five events the Streamwake reliability agent executed without a human in the loop. Each one was a governed action that cleared the gate — never an override.
  • classify · ranked three hypotheses with confidence in 90 s; cdn_cache_segment_miss · ruled out by tagged (cache leg stayed green)
  • Tier 0 ladder reissue · reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill under the top_confidence ≥ 0.80 + ladder_segment_index_delta ≥ 3 + multicdn.winner_RTT.pass gate
  • T+5 m status · observed ladder-side settle ahead of cohort-side settle; cohort_join_stall_ratio tolerance holding
  • T+18 m re-probe · ladder-segment-index delta k=6 → k=2 within the first same-length ladder window (T+30 s staged gate)
  • T+25 m status · cohort ladder-rendition re-aggregate pass on every monitored rendition; cohort_startup_time_p95_ms ahead of the next apac/teal weekday primetime window
Acts the agent surfaced to humans
Three events the agent did not act on its own — each one needed a cdn-B operator, a reliability-team owner, or a configuration-owner sign-off forward. The agent held off emitting a CDN-side cache re-route that would have masked the ladder-alignment root cause.
  • cdn-B operator team · Tier 1 surface — configured the ladder anchor / segment_index backfill before cdn-B's manifest tail fully publishes
  • reliability team · Tier 2 surface — sign-off forward on pinning cdn.pop.ladder_segment_index_delta + cohort.ladder_rendition_reaggregate_probe into the cohort probe fan-in
  • reliability team · assigned the public postmortem write-up (this page) — recovery criteria on the audit step is the cohort-side + ladder-side close-out signal, NOT the cache-side re-anchoring
Staged T+30 s → T+15 m viewer-level recovery

What recovery looks like across the ladder-side + cohort-side probes

Recovery on this incident is verified across the staged T+30 s → T+15 m window on the apac/teal cohort + on the cross-CDN ladder alignment probe — NOT on the cache layer turning green again. The cache leg stayed green the entire window; if recovery were the cache turning green, the postmortem would land on a CDN-side lane that was never the failure. The audit step writes the close-out signal into the playbook as a four-probe verification — ladder probe re-read at T+30 s → joined-cohort rebuffer settle at T+2 m → join-stall settle at T+5 m → full cohort re-anchor at T+15 m.

T+30 s — ladder probe re-read
first ladder probe
First ladder-segment-index delta re-read on the affected cohort within the first same-length ladder window (30 s cadence).

cdn.pop.ladder_segment_index_delta drops from k=6 → k=2 within the first same-length ladder window after the Tier 0 reissue lands (cadence 30 s). Ladder-side settle leads the cohort-side settle, as expected when the replay-origin reanchor lands first.

T+2 m — joined-cohort rebuffer settle
rebuffer settle
Joined-cohort rebuffer settle on the affected cohort within ten consecutive 30 s windows (NOT ladder-side settling).

cohort.cohort_rebuffer_ratio settles ≤ 0.05 within ten consecutive 30 s windows on the apac/teal affected cohort — NOT ladder-side settling; cohort-side settle reads the per-segment PDAT retune on the affected cohort.

T+5 m — join-stall settle
join-stall settle
Join-stall settle on the affected cohort within ten consecutive 30 s windows (NOT ladder-side settling); mid-stream rejoin validates.

cohort.cohort_join_stall_ratio settles ≤ 0.05 within ten consecutive 30 s windows on the apac/teal affected cohort — NOT ladder-side settling; the mid-stream rejoin (PDAT retune on viewers already joined at the point of the failover) validates on the same cadence.

T+15 m — full cohort re-anchor
full cohort re-anchor
Full cohort re-anchor; cohort ladder-rendition re-aggregate probe pass; cross-CDN ladder alignment probe ≤ 0 within the same window; cohort_startup_time_p95_ms ahead of the next weekday primetime window.

cohort.ladder_rendition_reaggregate_probe reads pass on every monitored rendition (240p / 540p / 720p / 1080p); cdn.pop.ladder_segment_index_delta sits at ≤ 0 within the same window; cohort.cohort_startup_time_p95_ms reads ahead of the next apac/teal weekday primetime broadcast window open.

NOT a close-out signal
The cache leg staying green · the multi-CDN health probe reading pass
The cache.availability / segment-leg cache-hit / multi-CDN health probe readings returning to pass are geometric — they read pass on the affected cohort across the entire window, NOT a cohort-side close-out. Reading those probes as the close-out signal is the read that misses the lane forward.

cache.availability returning to X-Cache: HIT at zero egress change is NOT a close-out signal — the cache leg was green the entire window on both CDNs. cache.segment_leg_cache_hit staying pass is NOT a close-out signal — the segment leg was green the entire window. multicdn.winner_RTT reading pass on cdn-a / cdn-b / cdn-c (40/41/42ms flat at baseline) is NOT a close-out signal — the multi-CDN health probe proves the CDN leg is healthy, but does not prove the cohort delivers green playback when the cross-CDN ladder manifest is misanchored.

Anatomy

Anatomy of the evidence packet

The two packets on the failing source — a multi-CDN cycle 0 (CDN-A live-edge tail) + cycle 1 (CDN-B replay-origin anchor) ABR ladder probe packet on the apac/teal affected cohort (with the cross-CDN ladder alignment probe failing on the cohort while cache.availability + segment-leg cache-hit + license-rollout posture stay pass on both CDNs), and the agent timeline response with the three ranked hypotheses, the authorization-tiered remediation policy, and the staged T+30 s → T+15 m viewer-level recovery verification window. The probe packet is what the agent decided on; the timeline response is what the agent emitted.

ABR ladder probe packet (multi-CDN cycle 0 + cycle 1, apac/teal cohort, T+0m)
GET /live/event/stream.m3u8 HTTP/1.1
host: cdn.example.com
accept: application/vnd.apple.mpegurl

----- cycle 0 (T+0m, mid-stream · CDN-A live-edge tail · apac/teal cohort) -----
# CDN-A continues the live-edge tail — the breadcrumb viewers already joined against
HTTP/2 200
content-type: application/vnd.apple.mpegurl
server-timing: manifest-fetch;dur=124
x-cdn: cdn-A/sin02
x-pop: cdn-A/sin02
x-cache: HIT
x-multicdn-route: cdn-A

#EXTM3U
#EXT-X-VERSION:9
#EXT-X-TARGETDURATION:6
#EXT-X-MEDIA-SEQUENCE:84217
#EXT-X-PROGRAM-DATE-TIME:2026-08-20T18:42:11.000Z
#EXTINF:6.000,
084217.ts

# ladder_segment_index_delta:                  fail  (cdn-A continuing live-edge tail at #EXT-X-MEDIA-SEQUENCE:84217)

----- cycle 1 (T+0m, mid-stream · CDN-B replay-origin anchor · apac/teal slice) -----
# CDN-B serves a fresh replay-origin anchor — the new timeline for the failover slice
HTTP/2 200
content-type: application/vnd.apple.mpegurl
server-timing: manifest-fetch;dur=131
x-cdn: cdn-B/hkg01
x-pop: cdn-B/hkg01
x-cache: HIT
x-multicdn-route: cdn-B
x-failover-mode: replay-origin-anchor-at-t-mid

#EXTM3U
#EXT-X-VERSION:9
#EXT-X-TARGETDURATION:6
#EXT-X-MEDIA-SEQUENCE:84223
#EXT-X-PROGRAM-DATE-TIME:2026-08-20T18:42:47.040Z          ← 36.04 s ahead of CDN-A's PDAT
#EXTINF:6.000,
084223.ts

# cdn.pop.ladder_segment_index_delta:          fail  (k=6 across all 4 renditions on cdn-A/sin02 AND cdn-B/hkg01 —
#                                                    mid-stream manifest timeline divergence;
#                                                    CDN-A live-edge tail vs CDN-B replay-origin anchor)
# cdn.pop.ladder_segment_index_delta.k:        6     (vs 0 baseline — ~36 s drift)
# cdn.pop.affected_renditions:                 [240p, 540p, 720p, 1080p]
# cdn.pop.cross_cdn_alignment_probe:           fail  (CDN-A continuing against CDN-B anchor — ladders diverged)
# cdn.pop.affected_pops:                       [cdn-A/sin02, cdn-B/hkg01]
# cohort.ladder_rendition_reaggregate_probe:   fail  (240p / 540p / 720p / 1080p re-aggregate probe fan-in read fail
#                                                    across renditions; rendition-level segment_index drift != 0)
# cohort.cohort_join_stall_ratio:              0.124 (vs 0.018 baseline, 90s cohort window — late-join stalled behind
#                                                    the ladder mismatch; viewer PDAT drift past expected join window)
# cohort.cohort_rebuffer_ratio:                0.082 (vs 0.012 baseline — playback stalled on the affected cohort)
# cohort.cohort_startup_time_p95_ms:           6210  (vs 1840 baseline — TTFF lifted above the 4 s threshold)
# cohort.affected_cohort_id:                   apac-teal-primetime-broadcast
# cache.availability on the affected edge POP: pass  (X-Cache: HIT on cdn-A/sin02 AND cdn-B/hkg01 — segment leg green)
# cache.segment_leg_cache_hit:                 pass  (segment leg green on the affected cohort on both CDNs)
# cdn.license_rollout_posture_check:           pass  (license-rollout posture on the affected ladder is green —
#                                                    ladder rendered correctly across all renditions)
# multicdn.winner_RTT on cdn-a | cdn-b | cdn-c: pass  (cdn-a 40ms, cdn-b 41ms, cdn-c 42ms — flat at baseline;
#                                                     no rtt-side failover stampede; the failure is on-segment-index delta,
#                                                     not on rtt_ms signature)
# edge.egress_kbps:                            6421  (at-or-above expected — cohort IS delivering traffic; the failure
#                                                    is on the ladder manifest, not on edge egress shape)
# cdn_pop.partial_segment_warmup_state:        primed (segment leg green — not a partial-segment warm-up failure)

----- cycle 2 (T+~12 m, after reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill) -----
HTTP/2 200
content-type: application/vnd.apple.mpegurl
server-timing: manifest-fetch;dur=120
x-cdn: cdn-B/hkg01
x-pop: cdn-B/hkg01
x-cache: HIT
x-multicdn-route: cdn-B
x-failover-mode: replay-origin-anchor-at-t-mid (re-issued)

#EXTM3U
#EXT-X-VERSION:9
#EXT-X-TARGETDURATION:6
#EXT-X-MEDIA-SEQUENCE:84228
#EXT-X-PROGRAM-DATE-TIME:2026-08-20T18:55:24.040Z
#EXTINF:6.000,
084228.ts

# cdn.pop.ladder_segment_index_delta:          fail→recover  (k dropping toward 0 within the same-length cohort window)
# cohort.ladder_rendition_reaggregate_probe:   pass  (240p / 540p / 720p / 1080p ladder-rendition re-aggregate fan-in pass
#                                                    after the segment_index backfill)
# cohort.cohort_join_stall_ratio:              0.024 (within cohort baseline)
# cohort.cohort_rebuffer_ratio:                0.013 (within cohort baseline)
# cohort.cohort_startup_time_p95_ms:           1920  (within startup-time baseline)
# cache.availability:                          pass  (X-Cache: HIT — cache leg still clean through the entire window)
# cache.segment_leg_cache_hit:                 pass  (segment leg still clean — NOT a segment-leg cache miss)
# multicdn.winner_RTT:                         pass  (cdn-a / cdn-b / cdn-c flat — multi-CDN health probe healthy)
# governed_action_emitted:                     reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill,
#                                              surface_cdn_failover_evidence_to_cdn_b_for_ladder_resync_via_origin_clip,
#                                              surface_cdn.pop.ladder_segment_index_delta_probe_to_cohort_fanin
ABR ladder smoking gun (cross-CDN timeline divergence on the mid-stream failover)
These are the four signals that, together, file the cdn_abr_ladder_drift hypothesis — with the cache leg + multi-CDN health probe ruling cdn_cache_segment_miss out by name.
  • cdn.pop.ladder_segment_index_delta k=6 (vs 0 baseline — mid-stream manifest timeline divergence, ~36.04 s PDAT drift on the apac/teal affected cohort across 4 renditions on cdn-A/sin02 AND cdn-B/hkg01)
  • cdn.pop.cross_cdn_alignment_probe fail (CDN-A live-edge tail at #EXT-X-MEDIA-SEQUENCE:84217 vs CDN-B replay-origin anchor at #EXT-X-MEDIA-SEQUENCE:84223)
  • cohort.ladder_rendition_reaggregate_probe fail across 240p / 540p / 720p / 1080p on the affected renditions
  • cohort.cohort_join_stall_ratio 0.124 vs 0.018 baseline (90 s cohort window — viewer PDAT drift past expected join window on the affected cohort)
  • cohort.cohort_rebuffer_ratio 0.082 vs 0.012 baseline (90 s cohort window)
  • cache.availability → pass; cache.segment_leg_cache_hit → pass; cdn.license_rollout_posture_check → pass; multicdn.winner_RTT → pass (40 / 41 / 42 ms flat on cdn-a / cdn-b / cdn-c). The cache leg is healthy at zero egress change — NOT a cache-miss, NOT a segment-leg hit-rate miss, NOT a multi-CDN routing flip.
Agent timeline response (ranked hypotheses + authorization-tiered remediation policy)
{
  "stream_id": "ckabrpackagelistdrift5189",
  "source": "https://cdn.example.com/live/event/stream.m3u8",
  "protocol": "HLS / CMAF / cdn-A/sin02 | cdn-B/hkg01 / multi-CDN cdn-a|cdn-b|cdn-c / ABR ladder / mid-stream failover",
  "checked_at": "2026-08-20T18:42:11Z",
  "ranked_hypotheses": [
    {
      "rank": 1,
      "hypothesis": "cdn_abr_ladder_drift",
      "tag": "dominant",
      "confidence": 0.81,
      "evidence_signals": [
        "cdn.pop.ladder_segment_index_delta → fail (k=6 vs 0 baseline across 4 renditions on cdn-A/sin02 AND cdn-B/hkg01)",
        "cdn.pop.cross_cdn_alignment_probe → fail — mid-stream manifest timeline divergence (CDN-A live-edge tail at #EXT-X-MEDIA-SEQUENCE:84217 vs CDN-B replay-origin anchor at #EXT-X-MEDIA-SEQUENCE:84223; #EXT-X-PROGRAM-DATE-TIME drift ~36.04 s)",
        "cohort.ladder_rendition_reaggregate_probe → fail (240p / 540p / 720p / 1080p rendition-level segment_index drift != 0 across the affected renditions on cdn-A/sin02 and cdn-B/hkg01)",
        "cohort.cohort_join_stall_ratio → fail (0.124 vs 0.018 baseline, 90s cohort window — late-join stalled behind the ladder mismatch; viewer PDAT drift past the expected join window)",
        "cohort.cohort_rebuffer_ratio → fail (0.082 vs 0.012 baseline — playback stalled on the affected cohort)",
        "cohort.cohort_startup_time_p95_ms → fail (6210ms vs 1840 baseline — TTFF past the 4 s threshold)",
        "edge.egress_kbps → pass (6421 Kbps at-or-above expected — cohort IS delivering traffic; the failure is on the ladder manifest, not on edge egress shape)",
        "cache.availability → pass on the affected edge POP — X-Cache: HIT on cdn-A/sin02 AND cdn-B/hkg01 under rebuffer load; no edge miss posture",
        "cache.segment_leg_cache_hit → pass on the affected cohort across both CDNs",
        "cdn.license_rollout_posture_check → pass on the affected ladder; the failure is NOT on the segment-leg cache-side",
        "multicdn.winner_RTT → pass on cdn-a|cdn-b|cdn-c (40ms / 41ms / 42ms — flat at baseline; no rtt-side failover stampede; the multi-CDN health probe proves the CDN leg is healthy)"
      ]
    },
    {
      "rank": 2,
      "hypothesis": "cdn_cache_segment_miss",
      "tag": "ruled_out_by",
      "confidence": 0.24,
      "evidence_signals": [
        "cache.availability probe reads pass on the affected edge POP across both CDNs (cdn-A/sin02 AND cdn-B/hkg01); X-Cache: HIT under rebuffer load — no edge miss posture; the cache leg is healthy",
        "cache.segment_leg_cache_hit probe reads pass on the affected cohort across both CDNs; if a segment-leg cache miss were stateful, the cohort resolve path across both CDNs and the same breadcrumb ladder would all misfire at the cache leg — they do not",
        "cdn.license_rollout_posture_check reads pass on the affected ladder; the ladder rendered correctly across all renditions — the failure is on ladder alignment, not on the cache-locale",
        "edge.egress_kbps reads at-or-above expected — the cohort is delivering traffic, the failure is on the ladder manifest timeline, not on segment availability",
        "if the cache-locale had a stateful miss, the warm breadcrumb cohort on cdn-A/sin02 would misfire on the same rung — it does not; the failure is on the cdn_ABR package-list drift between CDN-A live-edge tail AND CDN-B replay-origin anchor"
      ]
    },
    {
      "rank": 3,
      "hypothesis": "cdn_pop_route_bouncing_under_failover",
      "tag": "alternate",
      "confidence": 0.21,
      "evidence_signals": [
        "multicdn.winner_RTT probe reads pass on cdn-a | cdn-b | cdn-c (40ms / 41ms / 42ms — flat at baseline); no rtt-side failover stampede",
        "the multi-CDN health probe proves the CDN leg is healthy at zero egress change; the probe-side shape is on-segment-index delta, not on rtt_ms signature — the discriminator for the cdn_ABR package-list drift failure lane",
        "if rtt-side route bouncing were stateful, three AS populations on adjacent multicast-cdn neighborhoods would flip together under the failover — they did not; the failure is on the ladder manifest timeline alignment, not on rtt routing"
      ]
    }
  ],
  "governed_actions": [
    {
      "action":       "reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill",
      "type":         "governed",
      "decision_lane": "autonomous",
      "authorization_tier":   "Tier 0 (autonomous under gate)",
      "gating":       "top_confidence >= 0.80 AND cdn.pop.ladder_segment_index_delta >= 3 AND multicdn.winner_RTT.pass == true on the affected cohort",
      "evidence":     "top-hypothesis confidence 0.81; ladder_segment_index_delta k=6 >= 3; multicdn.winner_RTT pass (40/41/42ms flat)",
      "expected_effect": "ladder_segment_index_delta collapses toward 0 within the same-length cohort window; cohort.ladder_rendition_reaggregate_probe settles to pass; cohort_join_stall_ratio / cohort_rebuffer_ratio clears within tolerance"
    },
    {
      "action":       "surface_cdn_failover_evidence_to_cdn_b_for_ladder_resync_via_origin_clip",
      "type":         "governed",
      "decision_lane": "surfaced_to_humans",
      "authorization_tier":   "Tier 1 (surfaced to humans)",
      "gating":       "cdn_abr_ladder_drift >= 3 segments AND top_confidence >= 0.75",
      "evidence":     "ladder_segment_index_delta k=6 >= 3; top-hypothesis confidence 0.81 >= 0.75; the operator at the cdn-b side configures the ladder anchor and segment_index backfill before cdn-b's manifest tail fully publishes, so a misanchor that surfaces on a future primetime isn't masked by a 'we reconciled' close-out",
      "expected_effect": "cdn-B operator team configures the ladder anchor / segment_index backfill before the next apac/teal weekday primetime broadcast window opens"
    },
    {
      "action":       "surface_cdn.pop.ladder_segment_index_delta_probe_to_cohort_fanin",
      "type":         "governed",
      "decision_lane": "surfaced_to_humans_for_approval_forward",
      "authorization_tier":   "Tier 2 (operator-team sign-off forward)",
      "gating":       "incident_resolved AND cohort_side close-out signals cleared",
      "evidence":     "incident_resolved on T+41 m; cohort-side close-out signals (cohort_join_stall_ratio + cohort_rebuffer_ratio + cohort_startup_time_p95_ms + ladder_rendition_reaggregate_probe) cleared within tolerance",
      "expected_effect": "operator-team signs off on pinning the new probe into the cohort fan-in for future mid-stream failover windows; learn-loop codification"
    }
  ],
  "verification_window": {
    "close_out_signal": "cohort_side_plus_ladder_side",
    "probes_t_plus_30s": [
      "cdn.pop.ladder_segment_index_delta drops from k=6 → k=2 within the first same-length ladder window (30s cadence)"
    ],
    "probes_t_plus_2m": [
      "cohort.cohort_rebuffer_ratio clears <= 0.05 within ten consecutive 30s windows (NOT ladder-side settling)"
    ],
    "probes_t_plus_5m": [
      "cohort.cohort_join_stall_ratio settles <= 0.05 within ten consecutive 30s windows on the affected cohort (NOT ladder-side settling); mid-stream rejoin validates"
    ],
    "probes_t_plus_15m": [
      "cohort.ladder_rendition_reaggregate_probe pass on every monitored rendition (240p / 540p / 720p / 1080p)",
      "cdn.pop.ladder_segment_index_delta <= 0 within the same window",
      "cohort.cohort_startup_time_p95_ms ahead of the next weekday primetime broadcast window opens"
    ],
    "NOT_close_out_signal": [
      "cache.availability returning to X-Cache: HIT at zero egress change — geometric (cache leg was green the entire window)",
      "cache.segment_leg_cache_hit staying pass — geometric, the cache leg was green",
      "multicdn.winner_RTT reading pass on cdn-a|cdn-b|cdn-c — geometric (proves CDN leg is healthy, but does not on its own prove the cohort delivers green playback when the ladder is misanchored)"
    ],
    "explicit_note": "'ladder aligned' != 'cohort delivers green playback'. Recovery is verified cohort-side + ladder-side."
  },
  "surfaced_to_humans": [
    {"owner": "cdn-B operator team",   "task": "configure the ladder anchor / segment_index backfill on the cdn-B side before the next apac/teal weekday primetime broadcast window opens"},
    {"owner": "reliability team",      "task": "approve the cdn.pop.ladder_segment_index_delta probe + the cohort.ladder_rendition_reaggregate_probe codification into the cohort fan-in"},
    {"owner": "viewer-platform team",  "task": "approve the apac/teal cohort-side profile post-mortem from this entry — ladder-side settle + cohort-side settle cadence pinned from the timeline"},
    {"owner": "on-call",               "task": "page on the cdn_ABR package-list drift root-cause review — mid-stream failover + live-edge tail + replay-origin anchor"}
  ]
}
What the agent emitted (authorization-tiered)
The governed actions ship split across three authorization tiers: Tier 0 autonomous under an explicit gate; Tier 1 surfaced because cdn-B side coordination required; Tier 2 surfaced for operator-team sign-off forward on the new probe.
  • Tier 0 reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill (autonomous under gate: top_confidence ≥ 0.80 + ladder_segment_index_delta ≥ 3 + multicdn.winner_RTT.pass)
  • Tier 1 surface_cdn_failover_evidence_to_cdn_b_for_ladder_resync_via_origin_clip (surfaced: cdn_abr_ladder_drift ≥ 3 segments + top_confidence ≥ 0.75 + cdn-B operator configures the ladder anchor / segment_index backfill)
  • Tier 2 surface_cdn.pop.ladder_segment_index_delta_probe_to_cohort_fanin (surfaced for operator-team approval forward: incident_resolved + cohort-side close-out signals cleared; the learn-loop act)
  • surfaced → paged the cdn-B operator team + the reliability team on the cdn_ABR package-list drift root-cause review
  • close-out signal → cohort-side + ladder-side: cohort_join_stall_ratio, cohort_rebuffer_ratio, cohort.ladder_rendition_reaggregate_probe, cdn.pop.ladder_segment_index_delta within tolerance over the next same-length cohort window — NOT cache.availability / cache.segment_leg_cache_hit / multicdn.winner_RTT returning to pass
Timing

Detection, classify, mitigate, recover (simulated telemetry)

Four timing windows on the postmortem timeline, each read off the cohort probe cadence. The figures are simulated telemetry — the disclosure near the top of this page applies to every figure on this list. Note that the close-out window is verified cohort-side + ladder-side on the apac/teal cohort for the affected window (NOT cache-side, which was green the entire window).

Detection

~5 s

cdn.pop.ladder_segment_index_delta crossed 0→6 within a 30 s window on the apac/teal affected cohort across both CDNs; the agent surfaced the detector from the cross-CDN alignment probe at T+5 s on the cohort.

Time to classify

~1 m

cdn_abr_ladder_drift · dominant ranked at 0.81 confidence with three ranked hypotheses at T+1 m — discriminator is cache.availability + cache.segment_leg_cache_hit + license-rollout posture reading pass on the affected cohort.

Time to mitigate

~7 m

Tier 0 reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill queued at T+2 m; cdn-B operator-side ladder resync configured by T+10 m; cohort re-anchored by T+18 m.

Time to full recovery

~22 m

cohort.cohort_join_stall_ratio + cohort.cohort_rebuffer_ratio + cohort.ladder_rendition_reaggregate_probe + cdn.pop.ladder_segment_index_delta cleared within tolerance at T+25 m — cohort-side + ladder-side close-out, NOT cache-side re-anchoring.

Detect → Classify → Governed Fix

The loop on this incident

Three steps close the lane on an apac/teal mid-stream CDN failover ABR package-list drift incident. The fix is split explicitly into the authorization-tiered governed-action branches — Tier 0 acts on its own under the replay-origin reanchor gate, Tier 1 surfaces to the cdn-B operator team, Tier 2 surfaces for operator-team sign-off forward on the new probe.

01 · Detect
Cross-CDN ladder-mismatch probe fan-in (lockstep)
Cross-CDN ladder alignment probe, ladder-segment-index delta probe, cohort-side ladder-rendition re-aggregate probe, and the cache leg probes (availability, segment-leg cache-hit, license-rollout posture) running lockstep at 30 s cadence so the discriminator lands.

The probe set fans in across the apac/teal affected cohort on both CDNs. cdn.pop.ladder_segment_index_delta reads k=6 (vs 0 baseline); cdn.pop.cross_cdn_alignment_probe reads fail (CDN-A live-edge tail at #EXT-X-MEDIA-SEQUENCE:84217 vs CDN-B replay-origin anchor at #EXT-X-MEDIA-SEQUENCE:84223 — ~36.04 s PDAT drift); cohort.ladder_rendition_reaggregate_probe reads fail across 240p / 540p / 720p / 1080p; the cache leg probes read pass on both CDNs; multicdn.winner_RTT reads 40 / 41 / 42 ms on cdn-a / cdn-b / cdn-c — all three CDNs flat at baseline.

02 · Classify
cdn_abr_ladder_drift · dominant @ 0.81
Ranked with cdn_cache_segment_miss · ruled out by (0.24) and cdn_pop_route_bouncing_under_failover (0.21) alternate.

The discriminator against the cache-side lane is the cache.availability + segment-leg cache-hit + license-rollout posture pattern — the cache leg reads pass on the affected cohort across both CDNs, and the multi-CDN health probe proves the CDN leg is healthy at zero egress change; the failure is on the cross-CDN ladder manifest timeline alignment, not cache-locale-shaped and not rtt-side-shaped. The discriminator against the rtt-side lane is the multi-CDN health probe — flat at baseline on cdn-a / cdn-b / cdn-c.

03 · Governed Fix
Tier 0 ladder reissue (autonomous) · Tier 1 cdn-B resync · Tier 2 codify probe
Three authorization-tiered branches — Tier 0 autonomous under an explicit gate; Tier 1 surfaces to the cdn-B operator team; Tier 2 surfaces for operator-team sign-off forward on the new probe.

The Tier 0 autonomous branch reissues the ABR ladder via replay-origin reanchor at T_mid plus a CDN-side segment_index backfill — clears the ladder probe within T+30 s and the cohort side within T+18 m. The Tier 1 surface branch configures the cdn-B side ladder anchor / segment_index backfill before cdn-B's manifest tail fully publishes. The Tier 2 forward branch codifies the new probe (cdn.pop.ladder_segment_index_delta + cohort.ladder_ rendition_reaggregate_probe) into the cohort probe fan-in for future mid- stream failover windows. Recovery is verified cohort-side + ladder-side across T+30 s → T+15 m on the apac/teal affected cohort.

Recommended fix (Governed)

Reissue the ABR ladder — or sync the cdn-B anchor.

On an apac/teal mid-stream CDN failover ABR package-list drift incident, the governed fix is a two-arm branch — an agent-side reissue arm that clears the cross-CDN ladder with replay-origin reanchor (Tier 0 autonomous under gate); a cdn-B operator-side ladder resync arm that configures the cdn-B side anchor + segment_index backfill (Tier 1 surfaced to humans); and a forward branch that codifies the new probe into the cohort probe fan-in (Tier 2 surfaced for operator-team sign-off forward). The three arms close the lane in the same incident window and forward.

01Agent — reissue the ABR ladder (Tier 0 autonomous)
Tier 0 · agent-emitted

Emit reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill — replay-origin reanchors the affected ladder at T_mid and the cdn-side segment_index backfill resets the affected renditions' index against the anchor. The fix is Tier 0 autonomous under gate (top_confidence >= 0.80 AND cdn.pop.ladder_segment_index_delta >= 3 AND multicdn.winner_RTT.pass == true on the affected cohort) and surfaced for operator-team sign-off forward outside that gate so the operator team can review the false-positive rate before the reissue lands.

Verification (cohort-side + ladder-side)

cdn.pop.ladder_segment_index_delta drops from k=6 → k=2 within the first same-length ladder window (T+30 s staged gate); cohort.cohort_rebuffer_ratio settles ≤ 0.05 within ten consecutive 30 s windows (T+2 m staged gate); cohort.cohort_join_stall_ratio settles ≤ 0.05 within ten consecutive 30 s windows (T+5 m staged gate); cohort.ladder_rendition_reaggregate_probe reads pass on every monitored rendition; cohort.cohort_startup_time_p95_ms reads ahead of the next apac/teal weekday primetime broadcast window opens.

NOT a close-out signal: cache.availability returning to X-Cache: HIT at zero egress change, cache.segment_leg_cache_hit staying pass, or multicdn.winner_RTT reading pass on cdn-a / cdn-b / cdn-c. The cache leg was green the entire window; reading the cache-side probes as the close-out signal misses the lane forward.

02cdn-B operator — sync the ladder anchor (Tier 1 surfaced)
Tier 1 · surfaced to humans

Configure the cdn-B side ladder anchor + segment_index backfill before cdn-B's manifest tail fully publishes — the operator team at the cdn-B side configures the ladder anchor / segment_index backfill on the affected cohort so a misanchor that surfaces on a future primetime isn't masked by a "we reconciled" close-out. The fix is surfaced to humans because it requires CDN-side coordination the agent cannot perform — the multi-CDN health probe proves the CDN leg is healthy at zero egress change, but the cdn-B side ladder anchor must be reconfigured by the cdn-B operator team.

Verification (ladder-side + cohort-side)

cdn.pop.ladder_segment_index_delta settles ≤ 0 within the same window after the Tier 0 reissue + Tier 1 cdn-B resync configuration (T+15 m staged gate); cohort.ladder_rendition_reaggregate_probe reads pass on every monitored rendition (240p / 540p / 720p / 1080p); the cdn-B side ladder anchor ship configured for the next apac/teal weekday primetime broadcast window so a re-incident on the affected scaffold lands already absorbed.

NOT a close-out signal: cache.availability / cache.segment_leg_cache_hit staying pass at zero egress change — those probes are geometric and do not prove the cohort delivers green playback when the ladder is misanchored.

Caveat — recovery is cohort-side + ladder-side, not cache-side
'CDN-side cache healthy' ≠ 'cohort delivers green playback'
The cache leg returning green (or staying green) is geometric close-out, not cohort-side close-out. The actual close-out signal is the cross-CDN ladder alignment probe plus the cohort's join_stall_ratio + rebuffer_ratio + ladder-rendition re-aggregate probe clearing tolerance over the next same-length cohort window.

A fix that normalizes cache.availability + cache.segment_leg_cache_hit + multicdn.winner_RTT without clearing the affected cohort's join stall, cohort rebuffer, or cross- CDN ladder alignment probe within tolerance is a fix that didn't reach the cohort. The audit step on this incident writes the close-out signal into the playbook as "verify cdn.pop.ladder_segment_index_delta + cohort.ladder_rendition_ reaggregate_probe + cohort_join_stall_ratio + cohort_rebuffer_ratio settle within tolerance over the next same-length cohort window" — not "verify the cache leg returning to green / verify cache.availability + segment-leg cache-hit returning to pass".

Learn-loop note

Codifying the new probe into the cohort fan-in

The Tier 2 surface for operator-team sign-off forward on this incident codifies two probes into the cohort probe fan-in — cdn.pop.ladder_segment_index_delta pinned into the cohort fan-in for the next weekday primetime broadcast window (operator-approved; surfaces on every mid-stream failover cohort); cohort.ladder_rendition_reaggregate_probe pinned for the next four weekday primetime windows so the false-positive rate can be measured before promotion.

Learn-loop · audit step on this incident
Tier 2 · operator-team sign-off forward
'Verify ladder-segment-index delta + cohort ladder-rendition re-aggregate settle within tolerance over the next same-length cohort window'
NOT 'verify cache.availability returning to green at zero egress change'; NOT 'verify cache.segment_leg_cache_hit returning to pass'; NOT 'verify multicdn.winner_RTT reading pass on cdn-a|cdn-b|cdn-c'; NOT 'verify the CDN-cache re-anchor'.

The audit step codifies that cdn.pop.ladder_segment_index_delta +cohort.ladder_rendition_reaggregate_probe settle within tolerance over the next same-length cohort window — meaning the cross-CDN ladder probe fan-in tracks across every mid-stream failover window, the operator-team sign-off goes on the new probe codification, and the cohort fan-in carries the ladder-mismatch signal lane forward. The cache.availability / segment-leg cache-hit / multi-CDN health probe readings stay pass through the entire window — those are geometric, NOT cohort-side close-out signals.

Next step

Want Streamwake to disambiguate ABR ladder drift on your cohort?

Sign up, register a multi-CDN probe, and the same cdn.pop.ladder_segment_index_delta · cdn.pop.cross_cdn_alignment_probe · cohort.ladder_rendition_reaggregate_probe · cohort.cohort_join_stall_ratio probes that produced the timeline above run on every prime-cohort refresh — and surface in a Slack channel, a webhook, or the streams dashboard.

Synthetic Incident — This scenario uses simulated telemetry constructed from documented streaming behaviors. It does not represent a Streamwake customer outage.

Need Streamwake on one of your incidents?
Would you like Streamwake to analyze one of your historical incidents and show where AI could reduce investigation time? (Filed under: ABR ladder drift on mid-stream CDN failover.)
Incident analysis
  • Pick a recent on-call incident — manifest stall, edge miss, player-side stall, or peer congestion.
  • We replay it through the same reliability-agent probe cascade used on the postmortem above.
  • You walk away with a written what-could-have-been-Automated readout, not a sales deck.
Read the next

Related writeups

The closest siblings cover the player-visible symptom lands: the broader ISP-vs-CDN triage guide at /troubleshooting/isp-congestion-vs-cdn-failure (the eight-symptom read / ISP-vs-CDN discriminator), the ISP-vs-CDN recovery criteria entry (cdn_failure · ruled out by name + cohort-side + last-mile close-out), the origin-shield queue saturation write-up (the shield queue p99 waiting past the cohort's expected join-window profile), the manifest fetch timeout storm at a regional edge POP (a regional edge POP returning manifest-timeouts above baseline during a quiet pre-peak window), and the live-event scale-out buffering postmortem (a marquee broadcast with a viewer-spike that pushes the cohort beyond the pre-provisioned capacity envelope). Together they cover the four failure-mode lanes Streamwake reliability agents are tuned for alongside ABR ladder drift on mid-stream CDN failover.