Fasil PM logoFasil PMProject & IT ConsultancyBook a call

The Log Did Not Fix the Outage. It Found It.

A Severity-1 latency incident three weeks before certification. Detection took four minutes. The change control log found the cause in nine.

Programme
Project Falcon
Organisation
Atlas Bank
Phase
Control
Template
Change Control Log
Post 19 of 22 following Project Falcon at Atlas Bank. Full program context is in Post 01. Posts 17 and 18 stepped back to show the quarter that re-baselined the money and the night steering moved to pause the program. This post returns to forward chronology and tells the story both of those posts kept deferring: the 72-hour incident of January 2026, and the unglamorous artifact that ended it.

The Pager Does Not Care That You Are Three Weeks From Certification

At 02:47 on a Monday, the on-call phone for the txn-screen group lights up, and the alert is the one nobody wanted: latency on the screening path climbing through 180 milliseconds against a 200 millisecond regulated SLA, and still rising. The service is in shadow mode against live Market A volume, so no customer is harmed yet. But the Market A onboarding peak begins at 08:00, the same Monday peak that materialized R04 in December, and the math is brutal: at the current climb, the path breaches its SLA under load in roughly five hours, on the exact certification-path control the central bank audits in eight weeks. This is not a customer incident. It is something more dangerous to the program: a certification-evidence incident, live, with the audit booked.

The alert exists because of December. The trigger Priya Raman's team wired into Dynatrace after R04 (Post 15) fires on the screening path's p99, and the register-wide instrumentation audit of 9 January had wired this specific alert eleven days before it was needed. Detection latency: four minutes from first breach of the warning threshold to pager. The fined bank in Post 18 found its incident on a customer complaint. Falcon found this one before sunrise.

The on-call SRE declares a Severity-1 at 03:04 and opens the PMO-CC-003 incident structure, the template designed back in Post 07 for exactly the event that had not happened until now. PMO-CC-003 does three things automatically: it stands up a single incident commander role (not the most senior person, the designated one), it opens a unified change control log that every team writes to, and it sets a communications cadence that protects the responders from the responders' own management. By 03:20, Jin-ho Park is incident commander, the war room is virtual and staffed, and the change log has its first entries.

And immediately, the incident becomes a whodunit, because three teams shipped change in the preceding 72 hours. Park's squad deployed a gRPC contract update to two services on Friday. The platform team rotated a set of Kubernetes node pools on Saturday. The vendor, under its own change calendar and its own contractual autonomy, pushed an update to its five contracted services on Sunday evening. Each team knows about its own change. No single person, at 03:30 on Monday, knows about all three at once, except that one artifact does, because all three were required to log to it.

Everyone remembers their own deploy. The log is the only thing that remembers everyone's.

Three Teams, Three Confessions, One Timeline

The first hour of a Severity-1 has a characteristic failure mode: everyone investigates their own most recent change, because it is the change they understand, and the actual cause hides in the seam between teams that no one owns. Falcon nearly falls into it. Park's instinct is that Friday's contract update is the culprit; he starts a rollback. The platform lead suspects Saturday's node rotation and begins draining pools. Both are about to spend two hours disproving their own innocence while the clock runs toward 08:00.

The incident commander discipline stops it. Park, in the commander role rather than the engineer role, makes the call that the post turns on: before anyone rolls back anything, the log gets read end to end, all changes, all teams, against the symptom-onset timestamp. The reading takes nine minutes and ends the investigation.

TimestampChange / EventTeamType Fri 17 Jan 16:20gRPC contract v1.4 deployed: api-orchestrator, txn-routerFalcon buildInternal Sat 18 Jan 11:05Kubernetes node pool rotation, screening namespace (3 pools)PlatformInternal Sun 19 Jan 21:30Vendor config push: shared rate-limiter policy updated across 5 contracted services (global limit lowered 40%)VendorVendor Mon 19 Jan 00:10Overnight batch onboarding load begins ramp (routine)—Event Mon 19 Jan 02:43Screening-path p99 first crosses 150ms warning threshold—Onset Mon 19 Jan 02:47Dynatrace alert fires; on-call pagedtxn-screenDetect Mon 19 Jan 03:04Severity-1 declared; PMO-CC-003 openedPMODetect

The timeline does the diagnosis that three teams arguing could not. The two internal changes landed Friday and Saturday; the path ran clean through both, including all of Sunday daytime. The symptom onset is 02:43 Monday. The only change between clean operation and onset is the vendor's Sunday-evening rate-limiter update, which lowered a shared throttle by 40 percent across the five contracted services, throttling exactly the calls the gRPC mesh makes into those services under the overnight load ramp. Friday's and Saturday's changes are not suspects; they are alibis, because the path worked after them. Park stops his rollback. The platform lead stops draining pools. Two hours of disproving innocence, saved by nine minutes of reading.

Jin-ho Park (incident commander): “Rollbacks halted. Nothing internal correlates. The only change on the timeline that fits the onset is the vendor rate-limiter at 21:30 Sunday. Dmitri, I need your team in this channel now, and I need that config reverted or whitelisted for our service accounts inside the hour.”

Dmitri Volkov joins the war room at 04:10 and does something that, after five posts of contractual caution, lands differently: he does not argue. The log is the log. His team's change is on it, timestamped, against an onset his own monitoring confirms. Contractual people are not obstructive people; they are people who respect documents, and the document is unambiguous. The argument that would have consumed the next three hours never happens, because there is nothing to argue about. The vendor begins remediation at 04:25.

The Sponsor Wants Hourly Updates. The Commander Needs Quiet.

By 06:00, the cause is known, remediation is underway, and the incident enters its most politically fraught phase: management wakes up. Fatima Idris, briefed at 05:30 per protocol, makes a request that is entirely reasonable and exactly wrong.

Fatima Idris: “The board chair is already nervous after the sector news. I want hourly updates from the war room, and I want Park briefing the executive group directly every hour until this closes. Visibility reassures.”

Visibility reassures the watchers and destroys the watched. An incident commander pulled out of the war room every sixty minutes to brief executives is an incident commander not commanding for the worst hour of the event. This is the precise failure PMO-CC-003 was designed to prevent, and the design is the defense. Fasil Alemeye Abate does not refuse Idris; he routes her.

Fasil Alemeye Abate: “The protocol already answers this, and it answers it in your favor. Park commands; he does not brief. I am the incident communications lead, that is the PMO-CC-003 split, and you will get a written situation report every ninety minutes from me, board-ready, that you can forward to the chair without translation. The executive group gets one scheduled brief at 09:00 from me, not an hourly interruption of the person holding the fix. If you pull Park into a briefing at 07:30, you are trading the resolution for the reassurance, and the reassurance does not survive a missed SLA at 08:00.”

Idris accepts the 90-minute SITREP cadence, because it gives her something better than hourly verbal updates: a written, forwardable, board-grade record she does not have to summarize under pressure. The split holds. Park stays in the war room. The first SITREP goes out at 06:15, the Market C regulator is notified at 07:00 through the standing PMO-CC-002 channel (24 hours ahead of any requirement, which is the move that mattered most in Post 18's defense), and the 08:00 peak arrives.

Screening-Path p99 · The 72-Hour Window
Latency against the 200ms SLA from onset to closure. The vendor revert lands before the 08:00 peak; the remaining 70 hours are hardening, monitoring, and the structured close, not firefighting.
220ms 200ms 130ms 100ms Mon 03:00 Mon 08:00 Mon 18:00 Wed Thu close 200ms regulated SLA 130ms normal vendor revert 07:05 peak held: 188ms Sev-1 closed Thu

The peak held at 188 milliseconds: inside the SLA, with twelve milliseconds to spare, because the revert landed at 07:05 and the path had recovered to baseline before the load arrived. No SLA breach was ever recorded. The remaining 70 hours of the 72 were not crisis; they were the structured close that separates a program from a lucky team.

The outage lasted four hours. The discipline lasted seventy-two. Only one of those is why the program survived February.

Drafting the Post-Incident Review with STAR

The incident closed Thursday. The post-incident review is what turns 72 hours of adrenaline into an artifact a regulator, a board, and a successor can all read, and it had to be written while the responders still remembered the timestamps and before they started remembering them flatteringly. The framework is STAR, Situation, Task, Action, Result: standard in structured interviewing and incident retrospectives alike, well-documented, no adoption caveat required. STAR suits a post-incident review because it enforces the one sequence incident write-ups most often corrupt: Situation and Task before Action, so the document establishes what was true and what was required before it narrates what anyone did, which is the order that keeps blame out and causation in. The Result section then carries both the resolution and the honest residue: what the program got right, and what the log revealed about a gap nobody had closed.

Prompt · Post-Incident Review · 23 January 2026
You are drafting a post-incident review (PIR) for a Severity-1 incident on a regulated banking program, using STAR (Situation, Task, Action, Result). The PIR will be read by the program board, the independent assurance reviewer, and is evidence for a regulatory certification audit in March. Tone: factual, blameless, causally precise. Name systems and changes; do not assign personal fault. SITUATION: txn-screen (real-time transaction screening, 200ms regulated SLA) in shadow pilot against live Market A volume. 19 Jan 2026, 02:43, screening-path p99 began climbing through a 150ms warning threshold toward the SLA, hours before the 08:00 onboarding peak. Three changes in preceding 72h: gRPC contract v1.4 (Fri, internal); K8s node-pool rotation (Sat, internal); vendor shared rate-limiter policy lowered 40% across 5 contracted services (Sun 21:30, vendor). TASK: prevent an SLA breach on a certification-path control before the 08:00 peak; identify cause without halting on the wrong change; preserve evidence for certification. ACTION: Sev-1 declared 03:04, PMO-CC-003 opened, single incident commander, unified change log. Log read end-to-end against onset timestamp (02:43): internal changes preceded clean operation through Sunday; only the vendor 21:30 change correlated with onset. Internal rollbacks halted; vendor engaged 04:10, rate-limiter reverted/whitelisted for Falcon service accounts 07:05. Comms: PMO-CC-003 commander/communications split; 90-min SITREPs; Market C regulator notified 07:00 via PMO-CC-002. RESULT: peak held at 188ms p99, no SLA breach recorded. Sev-1 closed Thu 22 Jan. Root cause: vendor change to a SHARED rate-limiter affecting the gRPC mesh, deployed under the vendor's autonomous change calendar with no cross-party change notification to the program. TASK FOR YOU: draft the PIR with STAR sections, a root-cause statement, and a corrective-actions table (owners, dates). Then identify the single systemic gap this incident exposes that is NOT "the vendor made a bad change," and flag it for human decision.

The Gap Behind the Cause

The STAR draft was clean and the root-cause statement accurate. The instruction that mattered was the last one: find the systemic gap that is not the vendor's bad change. The machine found it, and it indicted the program, not the vendor.

AI Draft · PIR · Result Section (Extract) and Systemic-Gap Flag

RESULT. The incident was detected in 4 minutes, diagnosed via the unified change log in approximately 9 minutes of timeline analysis, and resolved before the 08:00 peak with no SLA breach recorded. Detection, command structure, cross-team logging, regulator notification, and communications discipline all functioned as designed. The screening path returned to its 130ms baseline by 07:20 and held through the peak at 188ms p99.

AI flagged (systemic gap): “The proximate cause is the vendor change. The systemic gap is that the program had no cross-party change freeze or change-visibility requirement covering the vendor's autonomous change calendar, despite the vendor's five services sharing infrastructure (the rate-limiter) with the program's certification-path control. The unified change log captured the vendor change after the fact, which is why diagnosis was fast; but nothing required the vendor to notify the program before deploying a shared-resource change, and nothing imposed a change freeze on the certification path during the shadow-evidence window. The program was protected by good detection, not by good prevention. The same class of incident recurs on go-live night unless cross-party change governance is established. This is a decision for the program, not a vendor defect: the vendor operated within its contract. The contract is the gap.”

“The program was protected by good detection, not by good prevention” is the sentence that turned a clean incident into an open action. A lesser review would have closed on “vendor reverted, SLA held, well done.” The flag refused the victory lap and named the uncomfortable structural truth: the change log made the program fast, but nothing made the program safe, and the difference would matter enormously on go-live night when the same shared rate-limiter sits under live customer traffic instead of shadow. The vendor did nothing wrong; it exercised a contractual right over its own change calendar. The defect was that the program never negotiated visibility into changes to a resource its regulated control depended on. That becomes corrective action CA-3, and it routes straight into the go-live transition plan of Post 20.

INC-001 Post-Incident Review, as Filed

Post-Incident ReviewINC-001 · v1.0
Severity-1: Screening-Path Latency Degradation (Shared Rate-Limiter)
Project Falcon · Atlas Bank · Incident Commander J. Park · PIR by Fasil Alemeye Abate
Status Resolved Declared 19 Jan 03:04 · Resolved 22 Jan · Duration 72h (impact ~4h) · SLA breach none

Situation and Task

txn-screen (200ms regulated SLA) in shadow pilot against live Market A volume. At 02:43 on 19 Jan, screening-path p99 began climbing toward the SLA ahead of the 08:00 onboarding peak. Task: prevent breach on a certification-path control, diagnose without halting on the wrong change, preserve certification evidence. Link: txn-screen SLA (CR-001 v1.1); RR-R04 alert instrumentation.

Action and Timeline

Detection02:47 Dynatrace alert (4 min from onset); Sev-1 declared 03:04; PMO-CC-003 opened, IC assigned 03:20
DiagnosisUnified change log read end-to-end vs 02:43 onset; internal Fri/Sat changes excluded (clean operation followed); vendor 21:30 Sun rate-limiter change isolated as sole correlate
ResolutionInternal rollbacks halted; vendor engaged 04:10; rate-limiter reverted and Falcon service accounts whitelisted 07:05; baseline restored 07:20; peak held 188ms
CommunicationsIC/comms split held; 90-min SITREPs to sponsor; Market C regulator notified 07:00 via PMO-CC-002 (24h+ ahead of obligation); executive brief 09:00

Root Cause

Vendor lowered a shared rate-limiter global threshold by 40% across its five contracted services under its autonomous change calendar (Sun 21:30). The gRPC mesh's calls into those services were throttled under the overnight load ramp, degrading screening-path latency. Vendor operated within contract ATL-PI-2024-11; no cross-party notification of shared-resource changes was required. Proximate cause: vendor change. Systemic cause: absent cross-party change governance over a shared resource underpinning a regulated control.

Corrective Actions

CA-1Falcon service accounts permanently exempted from the shared rate-limiter; dedicated limiter provisioned for the screening path. Done 24 Jan. Owner: J. Park
CA-2Screening-path p95 latency alert added below the existing p99 (earlier warning). Live 26 Jan. Owner: P. Raman
CA-3Cross-party change governance: vendor change notification (48h) for shared-resource changes + change freeze on the certification path during evidence windows and the go-live window. Negotiated via contract amendment; routed into the Post 20 transition plan. Target 15 Mar. Owner: L. Marquez
CA-4INC-001 added to the March certification evidence file as a detection-to-closure demonstration. Done 23 Jan. Owner: A. Okonkwo

What the Human Changed

What the Human Changed (AI Draft to Filed INC-001)
  1. Promoted the systemic-gap flag to corrective action CA-3. The draft logged the cross-party change gap as an observation in the Result section. Observations close with the incident; corrective actions outlive it. Fasil made it CA-3 with an owner (Marquez), a contract-amendment route, and a 15 March target, because the same shared rate-limiter sits under live traffic on go-live night.
  2. Kept the duration honest at 72 hours, not 4. The draft was tempted to headline the four-hour impact window. The incident was open for 72 hours through monitoring, hardening, and structured close, and the PIR says 72 with the impact window noted separately. A program that reports the flattering number teaches itself to stop early.
  3. Named the vendor's innocence explicitly. The draft's root cause stopped at “vendor change.” Fasil added the clause that the vendor operated within contract and the defect was the program's missing governance. Blaming the vendor would have felt good and prevented nothing; naming the contractual gap is what produced CA-3.
  4. Added the p95 alert below the p99. The existing alert fired at p99 against a 150ms threshold, giving four minutes. The incident review showed that an earlier p95 signal would have given closer to fifteen. CA-2 adds it: the program that just survived on four minutes of warning bought itself more, because next time the revert might not be a phone call away.
  5. Routed INC-001 into the certification file the same day. The draft treated the PIR as an internal close-out. Fasil filed it as CA-4 into the March evidence pack within 24 hours. A regulator auditing incident management is best answered with an incident that was detected in minutes, diagnosed by a documented log, and closed without a breach, which is exactly the story Post 18's defense and Post 16's readiness board both drew on.
The Future Payoff

This incident is the single most reused artifact in the back half of the series. Its evidence pack is exhibit one in Post 18's cancellation defense (detected in minutes, regulator briefed in 24 hours, no breach). It is a named line in Post 16's readiness grid, the reason the operations dimension could be argued at all. CA-3's cross-party change freeze becomes a load-bearing clause in Post 20's go-live transition plan, the night the same shared rate-limiter finally sits under real customer traffic. And the whole episode anchors the lessons register (Post 21) under the heading the flag wrote for it: protected by detection, not by prevention, until we fixed the contract.

The Severity-1 closed on Thursday 22 January with no SLA breach, a four-action plan, and a regulator who had heard about it from the program before hearing about it from anyone else. Eleven days later, steering would try to pause the program partly because this incident happened. The defense that survived that night was written in this room, in a change log nobody enjoyed maintaining. The series now moves to the day all of it was for: go-live, 14 July, when the rehearsals stop and the customers arrive. That is Post 20.

The Takeaway
In an outage, the most valuable record is the one that proves what you did not break.
A change log feels like overhead every single day it is maintained and like salvation for the one hour it is needed. When three teams each have a recent deploy and the clock is running toward a regulated peak, the instinct is for everyone to investigate their own change, and that instinct sends two of the three teams to disprove their innocence while the real cause runs free in the seam between them. The unified log collapses that hour into nine minutes, not by being clever but by being complete: one timeline, every team's change, against the moment the symptom began. It exonerated the internal work, isolated the shared-resource change, and ended the argument before it started, because there is nothing to argue with a timestamp. But the log's deeper gift was the flag it made possible afterward: the honest finding that the program had been fast, not safe, and that good detection had quietly substituted for the cross-party prevention nobody had negotiated. Maintain the log. Then read what it tells you about the gap you have not closed yet.

A fictional case study for teaching purposes. Atlas Bank, Project Falcon and all named individuals are invented. Technologies are industry-standard and publicly available.