The Dashboard Said Fine. The Average Was Doing the Lying.
A dashboard averaging away a live problem, the risk it hid for weeks, and what it cost to wire the register's trigger conditions into real alerts.
Month 11. The Program Is Boring. That Should Have Been the Tell.
By early December, Falcon is in the most dangerous condition a program can occupy: quietly on track. The latency evidence plan landed with the PMO on 28 November, on time, built on ADR-002's 70ms headroom. The gRPC adoption is proceeding inside its boundary; eight services are on protobuf contracts, zero migration-only work has leaked into the schedule, exactly as written. Phase 3 starts in three weeks. The certification audit is in March. Steering in December will be the shortest of the year.
And since 17 November, one piece of Falcon has been touching reality. The async-KYC service, kyc-verify, has been running in shadow mode against live Market A onboarding volume: real applications, real documents, real verification calls, with its verdicts logged but not acted on. The shadow pilot exists to gather certification evidence. It produces telemetry into Splunk and Dynatrace, and that telemetry feeds, among other things, the monthly risk review prep that the PMO runs with AI assistance: register entries on one side, three weeks of operational data on the other, with a standing instruction to surface anything in the data that looks like a register entry waking up.
Risk R04 has been on the register since April: async-KYC performance degradation under sustained load, owner Priya Raman, trigger condition written as kyc-verify queue depth > 50K sustained > 4 hours. It has sat there for eight months, probability Medium, doing nothing. The internal audit in Post 09 scored the program's controls 4.7 out of 5, and the 0.3 it withheld was a single finding: R04's performance testing stopped at twice average load and never exercised a sustained peak. The program logged the finding, scheduled the deeper testing for Phase 3, and moved on. The gap was known, dated, and parked.
On 5 December, the AI prep run reads three weeks of shadow telemetry and stops parking it.
The pattern is invisible at the resolution the dashboards use. The kyc-verify operations view reports daily mean queue depth: 11K messages, comfortably nominal, green for nineteen consecutive days. The AI's prep run does not read the day; it reads the hours, and the hours tell a different story. Every Monday, Market A's onboarding peak pushes queue depth past the 50K trigger by mid-morning, holds it above the line for nine hours or more, and the queue does not finish draining until close to 23:00. The register's materialization condition, the exact sentence Raman wrote in April, has been true every Monday since the pilot started. Nobody saw it, because the average ate it.
Raman Does Not Believe the Machine
The flag reaches Fasil Alemeye Abate on the morning of 8 December. He takes it to the risk owner before he takes it anywhere else, and the risk owner pushes back hard.
Raman's skepticism is not a flaw in the story; it is the control working. An AI flag is a hypothesis, not a finding, and the program's standing rule since Post 09 is that machine-surfaced signals get verified by a human against raw data before they touch the register. Fasil does not argue the pattern. He asks Raman to re-cut her own telemetry one way: drop the daily mean, plot hourly depth for the last three Mondays against her own trigger line. The re-cut takes forty minutes.
Raman looks at her own data at the resolution the AI used and concedes in one sentence, which is one more sentence than most technical leads manage:
The 0.3 audit gap from Post 09 turns out to have been a map. The finding said: your testing never exercised sustained peak load. The materialization arrived precisely inside the untested region. Audits do not predict the future; they predict where you are blind, which is the half of the future that hurts.
The Fix Is Easy. The Interim Is Political.
The permanent fix is genuinely unexciting: scale the consumers. Raman's plan adds queue-depth-driven autoscaling to the kyc-verify consumer pool with backpressure signaling, priced at $52K including the alert instrumentation, deployable by 29 December. Within the PM's single-item authority; Ahmed Hassan logs it into December's discretionary aggregate. The problem is the three weeks between now and then, which contain two more Monday peaks, a year-end onboarding push, and a shadow pilot whose entire purpose is generating clean certification evidence. Nine-hour trigger breaches in the certification evidence are not clean.
The interim option on the table is an intake throttle: cap shadow-pilot onboarding intake during the Monday peak window so the queue stays under the trigger until the scaling lands. Operationally trivial. Politically, it walks straight into Samuel Osei, Head of Retail Banking and owner of the Revenue benefit category from Post 08.
Fasil puts the objection in the document, and then does something with it: the throttle goes in time-boxed, with a written expiry of 5 January 2026, a benefit-impact note co-drafted with Osei confirming zero effect on the Revenue category baseline (shadow verdicts are not acted on; no customer experiences the throttle), and a commitment that any throttle touching live traffic, ever, returns to steering. Osei withdraws the objection and keeps the note. Nadia Benali, asked informally whether Accept-until-Phase-3 was ever an option, answers in her capacity as CRO with veto on risk acceptance: a materialized risk on a certification-path service, three months before the audit, is not acceptable by anyone, including her.
The Four Responses, Assessed
| Response | Assessment | Verdict |
|---|---|---|
| Avoid | Redesign onboarding flow to make KYC verification synchronous, removing the queue entirely. | Rejected: re-architecture of a feature-complete service in month 11; contradicts ADR-002's no-migration-only rule; cost and blast radius unjustifiable. |
| Transfer | Move verification workload to the vendor's managed KYC capability under a contract extension. | Rejected: same regulated-control vendor-dependency position steering declined in CR-001 v1.1; clause 3.4 scope change for a capacity problem. |
| Mitigate | Queue-depth-driven consumer autoscaling with backpressure; interim time-boxed intake throttle in shadow; trigger condition wired as a live alert. | Adopted: $52K, deployable 29 Dec, removes the load-shape failure mode and the detection gap that hid it. |
| Accept | Tolerate Monday breaches until Phase 3 performance testing addresses capacity holistically. | Rejected: materialized risk on a certification-path service inside the audit window; CRO veto on risk acceptance confirmed unavailable (N. Benali, 9 Dec). |
Drafting the Risk Response with OODA
A risk response document has exactly the shape of a decision made under incoming information, which is why this one is drafted on OODA. Observe, Orient, Decide, Act: John Boyd's decision loop, originating in military strategy and since adopted widely in incident response and security operations, well-documented in both. Its application as a drafting structure for a program risk response is the practitioner adaptation here, and the fit is unusually literal. Observe is the materialization signal and its validation. Orient is the signal placed against the register, the audit finding, and the certification timeline. Decide is the response choice with the rejected options preserved. Act is the execution plan with owners, dates, and the verification that closes the loop. A framework built for contact with reality suits a document triggered by it.
The Draft, and the Flag About the Register Itself
An extract from the Observe and Orient sections, and a flag that widened the document's job from one risk to the whole register.
OBSERVE. Materialization signal: R04's written trigger condition (kyc-verify queue depth > 50K sustained > 4 hours) has been met on every Monday of the shadow pilot, peaking at 78K with drain completing near 22:40, approximately nine hours above trigger. Source: Splunk/Dynatrace hourly telemetry, 17 Nov to 8 Dec. Detection: AI-assisted risk review preparation, 5 Dec; the breach is not visible in daily-mean reporting (mean 11K). Validation: risk owner re-cut raw hourly data for three consecutive Mondays and confirmed the breach, 9 Dec.
ORIENT. R04 has been on the register since 23 April with this exact trigger language; the condition was authored, approved, and then never connected to an alerting rule. Internal audit finding (controls score 4.7/5) identified the precise blind spot: performance testing bounded at 2x average load, sustained-peak behavior unexercised. The materialization occurred inside the untested region, on a certification-path service, fourteen weeks before the certification audit. The detection delay was 18 days, bounded only by the pilot's start date and the monthly review cadence.
The flag is the most valuable sentence the machine produced, and it is not about R04. It is about the register as a system: eleven months of well-written trigger conditions, of which an unknown number are promises with no wiring behind them. Fasil commissions the register-wide instrumentation audit the same day, due 9 January, owner shared between Raman and Hassan. Eight of the register's open risks turn out to have machine-checkable trigger conditions; before this week, exactly one of them was wired. The risk register was a library. It is about to become a smoke detector.
RR-R04, as Issued
Response Choice
MITIGATE. Avoid, Transfer, and Accept assessed and rejected with recorded verdicts (table on file; Accept ruled out under CRO risk-acceptance veto, N. Benali, 9 Dec 2025). Link: Risk Register v1.0 (23 Apr 2025), R04; CR-001 v1.1 vendor-dependency rationale; ADR-002 migration rule.
Execution Plan
| Permanent fix | Queue-depth-driven consumer autoscaling with backpressure signaling on kyc-verify; $52K within PM single-item authority (December aggregate updated by A. Hassan). Deploy by 29 Dec 2025. Owner: P. Raman |
| Interim measure | Shadow-pilot intake throttle, Monday peak window only, effective 10 Dec 2025, hard expiry 5 Jan 2026. Shadow verdicts are not customer-facing; zero live-traffic impact. Any future throttle touching live traffic requires steering approval. Owner: J. Park |
| Detection fix | R04 trigger condition instrumented as live Dynatrace alert with named on-call ownership; alert-to-acknowledgment target 15 minutes. Live by 19 Dec 2025. Owner: P. Raman |
| Register action | Register-wide trigger-instrumentation audit: every open risk with a machine-checkable trigger condition gets wired or gets a written reason why not. Due 9 Jan 2026. Owners: P. Raman, A. Hassan |
Verification Criteria
Response closes when two consecutive peak Mondays complete under the 50K trigger with the throttle lifted, evidenced from hourly telemetry by 12 Jan 2026. Evidence pack enters the certification file as demonstration of detection, response, and closure on a certification-path service. Link: Latency evidence plan (28 Nov 2025); Phase 3 readiness criteria.
Benefit Impact
Revenue category (B01-B03) baseline unaffected: throttle applies to shadow intake only and expires before live onboarding scale-up. Note co-signed S. Osei, 10 Dec 2025, attached. Link: Benefits Plan (Post 08), Revenue category baseline.
Residual Risk
R04 remains open at probability Low post-mitigation pending Phase 3 sustained-peak test execution (the Post 09 audit finding's original remediation), scheduled February 2026. The register entry is updated, not closed; mitigated is not finished. Link: Internal audit report (Post 09), finding 1.
What the Human Changed
- Held the document until the human validation existed. The AI's draft was ready on 8 December with the signal marked “detected.” Fasil refused to issue a risk response on a model-surfaced pattern alone; the Observe section was rewritten to record Raman's independent re-cut as the validation event, dated, with the raw-data method named. A register that logs machine hypotheses as findings will eventually cry wolf, and this register cannot afford to.
- Promoted the instrumentation flag from observation to commitment. The draft recommended a register-wide audit; recommendations are where good ideas go to wait. The issued document converts it into a dated, owned register action due 9 January, because the difference between this post and a postmortem is detection latency.
- Time-boxed the throttle with a hard expiry and a steering tripwire. The draft's interim measure was open-ended “until mitigation deploys.” Open-ended interim measures become permanent furniture. The issued version expires 5 January in writing, and any throttle ever touching live traffic goes to steering, which is the sentence that converted Osei's objection into a co-signature.
- Kept R04 open at reduced probability instead of closing it. The draft moved R04 to Closed on mitigation deployment. The Post 09 audit finding's actual remediation, sustained-peak performance testing, has not happened yet; it happens in February. Closing a risk because you fixed the instance of it is how registers learn to lie politely.
- Added the certification evidence framing. The draft treated the episode as an internal matter. Fasil added the verification pack's routing into the March certification file: a regulator auditing a program's risk management is best answered with a risk that materialized, was caught, was fixed, and is documented end to end. The incident is the evidence.
The register-wide instrumentation audit lands 9 January and wires seven more trigger conditions into live alerts. That early-warning mesh is the difference between an incident and a catastrophe when the 72-hour event arrives (Post 19), where detection latency is measured in minutes because December made it so. The RR-R04 verification pack becomes a named exhibit in the March readiness assessment (Post 16). And the episode's core lesson, that the program's risk register had been a library rather than a smoke detector for eleven months, is written up in the lessons register (Post 21) as the cheapest expensive lesson Falcon ever bought: $52K, zero customer impact, caught in shadow.
The verification closed on schedule: 5 and 12 January, two clean Mondays, no throttle, peak depth 31K. Phase 3 opens with the register wired for sound. What it opens into is the densest artifact of the series: the readiness assessment that decides whether this program is allowed to meet its regulator. That is Post 16.
A fictional case study for teaching purposes. Atlas Bank, Project Falcon and all named individuals are invented. Technologies are industry-standard and publicly available.