Here's what a real incident looks like.
Hour 0. 3.47 AM. Pagerduty fires. RPC node down in us-east-1 (primary). Secondary in us-west-2 should failover automatically. It doesn't. The dashboard shows the primary is hung, syncing but not responding. Mempool backed up. Transactions timing out.
First mistake. The on-call engineer assumes the problem is where the alert says. But the RPC is fine. The problem is network. Fifteen minutes of diagnostics. The secondary is unreachable. DNS record pointing to a stale IP. A deployment six hours earlier updated RPC infrastructure but didn't update DNS.
The engineer realizes the mistake. Rollback or move forward? Moving forward is faster. DNS gets fixed.
Hour 2. 5.30 AM. DNS is fixed. Secondary RPC is reachable. But massive mempool backlog. 15,000 transactions. The RPC is falling behind. Blocks are created but many unconfirmed.
First support email. "Where is my transaction?" The team has a template response. Customers get added to a manual queue (a spreadsheet). No actual handling. Just a promise to look later.
Hour 4. 7.00 AM. Immediate crisis is over. RPC recovering. Processing is slow. Primary isn't fully synced. The product manager asks "How long until normal?" Nobody knows. A guess. Two hours.
Hour 5. 8.30 AM. A customer's withdrawal is stuck. They had balance. It's rejected as insufficient. But the balance query was from the primary (stale). The transaction routed to the secondary (older state). State divergence. The two RPC nodes are out of sync.
One node has to reset and resync from zero. Eight to twelve hours. While syncing, it can't process transactions. Network is degraded.
Customer communication is drafted. Bland. No mention of lost transactions. Logs are manually searched to find affected transactions. Text editor. No tooling.
Second support email. "I've lost money. What are you doing?"
Nobody knows if customers lost money. Uncertainty is the worst part of an incident. There's no visibility into actual impact.
Hour 12. 3.30 PM. Secondary is clean. Primary catching up. Both processing correctly. Mempool clearing. All-clear gets sent. It's wrong.
Hour 18. 10.00 PM. A transaction from early in the incident finally confirms. A customer gets a credit they weren't expecting. They reissued it two hours ago. Duplicate ledger entry.
End-of-day reconciliation catches it. Payment count off by one. The team builds a deduplication process. Four duplicates total. Three were small. One was $240,000. Manual refund needed.
Hour 24. 4.00 AM. Post-mortem being written.
The root cause. DNS record not updated during infrastructure change. This triggered a failover. The failover succeeded but RPC nodes diverged. Transaction queries hit inconsistent state. Duplicates were created.
But that's not the full story. The real root cause is the failover was never tested. Redundancy was built but never exercised. If it'd been tested, the DNS issue would've surfaced weeks earlier. The second root cause is the API layer didn't handle RPC inconsistency. It should've had circuit breakers. Should've detected divergence. Should've rejected transactions if any node had insufficient balance. It didn't.
The third root cause is manual incident response. Spreadsheets and text files. No automated tooling. No dashboard showing which transactions are affected, which need reissues, which are valid. Nothing.
We decide on seven action items.
"Add failover testing to the deployment pipeline," the VP says. "Monthly exercises where we intentionally down the primary RPC and verify secondary failover."
"Add state synchronization monitoring. If the nodes diverge by more than one block, alert immediately."
"Add transaction state tracking. Every transaction has a status. Pending. Confirmed. Failed. We need to track this end-to-end. Not just check a blockchain explorer."
"Add a deduplication system. Scan for duplicate transactions on transaction commit. Reject if the transaction hash already exists in the ledger."
"Build an incident dashboard. During an incident, a single page shows all affected transactions, their status, which ones need reissue."
"Create an RPC consistency checking system. Periodically query all RPC nodes and compare their view of the latest block. Alert if they diverge."
"Build a customer communication protocol. During incidents, update affected customers automatically. Don't wait for support emails. Proactively tell them their transaction is affected."
The post-mortem is honest. Too honest. We messed up. Multiple times. The systems we built didn't handle the failure modes we should've anticipated.
What you should do differently is implement the seven things we implemented. But do it before you have an incident. Not after.
Test failover monthly. Not theoretically. Actually take down a primary system and verify redundancy works. You'll find problems. The problems won't destroy a customer's transaction during production. They'll destroy your test.
Build monitoring for state consistency. If a distributed system diverges, you need to know in seconds. Not hours. Not when a customer complains.
Track transaction state explicitly. Every transaction is an object with a status. Pending. Confirmed. Failed. You query this status, not a blockchain explorer. You control the truth.
Deduplicate. Even if it shouldn't happen. Build the deduplication. It will happen. Mempool congestion. Retries. Network delays. Some transaction will hit your system twice. Be ready.
Build an incident dashboard. Not Slack. Not spreadsheets. A single page showing every piece of information you need. All affected transactions. Their status. What action you need to take. This doesn't exist in most systems. It should.
Test customer communication automatically. During an incident, can you reach affected customers in less than a minute? With exact information about what's wrong? Most companies can't. Test this.
The final thing. Post-mortems are useless if you don't act on them. We wrote the post-mortem. We identified seven action items. We shipped six of them in the next month. The seventh, customer communication automation, took three months. It's still not perfect.
But when the next incident happened three months later (and there was a next incident), we found it in six minutes instead of two hours. We told customers in ninety seconds. We resolved it in four hours instead of twenty-four.
The playbook you write during an incident is a lie. The playbook you write after an incident works.
Barely.