The Playbook
Most Customer Success functions are still run on vibes: a spreadsheet, a gut feeling about which accounts are "fine," and a QBR deck built the night before. Every other function touching production risk — infrastructure, security, finance — long ago moved to systems: instrumentation, defined targets, standardized response procedures, and a blameless review when something breaks. Customer Success hasn't caught up, and the gap shows up as surprise churn.
The Playbook is an attempt to close that gap: treat the customer base the way a reliability engineer treats a fleet of production systems.
Why this framing
Twenty-plus years in enterprise infrastructure and managed services means being the person a customer calls right after they've already tried turning it off and on again. That job runs on the same discipline as Spacelift and the practice labs on this site: plan and apply are structurally separate, nothing touches production without a human confirming, and every failure gets a real postmortem instead of a shrug. The Playbook applies that same discipline to the customer relationship instead of the AWS account.
The other half of the framing: since 2024, building AI agents (see MAC) that now do the boring 60% of this job. The Playbook is written for a world where automation already handles detection and drafting — the open question isn't whether AI touches the customer relationship, it's exactly where the human has to stay in the loop. The Confirm Gate is the chapter that answers that.
The six plays
Instrumentation — knowing the state of the fleet before anything goes wrong:
- Observability (Customer Health Scoring) — the monitoring layer: what to measure and why
- SLOs for Customer Outcomes — what "healthy" actually means, as a number, per segment
Response — what happens when the monitor fires:
- The Play Library — standardized runbooks for recurring account signals
- The Confirm Gate — where automation stops and a human has to sign off
When it breaks — the failure path, done properly instead of ad hoc:
- Incident Response for At-Risk Accounts — declaring, staffing, and running a save motion
- The Postmortem (Churn and Save Retros) — the blameless review that feeds back into every other page here
Supporting concepts
- Customer Lifecycle Stages — the timeline the plays run against, from onboarding to advocacy
- CSM vs TAM vs Customer Engineer — who on the team actually owns which play
Who this is for
CS leaders scaling a motion past the point where tribal knowledge holds it together, and senior CSMs/TAMs who want a system instead of another "be customer-obsessed" slide. See The Playbook (book) for the long-form version of this argument.