DevOps / SRE · Playbook

DevOps / SRE Playbook

A practical guide to shipping and running software without heroics. Metrics describe health. Platforms describe the path. Incidents describe learning. Keep those ideas joined so you do not build a beautiful pipeline that nobody can operate, or a war-room culture that burns people out.

MAP-01 Reliability maps

MAP-01

Reliability & delivery maps

Pick the map that matches the decision in front of you.

DevOps and SRE mix delivery performance, reliability science, platform product thinking and security hygiene. Confusing them creates either endless tooling or a hero culture. Every framework below answers a different question.

Maps you will actually use

MapAnswersShape
DORA metricsIs delivery getting healthier?Deploy frequency · lead time · change fail · recovery
SLO / error budgetsHow reliable must we be, and what do we trade when we burn?SLI · SLO · budget · policy
Continuous DeliveryCan we release safely on demand?Trunk · pipeline · small batches · rollback
ObservabilityCan we ask new questions of a live system?Metrics · logs · traces · continuous profiling
Incident managementHow do we restore service and learn?Command · comms · timeline · review
Secure SDLC / supply chainAre we shipping known risk?SAST · SCA · secrets · provenance · least privilege

Choose in practice

  • Exec debate on speed vs stability: DORA trends plus SLO burn, not vibes.
  • Feature pressure eating reliability: write an error budget policy with Product Owner.
  • Painful releases: shrink batch size and automate rollback before adding more CAB theatre.
  • Pager storms: fix alert quality and SLOs before hiring more on-call bodies.
  • Government service: join operability evidence to Service Standard and NCSC good practice continuously.
If a metric never changes a release or reliability decision, stop performing it.