Incident and operations
Alerts, on-call triage, mitigation, runbooks, observability and postmortems.
Download all 26
- Instrument a service for observability
Plans and adds logs, metrics and traces using OpenTelemetry conventions, golden signals, useful log fields, cardinality limits and first dashboards. Use when a service is hard to debug in production.
- Plan a game day or chaos exercise
Plans a game day or chaos exercise with failure scenarios, hypotheses, blast-radius limits, abort criteria, roles, an observation checklist and a follow-up review. Use to test resilience.
- Triage a production alert
Turns a firing production alert into a severity call, the safest mitigation to try first, ranked hypotheses and the next checks. Use in the first minutes of an incident or page.
- Write an incident status update
Writes a clear status update for an ongoing incident, tuned to customers, internal teams or executives, without speculation or promises the team cannot keep. Use for status pages, Slack and email.
- Write a blameless postmortem
Turns incident notes, chat logs and timelines into a blameless postmortem with impact, timeline, contributing factors and owned action items. Use after an incident is resolved.
- Audit postmortem action items
Reviews action items across recent postmortems for done, stale and vague items and repeated systemic themes, and rewrites each open item to be specific, owned, dated and verifiable.
- Build an incident timeline
Builds a timestamped incident timeline from chat logs, alerts and deploy records, marking detection, escalation, mitigation and the gaps between them. Use when preparing a postmortem.
- Collect evidence for an ongoing incident, read-only
Gathers evidence for a live incident with read-only commands, covering recent deploys, error rates, logs and resource use, and writes a timestamped evidence summary. Use while responders work the fix.
- Define SLOs and burn-rate alerts
Defines SLIs, SLOs and an error-budget policy from a service's user journeys, with multi-window burn-rate alert rules. Use when alerting is noisy or reliability targets are vague.
- Design actionable alerting rules
Designs actionable alerts from SLOs and user-facing symptoms, with thresholds, routing, runbook links, and a list of noisy alerts to delete. Use when pages are noisy or real outages go unnoticed.
- Design an on-call rotation
Designs an on-call rotation with schedule, escalation, handoff, alert ownership, compensation norms and health checks. Use when starting on-call or when the current one burns people out.
- Design a service dashboard
Designs an operational dashboard for one service, with golden signals on top, dependencies, saturation, deploy annotations and drill-down order, plus the query and on-call action for each panel.
- Incident commander
Runs a live incident like an experienced incident commander, assigning roles, keeping a steady comms cadence and driving mitigation before root cause. Use as the coordinating voice during an outage.
- Investigate a latency spike
Walks an on-call engineer through a live latency spike one piece of evidence at a time, from percentile and endpoint to deploys, saturation or a slow dependency, and the safest mitigation.
- Logging rules
Standing rules for logs an assistant writes, with structured fields, meaningful levels, no secrets or personal data, correlation IDs and errors logged once where they are handled.
- Add logs, metrics and traces to a service
Instruments a service in gated steps with structured logs, metrics, traces, correlation ids, business metrics, dashboards as code and a verification run. Use when a service is a black box.
- On-call readiness track
Gets a new service ready for on-call in gated steps, from SLOs on user journeys to symptom alerts, a dashboard, runbooks per alert and an escalation and go-live check.
- Fix a simulated broken server
Hands the learner a simulated Linux server with a hidden fault such as a full disk, a dead service or bad permissions, answering their commands until they find and fix it.
- Prune noisy alerts
Analyses an alert history export to cut alert fatigue, finding alerts nobody acts on, flapping, duplicates and missing owners, with a delete, tune, route or keep decision per alert.
- Rehearse incident command
Runs a simulated incident with the learner as incident commander, injecting alerts, confused responders and executive pings turn by turn, then debriefs coordination and communication.
- Site reliability engineer
Acts as a site reliability engineer who thinks in SLOs and error budgets, automates toil, designs for failure and writes blameless reviews.
- Triage a mobile crash spike
Triages a crash-rate spike after a mobile release, isolating affected versions, devices and OS, server or flag causes, and deciding whether to halt rollout, kill-switch a feature or hotfix.
- Write an incident response plan
Writes an engineering incident response plan - severity levels with criteria, response targets, roles, escalation paths and copy-paste communication templates - as a quick reference for on-call.
- Write observability queries
Writes PromQL, LogQL, TraceQL, SQL or vendor queries for an operational question, explains each part and warns about traps such as rate windows, counter resets and label cardinality.
- Write an on-call handoff
Writes the end-of-shift on-call handoff covering open incidents, alerts that fired and why, silences and their expiry, risky changes in flight and what to watch. Use at every rotation change.
- Write an operational runbook
Writes a runbook for an alert or routine procedure with symptoms, diagnosis commands, ordered mitigations, verification and escalation. Use so on-call engineers can act without tribal knowledge.
Not: a defect found in development (debugging).