Do you actually need on-call?
Not every product needs 24/7 on-call. Before setting one up, ask:
- Is your product used outside business hours? (If 95% of usage is 9am–7pm IST, overnight on-call may not be worth it)
- Does downtime cost significant revenue or break SLAs?
- Do you have paying customers who expect availability?
If the answer to any of these is yes, you need some form of on-call. If not, a next-business-day response to incidents may be sufficient — and significantly better for team morale.
The minimal viable on-call setup
For a team of 3–5 engineers:
- One engineer is "on-call" per week — they carry the responsibility from Monday 9am to the following Monday 9am
- One engineer is "secondary on-call" — they're the backup if the primary can't be reached within 15 minutes
- Rotation: each engineer is primary for 1 week out of every N weeks (where N = number of engineers in the rotation)
- The on-call engineer gets a compensating day off the following week (or equivalent recognition)
What the on-call engineer is responsible for
Be explicit about this — ambiguity leads to incidents going unacknowledged:
- Acknowledge any production alert within 15 minutes
- Diagnose the incident and attempt resolution using the runbook
- If not resolved within 30 minutes, escalate to secondary on-call and/or team lead
- Post a status update in the #incidents Slack channel every 15 minutes during an active incident
- Write a brief incident report within 24 hours of any P1 incident
💡 "On-call" doesn't mean "at your desk all night." For most small teams, on-call means: keep your phone on, be reachable within 15 minutes, be able to VPN into the system and start diagnosing. It doesn't mean sitting in front of a laptop from 8pm to 8am.
What alerts to set up
Only alert on things that require human action right now. Alert fatigue kills on-call effectiveness — if the on-call engineer gets 30 alerts per night and 28 of them are noise, they'll start ignoring them.
Good things to alert on:
- API error rate above 5% for 5+ minutes
- Server CPU above 90% for 10+ minutes
- Database connection pool exhausted
- Payment processing failure rate above 0%
- Any 5xx error spike more than 2x baseline
Bad things to alert on:
- Every single 5xx error (1–2 per hour is normal)
- CPU spikes under 1 minute (transient, self-resolving)
- Informational events that don't require action
The on-call handoff
When on-call shifts change, the outgoing engineer should do a 10-minute handoff:
⚠️ New engineers shouldn't start on-call alone. For the first 2–3 weeks, have a new engineer shadow an experienced on-call engineer. They're on-call together — the new engineer handles the alert, the experienced one observes and steps in if needed. This builds confidence and prevents panic-driven mistakes.
Track on-call incidents in Resolvo
Issue tracking, incident logs, SLA monitoring. Built for engineering teams. Free to start.
Start Free →