A Guided Tour of Google's Site Reliability Engineering (Part 1)
Google’s Site Reliability Engineering, published in 2016 and available in full for free at sre.google, is a foundational text on modern distributed systems operations and reliability engineering.
Written by Google’s SRE team, the book has profoundly reshaped how the tech industry thinks about operations over the past decade.
For backend engineers moving into DevOps, it shows how Google used engineering and economic reasoning to turn system reliability from a matter of luck into an organizational discipline with measurements, budgets, and checks and balances.
This first installment covers Chapters 1, 3–6, 11, 14, 15, and 27 of the book’s 34 chapters, plus Appendix D. It spans SRE’s organizational origins, how reliability is quantified, toil, and four everyday practices—alerting, on-call, blameless postmortems, and launch reviews. For the relationship between SRE and DevOps, it draws on the 2018 follow-up, The Site Reliability Workbook.
1. What Is SRE? The Seven-Person Team of 2003 and an Organizational Redesign
SRE has a clear historical starting point. Chapter 1 was written by Ben Treynor Sloss, then Google’s VP of Engineering responsible for its production environment. He recalls that when he joined Google in 2003, he took over a seven-person “Production Team”—the precursor to Google’s SRE organization.
Treynor opens with the classic definition:
SRE is what happens when you ask a software engineer to design an operations team.
In Chinese-language tech communities, this is often loosely paraphrased as “SRE means getting software engineers to do what ops used to do.” But the emphasis is on “design an operations team”, not merely changing who performs the work. SRE is not about making software engineers endure traditional ops pain; it is about redesigning the organization and its processes from the ground up with a software engineering mindset.
Escaping the Trap of Linear Scaling
Google moved away from the traditional sysadmin model not because its practitioners lacked skill, but because of a structural problem: staffing had to grow in lockstep with service traffic and server count. When the business grows exponentially, operations that scale only by adding people eventually become untenable.
Scale vs Service Growth
[Traditional Sysadmin]
Scale : ------------------------> (10x)
Staff : ------------------------> (10x Linear Growth)
[Site Reliability Engineering]
Scale : ------------------------> (10x)
Staff : --------> (Sublinear Growth via Software)
To break that linear relationship, Google deliberately recruited engineers who could write software. The book describes two practical traits of these engineers:
- They quickly tire of repeating manual tasks.
- They can write software to automate those tasks.
About 50%–60% of Google SREs are software engineers (SWEs). The other 40%–50% are specialists who meet 85%–99% of the SWE hiring bar and bring expertise in UNIX systems and networking internals. According to the book, there is no material difference in their job performance; the two backgrounds complement each other.
Code or Drown: The 50% Line of Defense
This leads to a central organizational rule:
the team tasked with managing a service needs to code or it will drown.
To protect time for coding, Google caps SREs’ operational work at 50% (the work called toil, discussed in Section 4). When operational load exceeds that threshold, excess tickets, bugs, and troubleshooting go back to the product development team. Developers may even join the on-call rotation.
Handing problems back to the teams that created them breaks the pattern in which developers ship features while operations absorbs all the reliability debt.
That is why Chapter 1 opens with the epigraph “Hope is not a strategy,” which Chapter 3 calls an unofficial SRE motto. The point is simple: reliability cannot depend on developer caution or luck; it needs organizational safeguards and engineering constraints.
2. SRE and DevOps: class SRE implements DevOps
The best-known line about the relationship between SRE and DevOps is:
class SRE implements interface DevOps
This is a Java object-oriented metaphor. An interface specifies which methods must exist, not how they work; a class that implements it must supply the concrete behavior.
Applied to organizations, DevOps supplies principles for collaboration between development and operations without prescribing a particular implementation. SRE is one way Google puts those principles into practice. Other implementations are possible.
The line, however, is not in the 2016 book. It first appeared in Chapter 1 of the 2018 follow-up, The Site Reliability Workbook, directly beneath the chapter title. It addresses the industry’s debate over whether SRE and DevOps are competing approaches.
The Five DevOps Pillars in SRE Practice
The same chapter maps the five dimensions of the CALMS framework (Culture, Automation, Lean, Measurement, Sharing) to concrete SRE practices:
| DevOps principle | SRE practice | Underlying idea and original wording |
|---|---|---|
| Culture | Blameless postmortems | Replace fear and blame with learning. “A good culture can work around broken tooling, but the opposite rarely holds true.” |
| Automation | The 50% cap on toil | Limit operational work to protect time for engineering. |
| Lean | Small, frequent changes and shorter MTTR | Limit the blast radius of releases and reduce mean time to repair. |
| Measurement | SLOs and error budgets | “Measurement is absolutely key to how both DevOps and SRE work.” Base service decisions on measurable outcomes. |
| Sharing | Common tools and shared ownership | Developers and SREs use the same tools and share responsibility for production. |
The Workbook sums it up at the end of its Compare and Contrast section:
SRE believes in the same things as DevOps but for slightly different reasons.
3. Quantifying Reliability: SLIs, SLOs, SLAs, and the Economics of Error Budgets
Three Levels of Commitment
SRE replaces vague impressions of reliability with three related concepts:
[SLI: Service Level Indicator]
| Quantitative measure (e.g., latency, error rate)
v
[SLO: Service Level Objective]
| Target value or range (e.g., 99.9% success rate)
v
[SLA: Service Level Agreement]
Contract with users with explicit consequences
- SLI (service level indicator): A precise measure of a particular aspect of a service, such as request latency or success rate. Look at tail latency, not just averages: a P99 value, for example, is a threshold below which 99% of requests fall. A concrete indicator might ask whether Get RPC latency stays under 100 ms in a one-minute aggregation window.
- SLO (service level objective): A target value or range for an SLI, such as “99.9% of requests must succeed each month.”
- SLA (service level agreement): An explicit or implicit agreement with users that specifies consequences for missing the target, such as service credits or refunds.
Chapter 4 gives a straightforward test: without explicit business consequences for missing the target, it is an SLO, not an SLA.
Why 100% Reliability Is the Wrong Target
Chapter 3 makes an economic argument against pursuing extreme reliability—100%, or simply too many nines. Such a goal can be impractical and harmful to users and the business:
- Users may not notice the difference: Someone using a mobile network that is only 99% reliable cannot distinguish between 99.99% and 99.999% availability in a backend service.
- Marginal costs rise sharply: Adding another nine of availability could cost 100 times as much in engineering and redundant infrastructure.
Consider the book’s example: a service earns 900 of additional potential revenue a year:
The test is simple: adding another nine is worth the investment only if it costs less than $900; spend any more, and the cost outweighs the added potential revenue.
Google SRE therefore treats the availability target as both a floor and a ceiling: “we view the availability target as both a minimum and a maximum.” Reliability above the target may offer little additional benefit while suggesting that the team is being too cautious about shipping features.
How Error Budgets Work
An error budget turns that economic judgment into a rule for everyday decisions. It is the maximum unreliability a service can incur over a specified period:
For a 30-day rolling window:
- 99.9% SLO allows 43.2 minutes of downtime.
- 99.99% SLO allows just 4.32 minutes.
[Feature Release Queue]
|
v
[ Error Budget > 0? ]
/ \
(Yes) (No)
/ \
v v
[ Release Feature ] [ Freeze & Fix Reliability ]
The budget acts as a control valve for releases:
- When budget remains, developers can ship features and try refactors or architectural upgrades.
- When it runs out, Chapter 3, “Embracing Risk,” calls for a release freeze. Non-routine changes stop, and engineering effort shifts to writing integration tests, fixing performance bottlenecks, and improving disaster recovery (DR) until reliability recovers.
In practice, an SLO needs a written error budget policy approved by product managers, developers, and operations, as Chapter 2 of the Workbook emphasizes. Without an agreed response to an exhausted budget, an SLO risks becoming just another dashboard number.
Two Counterintuitive Rules for Setting SLOs
Chapter 4 offers two particularly counterintuitive rules:
- Leave a safety margin: Set the publicly stated SLO below the internal monitoring target. The stricter internal threshold gives the team time to detect and repair problems before users notice or SLA consequences apply.
- Don’t overachieve: If a service consistently performs far better than its stated SLO, downstream users may come to rely on that unstated level of reliability. When availability returns to the published target, their applications may fail.
The book’s example is Chubby, Google’s distributed lock service. After long stretches without failures, its operators would deliberately schedule brief outages to test whether dependent applications had implemented reconnection, retries, and graceful degradation.
⭐️4. Toil: SRE’s Organizational Enemy
Toil has a precise meaning in SRE. Chapter 5 defines it as follows:
Toil is the kind of work tied to running a production service that tends to be manual, repetitive, automatable, tactical, devoid of enduring value, and that scales linearly as a service grows.
Six Characteristics of Toil
Ask whether the work has these six characteristics:
- Manual: A person must type commands or click through a UI.
- Repetitive: The same kind of problem recurs.
- Automatable: Software or a script could do it instead.
- Tactical: It treats the immediate symptom, not the underlying cause.
- Devoid of enduring value: After the task, the system is no better than before.
- Scales linearly: Double the service traffic and the work doubles too.
The book is clear that toil is not merely “work I don’t like.” It also differs from necessary organizational overhead, such as team meetings, goal setting and grading, or HR paperwork.
Engineering Leverage and Sublinear Growth
The 50% cap in Section 1 is a long-term average. SREs should spend at least the other half of their time on engineering project work—writing automation, refactoring infrastructure, and building self-healing systems that eliminate future toil.
SRE Work Distribution
[Engineering Project Work (>= 50%)]
+-- Build Automation Frameworks
+-- Architectural Hardening
+-- Self-Healing & Scalability Engineering
[Operational Toil (< 50%)]
+-- Manual Interventions & Ticket Response
+-- Repetitive On-Call Routine Tasks
Engineering projects compound: they are what allow an SRE organization to grow sublinearly as its services scale. Regular internal Google surveys found that SREs spent about 33% of their time on toil on average—proof that the 50% cap works as a real line of defense.
5. Everyday Practice: Alerting, On-Call, Postmortems, and Launch Reviews
The ideas become concrete in four operational practices:
1. Monitoring and Alerting: The Four Golden Signals
Chapter 6 boils monitoring’s core job down to two basic questions: “What’s broken?” and “Why?”
For detecting symptoms, Google identifies the four golden signals:
- Latency: How long requests take; measure successful and failed requests separately.
- Traffic: The demand on the system, such as QPS or concurrent connections.
- Errors: The proportion of requests that fail, including HTTP 5xx responses, implicit business errors, and timeouts.
- Saturation: How close a resource is to its limit, such as memory pressure or connection-pool queues.
An alerting system should be reliable and simple. The book sets a strict bar for a page—an alert that summons the on-call engineer: it should be urgent, actionable, and already affecting users or about to do so.
If responding to an alert means following a fixed standard operating procedure (SOP), automate the procedure instead of paging someone.
Too many alerts cause alert fatigue. Chapter 11 urges teams to control alert fan-out—a single incident spawning multiple alerts—and work toward a roughly 1:1 ratio between alerts and real incidents, where nearly every alert corresponds to a real problem.
2. On-Call: A Sustainable Rotation
Chapter 11 spells out how to structure on-call work without burning engineers out:
- Roles: The primary responds to urgent pages; the secondary handles non-urgent production issues.
- Minimum team size: A single site covering a service 24/7 needs at least eight people. In a two-site follow-the-sun rotation, each site needs at least six.
- Limits on time and incidents: An engineer should spend no more than 25% of their time on call, and a 12-hour shift should see no more than two incidents. More suggests the operational load is unsustainable.
- The ultimate safeguard: If a service’s failure rate stays chronically high, SREs can give back the pager, returning its operations and on-call responsibilities to the development team.
Alert quality and on-call workload form a feedback loop. Oversensitive rules drive up pages; too many incidents force the team to improve the service—or hand the pager back.
For major incidents, Chapter 14 draws on the Incident Command System (ICS) used in emergency response, which separates four roles:
- Incident Commander: Maintains the overall picture and coordinates decisions.
- Operations Lead: Works on the technical fix.
- Communications Lead: Keeps stakeholders informed.
- Planning Lead: Supports Ops with longer-term issues, such as filing bugs, arranging handoffs, or even ordering dinner.
This division keeps remediation and communication from competing for one engineer’s attention during a crisis.
3. Blameless Postmortems
Chapter 15 and Appendix D describe both the culture and the format of a postmortem.
Its central assumption is:
A blamelessly written postmortem assumes that everyone involved in an incident had good intentions and did the right thing with the information they had.
Blame encourages engineers to hide mistakes and architectural weaknesses. Instead, organizations should remember: “You can’t fix people, but you can fix systems and processes.”
A postmortem should include:
- Impact: Quantified harm to users, such as affected queries or lost revenue.
- Root causes and trigger: The technical chain of failure.
- Timeline: A minute-by-minute record from the first anomaly through the alert to recovery.
- Action items: Concrete engineering tasks with a type, owner, and tracking issue.
- Lessons learned: What worked, what did not, and what succeeded only by luck.
4. Launch Reviews
Chapter 27 describes how Google developed its launch review process after forming the Launch Coordination Engineering (LCE) team in 2004.
A standardized launch checklist examines nine areas, including architectural dependencies, capacity planning, failure isolation, backups, and release strategy.
Each item needs a concrete reason for being there, ideally a painful past launch: “Every question’s importance must be substantiated, ideally by a previous launch disaster.”
Early reviews caused friction when inexperienced SREs were overly conservative and delayed launches. Over time, the process became more collaborative: SREs offered technical advice and support rather than acting as bureaucratic gatekeepers.
Conclusion and a Preview of Part 2
The core of Google’s SRE approach in the first half of the book can be understood in four layers:
- Definition: Redesign operations with a software engineering mindset so staffing does not have to grow in lockstep with traffic.
- Quantification: Use SLOs and error budgets to make reliability an economic decision, and set explicit conditions for releasing features.
- Structure: Protect engineering time with a 50% cap on toil and a mechanism for returning excess operational work to development.
- Practice: Put engineering discipline into daily operations through actionable alerting, sustainable on-call, blameless postmortems, and launch checklists.
These highly idealized practices, however, grew out of Google’s vast scale, deep engineering bench, and powerful internal tools. What can smaller organizations—or teams outside Google—take from them?
Part 2 steps beyond these core principles to examine:
- How systems fail: Cascading failures, overload protection, and retry storms, with comparisons to modern Kubernetes architectures.
- SRE outside Google: What Evernote and The Home Depot experienced while adopting SRE, including their difficulties and failure modes.
- A decade of change (2016–2026): Which SRE ideas still hold up amid platform engineering and mature observability tooling, and which assumptions need revisiting?