Honest Take — Before You Begin
The Google SRE book will challenge assumptions you did not know you had. The central idea -- error budgets -- is counterintuitive: if your service has been too reliable, you shoul…
CI/CD, SRE & Production Operations covers: Site Reliability Engineering. Build deployment pipelines, practice Site Reliability Engineering principles, and develop the skills to own production systems end-to-end.
On-call for Rails apps. Error budgets. SLIs/SLOs for response time and error rates. GitHub Actions for Rails: RSpec, Rubocop, security scanning, deploy to staging.
| SRE Concept | Rails Application |
|------------|------------------|
| SLI: latency | p99 response time for /api/orders |
| SLI: availability | Uptime of the Rails app (target 99.9%) |
| SLO: error rate | Less than 0.1% 5xx responses |
| Error budget | 43 minutes of downtime/month at 99.9% |
| Canary deploy | Kamal with --rolling or K8s rolling update |
| Incident response | PagerDuty alert → triage → fix → postmortem |
| Postmortem | "Why did the migration lock the users table for 45 minutes?" |
This course unlocks once you've finished its prerequisite. Open prerequisite →
The Google SRE book will challenge assumptions you did not know you had. The central idea -- error budgets -- is counterintuitive: if your service has been too reliable, you shoul…
Here is the most counterintuitive sentence in this course: if your app has not had an outage in months, you might be shipping too slowly.
Work through each item before the checkpoint.
Build the complete delivery pipeline for a Rails app in GitHub Actions: tests in parallel, Rubocop, Brakeman, a Docker build pushed to a registry, automatic deploy to staging, and…
4 lessons. Read in order; spiral back when you need to. By the end you'll have used the core ideas twice — once on the abstract, once on something you'll meet at work next week.