Site Reliability Engineer
Keep systems up at scale: SLOs, observability, incident response and automating toil away.
✦ Guide me on this path with AIRoadmap progress
0% 0 of 6 steps done
How to read the signals
Importance High Market demand Medium Automation risk Low
- 1
Systems & Linux
Deep OS, networking, performance. Operate with confidence.
Importance High Market demand High Automation risk Low - 2
SLOs & error budgets
Define reliability, measure it, spend the budget wisely.
Importance High Market demand Medium Automation risk Low - 3
Observability
Metrics, logs, traces, dashboards, alerting that doesn't cry wolf.
Importance High Market demand High Automation risk Low - 4
Incident management
On-call, runbooks, blameless postmortems.
Importance High Market demand High Automation risk Low - 5
Automation
Script away toil; if you do it twice, automate it.
Importance High Market demand High Automation risk Medium - 6
Capacity & resilience
Load testing, chaos, graceful degradation.
Importance Medium Market demand Medium Automation risk Low