Expert Reliability Engineering Role at IBM
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
By continuing, you agree to our Terms & Privacy Policy.
At IBM Software, take on an Expert Reliability Engineering role focused on incident command in multi-cloud settings. Drive improvements to enhance system reliability and incident response capabilities. This hybrid role combines deep engineering work with strategic oversight. You'll engage in hands-on projects such as improving tooling and analyzing failures while also mentoring teams in incident response. Your goal will be to enhance reliability across IBM’s Cloud services. Key Responsibilities:
- Analyze and design improvements to prevent incidents
- Oversee Rootly configurations and related integrations
- Maintain SLO/SLA standards for incident management
- Edit customer-facing incident documentation for clarity
- Coach teams through incident post-mortems and training
- 10+ years in SRE or reliability-focused roles
- Cloud expertise in major platforms like AWS, GCP, Azure
- Familiarity with incident management tools like PagerDuty
- In-depth knowledge of distributed systems performance
- Strong written and verbal communication skills