Reliability Engineer at IBM Cloud
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
Drive proactive reliability at IBM Software, focusing on incident command in multi-cloud environments. Leverage hands-on engineering skills to enhance automation and improve incident management practices. IBM Software seeks a seasoned Reliability Engineer with extensive experience in incident management and SRE. This role involves 75% technical work, including designing reliability improvements, and 25% coaching on incident response. You'll work closely with engineering leaders to elevate reliability across all teams. Key Responsibilities:
- Analyze failure patterns and propose reliability enhancements
- Manage Rootly configurations and integration with key tools
- Define SLO/SLA frameworks and guide reliability investments
- Foster continuous improvement of incident response processes
- Develop and deliver training programs on incident management
- 10+ years in SRE or reliability engineering
- Expertise in AWS, GCP, or Azure
- Proficient with incident management tools like Rootly
- Strong written communication skills
- Experience in large organizations with 500+ engineers