Manager of Site Reliability Engineering

5 days ago

Vancouver BC, Greater Vancouver Regional District, BC; British Columbia, Canada LayerZero Labs Ltd. Full-time

  • At LayerZero, our Site Reliability Engineering (SRE) team is at the intersection of software and systems engineering, dedicated to crafting and maintaining large-scale, resilient systems
  • Our goal is to ensure that all LayerZero services — ranging from critical internal systems to those external users interact with — are reliable, meet the uptime expectations of our users, and continuously evolve at a swift pace. Our SRE professionals will monitor our system’s capacity and performance to uphold these standards
  • As Manager of SRE, you’ll lead a team of engineers responsible for the reliability, performance, and scalability of our blockchain node infrastructure and platform services — while staying technically sharp enough to guide architecture decisions and jump into critical incidents
  • You’ll balance people leadership with hands-on technical judgment, shaping how the team works, grows, and scales alongside LayerZero
  • Lead and develop a team of SREs — setting technical direction, growth plans, and performance expectations
  • Own the reliability strategy for blockchain node infrastructure across a variety of DLTs, including SLOs, capacity planning, and incident response
  • Partner with Engineering leadership and Product/Platform teams to align reliability investments with business priorities
  • Drive infrastructure-as-code practices, with a focus on Kubernetes and Helm at scale
  • Establish and continuously improve on-call structure, incident detection/triage automation, and postmortem culture
  • Stay hands-on: review designs, dig into complex incidents, and set the technical bar for the team
  • 6+ years in SRE, DevOps, or infrastructure engineering, including 2+ years directly managing or leading a technical team
  • Deep familiarity with blockchain node infrastructure (validator/full/archive nodes, RPC optimization, etc.)
  • Bachelor’s degree in Computer Science, similar technical field of study, or equivalent practical experience
  • 3+ years running Kubernetes in production, including Helm chart authoring at scale
  • Advanced knowledge of Unix/Linux internals and distributed systems / high-availability design
  • Track record building or scaling an on-call/incident response process
  • Excellent communication skills — able to represent the team to leadership and hire/retain strong engineers
  • Strong proficiency in TypeScript or Golang, with the judgment to know when to write code vs. delegate
#J-18808-Ljbffr