Staff Site Reliability Engineer

4 weeks ago


Montreal Quebec GF, CA Lightspeed Full time

We’re looking for a Staff Site Reliability Engineer to join our NuOrder by Lightspeed team. NuORDER by Lightspeed builds software solutions that help merchants grow the size and profitability of their business. You'll join a team responsible for supporting the group in cross-cutting concerns, such as cloud infrastructure, reliability and incident management, data warehousing and analytics, cost transparency and efficiency, and much more. You will also be supporting our growing Dev teams with the infrastructure and tools needed to continue scaling. You will build and support multi-region infrastructures and networks, and help run our products in a reliable, efficient, and secure manner by implementing, advising, and advocating well-known DevOps principles.

Role:

  • Design, build, and maintain robust infrastructure on GCP, leveraging cloud-native technologies such as GKE, Cloud SQL, BigQuery, etc.
  • Develop and manage CI/CD pipelines for efficient deployment and release using various technologies (GitLab, GitHub, Helm, Terraform, etc.).
  • Work closely with development teams to provide tools and practices for monitoring software health in production, defining and measuring reliability metrics (SLI, SLO), and managing error budgets.
  • Build platform solutions and apply software engineering principles to improve software reliability and accelerate delivery.
  • Support the incident management process and conduct post-mortem analysis to prevent future outages.
  • Mentor junior SREs and developers, offering guidance on best practices in cloud architecture, data management, and software development.
  • Manage infrastructure changes through infrastructure as code (IaC) using Terraform.
  • Participate in the on-call rotation.
  • Stay current with industry trends and emerging technologies, advocating for the adoption of new technologies and practices to improve product quality and team efficiency.

What you need to bring:

  • Bachelor’s degree in Computer Science, Engineering, or equivalent real-world experience.
  • 6+ years of experience in site reliability engineering, systems administration, and/or software engineering.
  • Expertise in container orchestration platforms, specifically Kubernetes.
  • Strong understanding of both relational (e.g., PostgreSQL, MySQL) and NoSQL databases (e.g., MongoDB, Cassandra, Redis).
  • Familiarity with network protocols and IP networking, along with experience in network troubleshooting.
  • Proficiency in at least one programming language such as Bash, Python, Go, etc.
  • Proven track record of managing large-scale infrastructure in cloud environments like Google Cloud, AWS, or Azure.
  • Experience with monitoring tools (e.g., Prometheus, Grafana, Datadog) and logging solutions (e.g., ELK stack).
  • Strong understanding of security best practices.
  • Excellent problem-solving skills and the ability to work under pressure to troubleshoot and resolve complex issues.
  • Excellent communication skills for effective collaboration with cross-functional teams.
  • A keen eagerness to learn and embrace challenges.
#J-18808-Ljbffr

  • Montreal, Quebec, G4F, CA Lightspeed Commerce Full time

    Hi there! Thanks for stopping by We're looking for a Staff Site Reliability Engineer to join our NuOrder by Lightspeed team in North America. NuORDER by Lightspeed builds software solutions that help merchants grow the size and profitability of their business. You'll join a team responsible for supporting the group in cross-cutting concerns, such as...


  • Montreal, Quebec, G4F, CA Lightspeed Full time

    Hi there! Thanks for stopping by We’re looking for a Staff Site Reliability Engineer to join our NuOrder by Lightspeed team in North America. NuORDER by Lightspeed builds software solutions that help merchants grow the size and profitability of their business. You'll join a team responsible for supporting the group in cross-cutting concerns, such as...


  • Montreal, Quebec, G4F, CA Lightspeed Full time

    Hi there! Thanks for stopping by. Are you actively looking for a new opportunity? Or just checking the market? Well… you might just be in the right place! We’re looking for a Staff Site Reliability Engineer to join our NuOrder by Lightspeed team in North America. NuORDER by Lightspeed builds software solutions that help merchants grow the size and the...


  • Montreal, Quebec, G4F, CA Lightspeed Full time

    Hi there! Thanks for stopping by We’re looking for a Staff Site Reliability Engineer to join our NuOrder by Lightspeed team in North America. NuORDER by Lightspeed builds software solutions that help merchants grow the size and the profitability of their business. You'll join a team responsible for supporting the group in cross-cutting concerns, such...


  • Montreal, Quebec, G4F, CA Lightspeed Full time

    We’re looking for a Staff Site Reliability Engineer to join our NuOrder by Lightspeed team in North America. NuORDER by Lightspeed builds software solutions that help merchants grow the size and the profitability of their business. You'll join a team responsible for supporting the group in cross-cutting concerns, such as cloud infrastructure,...


  • Montreal, Quebec, G4F, CA SAP Full time

    We help the world run better At SAP, we enable you to bring out your best. Our company culture is focused on collaboration and a shared passion to help the world run better. We offer a highly collaborative, caring team environment with a strong focus on learning and development, recognition for your individual contributions, and a variety of benefit...


  • Montreal, Quebec, G4F, CA Socotra, Inc. Full time

    At Lyft, our mission is to improve people’s lives with the world’s best transportation. Imagine cities where streets are safe, communities thrive, and personal cars are a thing of the past. We envision a future where shared and active transportation modes are the norm, fostering vibrant, connected neighborhoods. As a leader in micromobility, Lyft powers...


  • Montreal, Quebec, G4F, CA SAP SE Full time

    We help the world run betterAt SAP, we enable you to bring out your best. Our company culture is focused on collaboration and a shared passion to help the world run better. We focus every day on building the foundation for tomorrow and creating a workplace that embraces differences, values flexibility, and is aligned to our purpose-driven and future-focused...


  • Montreal, Quebec, G4F, CA Lightspeed Full time

    Hi there! Thanks for stopping by Are you actively looking for a new opportunity? Or just checking the market? Well… you might just be in the right place! We’re looking for a Principal Site Reliability Engineer to join our NuOrder by Lightspeed team in North America. NuORDER by Lightspeed builds software solutions that help merchants grow the size and...


  • Montreal, Quebec, G4F, CA Lightspeed Full time

    ```html Hi there! Thanks for stopping by. Are you actively looking for a new opportunity? Or just checking the market? Well… you might just be in the right place! We’re looking for a Principal Site Reliability Engineer to join our NuOrder by Lightspeed team in North America. NuORDER by Lightspeed builds software solutions that help merchants grow the...


  • Montreal, Quebec, G4F, CA Lightspeed Full time

    Hi there! Thanks for stopping by Are you actively looking for a new opportunity? Or just checking the market? Well… you might just be in the right place! We’re looking for a Principal Site Reliability Engineer to join our NuOrder by Lightspeed team in North America. NuORDER by Lightspeed builds software solutions that help merchants grow the size and...


  • Montreal, Quebec, G4F, CA LanceSoft Full time

    Job Title: Site Reliability Specialist Years of experience: 5+ years Location: Montreal (Office attendance from Day 1 - Hybrid mode)Position Description: The Private Cloud SRE L3 team is part of the Enterprise Computing organization within ***. The team has presence in cities globally and is focused on supporting cloud and container-based platforms for...


  • Montreal, Quebec, G4F, CA Banque Nationale du Canada Full time

    Site Reliability Engineering Developer SREHybridJob Number: 21829Category: Senior ProfessionalStatus: PermanentSchedule: Full-TimeArea of Interest: Information technologyA career in technology at National Bank means participating in the transformation to have a direct impact on the client. As a System Reliability Specialist, you will be responsible for...


  • Montreal, Quebec, G4F, CA Banque Nationale du Canada Part time

    Site Reliability Engineering Developper SRE Hybrid Job Number 21241 Category Senior Professional Status: Permanent Type of Contract Permanent Schedule: Full-Time Full Time / Part Time? Full-Time 06-Jun-2024 City Montreal Province/State Area of Interest: Information technology A career in technology at National Bank means participating in...


  • Montreal, Quebec, G4F, CA Axelon Services Full time

    Job Title: Private Cloud Site Reliability Specialist 12 Months Contract Years of experience: 5+ years Location: Montreal (Office attendance from Day 1 - Hybrid mode)Position Description: The Private Cloud SRE L3 team is part of the Enterprise Computing organization within Brokerage. The team has a presence in cities globally and is focused on supporting...


  • Montreal, Quebec, G4F, CA Unity Full time

    The opportunityAt Unity, we are the world's leading platform for creating and operating real-time 3D (RT3D) content. We deeply understand the critical importance of reliability in today's fast-paced digital world. Our infrastructure, systems, and applications play a pivotal role in delivering seamless experiences to our customers.Our team of Site...


  • Montreal, Quebec, G4F, CA Pharmascience Inc. Full time

    Proud to be at the forefront of our industry since 1983, Pharmascience is a leader in generic medicines. We’re a Canadian company with global reach that has never lost sight of the human touch. Pharmascience, headquartered in Montreal, was named one of the top 300employers in 2022. Our 1,400employees are committed to the quality of our products, research...


  • Montreal, Quebec, G4F, CA Electronic Arts Full time

    Pour visualiser la description de poste en français, veuillez sélectionner le français, "Select Language" dans le menu déroulant au haut de la page.Player Quality Insights (PQI) is a world class team, delivering industry leading insights to create great experiences that players want to play. As a Site Reliability Engineer, you will report to the PQI Labs...


  • Montreal, Quebec, G4F, CA Ennuviz Full time

    Job Information Industry: IT Services Work Experience: 3-5 years City: Downtown Montreal Northeast State/Province: Quebec Country: Canada Zip/Postal Code: H2Z About Us We are a client-focused digital transformation expert with 50+ years of technology & industry experience. Our customized solutions empower organizations to streamline, optimize operating...


  • Montreal, Quebec, G4F, CA Toparo Full time

    We seek a Staff Software Engineer for one of our remote-first clients based in Montreal. This client is building a SaaS solution in the crypto space. In this role, you will use your technical expertise to manage project priorities, deadlines, and deliverables. You will design, develop, test, deploy, maintain, and enhance software solutions.As a Staff...