DevOps, Kubernetes, and Site Reliability Engineer

2 days ago

Montreal administrative region, QC, Canada Randstad Enterprise Full-time $90,000 - $130,000 Contract

We are looking for an experienced DevOps, Kubernetes, and Site Reliability Engineer with a minimum of five years of relevant industry experience. Experience within financial services or another highly regulated technology environment is preferred. The position requires strong hands-on experience with Kubernetes, Docker/Podman, Linux, networking, CI/CD pipelines, GitHub, Python, and shell scripting. The candidate should understand modern DevOps and SRE practices and be comfortable supporting production systems, troubleshooting complex application and infrastructure issues, and automating repetitive operational tasks.The ideal candidate will be passionate about automation, production reliability, and continuous improvement. The candidate should be organized, disciplined, detail-oriented, self-motivated, collaborative, and focused on delivering measurable engineering outcomes.

Location: Montreal (day 1 onboarding / onsite presence required 3x/week)

QUALIFICATIONS

Responsibilities

  • Design, build, maintain, and enhance CI/CD pipelines and supporting build, test, release, and deployment infrastructure.
  • Develop and maintain automated deployment solutions for applications running in Kubernetes and containerized environments.
  • Partner with application development teams to improve build, test, deployment, and release processes.
  • Support the deployment of applications, configuration changes, patches, and platform upgrades across development, testing, and production environments.
  • Build and maintain Kubernetes deployment artifacts, including YAML configuration, Helm charts, or equivalent packaging and configuration mechanisms.
  • Develop automation using Python, Linux shell scripting, and related tools.
  • Manage and improve GitHub repositories, branching strategies, pull-request workflows, access controls, and automated repository processes.
  • Integrate automated testing, code-quality validation, security scanning, dependency checks, and linting into CI/CD pipelines.
  • Troubleshoot application, infrastructure, container, Kubernetes, network, and deployment issues in complex environments.
  • Investigate production incidents, identify root causes, and implement preventive or corrective engineering solutions.
  • Apply SRE principles to improve system reliability, availability, scalability, observability, and operational readiness.
  • Define and improve monitoring, alerting, dashboards, operational metrics, and production support procedures.
  • Automate routine operational activities to reduce manual effort and operational risk.
  • Collaborate with infrastructure, network, database, cybersecurity, release management, and application teams.
  • Create and maintain technical documentation, operational runbooks, deployment procedures, and troubleshooting guides.
  • Participate in design reviews, production-readiness reviews, incident reviews, and continuous-improvement initiatives.
  • Participate in an after-hours production support and on-call rotation when required.

Required Skills

  • Minimum of 5 years of relevant experience in DevOps, SRE, production engineering, platform engineering, infrastructure engineering, or a related discipline.
  • Strong hands-on experience with Kubernetes, including application deployment, configuration, troubleshooting, scaling, services, ingress, secrets, and operational support.
  • Strong hands-on experience with Docker/Podman and containerized application environments.
  • Strong Linux and UNIX system administration and troubleshooting skills.
  • Strong experience with Linux shell scripting, such as Bash or KornShell.
  • Hands-on programming and automation experience using Python or a comparable language.
  • Strong understanding of CI/CD concepts, software delivery lifecycles, release automation, and deployment strategies.
  • Hands-on experience with CI/CD and artifact-management tools such as Jenkins and Artifactory, or equivalent platforms.
  • Hands-on experience with Git and GitHub, including repository management, pull requests, branching strategies, release workflows, and automated checks.
  • Experience developing automation using YAML, Ansible, or an equivalent automation framework.
  • Strong understanding of networking concepts, including DNS, TCP/IP, HTTP/HTTPS, TLS, proxies, firewalls, routing, load balancing, and network troubleshooting.
  • Experience supporting applications and resolving production issues in a fast-paced environment.
  • Understanding of SRE practices, including monitoring, incident response, root-cause analysis, service reliability, operational readiness, and continuous improvement.
  • Experience integrating code-quality tools, security scanning, automated testing, and policy controls into CI/CD pipelines.
  • Ability to troubleshoot