Staff Engineer

2 days ago

Calgary AB, Calgary Census Division, AB; Alberta, Canada Lever, Inc. Full-time €141,000 - €240,000

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Engineer - Cloud Platform based in Canada.

This is a senior individual contributor role focused on building and operating highly scalable, secure, and resilient cloud platforms for next-generation broadband services.
You will shape platform architecture, modernization, automation, and reliability strategies across complex cloud and telecom environments.
The role combines deep hands‑on engineering with technical leadership across cloud infrastructure, networking, messaging, and observability.
You will work extensively with Google Cloud, Kubernetes, infrastructure as code, CI/CD, and distributed systems.
Your work will help engineering teams deliver products faster while maintaining carrier‑grade reliability and operational excellence.
You will collaborate across Engineering, SRE, Network Operations, Product, and Customer Success teams in a highly distributed environment.
The position offers significant scope to influence technical direction, mentor engineers, and drive strategic platform initiatives.

Accountabilities

  • Design, architect, and operate highly available, scalable, secure, and resilient cloud platform infrastructure, while driving modernization through cloud-native technologies and automation.

  • Define and implement strategies for platform reliability, observability, disaster recovery, capacity management, and operational excellence.

  • Troubleshoot complex end-to-end service issues across cloud infrastructure, network services, broadband access environments, and customer-facing applications.

  • Design and operate NATS messaging infrastructure and HAProxy-based traffic management, including load balancing, routing, SSL termination, high availability, and performance optimization.

  • Build scalable service communication frameworks using event-driven architectures, service discovery, fault tolerance, and resiliency patterns.

  • Design and manage Google Cloud infrastructure, including GKE, Compute Engine, Cloud Storage, BigQuery, Pub/Sub, Cloud SQL, and Composer/Airflow.

  • Implement Infrastructure as Code using Terraform and Terragrunt, automating provisioning, deployment, compliance, and operational workflows.

  • Develop reusable platform services and self-service capabilities that improve engineering productivity and accelerate application delivery.

  • Design and maintain enterprise‑scale CI/CD pipelines using tools such as Jenkins, GitLab CI, and GitHub Actions.

  • Establish and enhance monitoring, logging, tracing, and alerting using technologies such as Prometheus, Grafana, ELK/OpenSearch, VictoriaMetrics/VictoriaLogs, and cloud monitoring services.

  • Lead root cause analysis and incident response for critical platform and network issues, identifying opportunities for proactive prevention.

  • Partner with Product, Engineering, SRE, Network Operations, technical support, and Customer Success teams to deliver strategic platform initiatives.

  • Develop technical roadmaps, contribute to architecture governance and design reviews, establish engineering standards, and mentor engineers across cloud, networking, reliability, and telecom domains.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Telecommunications, or a related discipline, or equivalent practical experience.

  • 8+ years of experience designing and operating large‑scale distributed systems in production environments.

  • 5+ years of hands‑on experience with public cloud platforms, preferably Google Cloud Platform.

  • Strong hands‑on expertise with Kubernetes and containerized platforms.

  • Deep understanding of networking fundamentals, including TCP/IP, routing and switching, NAT, DNS, DHCP, firewalls, VPN technologies, and load balancing.

  • Practical experience configuring, tuning, troubleshooting, and operating HAProxy in highly available environments.

  • Hands‑on experience with NATS and event‑driven architectures.

  • Experience supp