Senior MLOps Engineer | Ingénieur·e MLOps senior

19 hours ago

Montreal, Quebec, Canada Jesta I.S. Inc. Full-time

Company Overview

Jesta I.S. builds enterprise retail technology used by apparel and footwear brands with complex, multi-site operations. Our data environment spans ERP and cloud platforms, and our engineering culture is hands-on, pragmatic, and fast-moving.

You’ll work in a production environment integrating Oracle, Snowflake, AWS, and Azure, supported by strong security standards, modern CI/CD practices, and close collaboration across Data Science, Engineering, Frontend, and Product teams.

Position Summary

We are looking for a Senior MLOps Engineer to design, build, and maintain the data and machine learning pipelines that power our AI and analytics platforms.

This is a deeply hands-on engineering role responsible for the full ML lifecycle, from data ingestion and transformation through model training, deployment, monitoring, retraining, and rollback.

You will bridge data engineering, ML automation, infrastructure, observability, and application deployment to help build scalable, secure, multi-tenant AI infrastructure with a strong focus on reliability, performance, and cost-efficient design.

Responsibilities

  • Build and automate ML pipelines for data preparation, training, inference, monitoring, and retraining.
  • Develop production data flows across Oracle ERP, Snowflake, AWS, and Azure environments.
  • Create reusable Kedro pipelines and manage scalable workloads through AWS Batch, EKS, Karpenter, Kueue, and Fargate.
  • Implement MLflow-based experiment tracking, model versioning, lineage, quality gates, staged promotion, and rollback.
  • Capture reproducible run manifests, validate prediction completeness, and support safe partial-run recovery.
  • Provision secure, multi-tenant cloud infrastructure using Terraform or OpenTofu.
  • Implement CI/CD workflows using Azure DevOps and GitHub Actions, including testing, scanning, immutable images, and rollback strategies.
  • Build observability for run success, completeness, freshness, duration, drift, failures, and infrastructure cost.
  • Maintain Dockerized, Kubernetes-native environments using ECR and EKS, with appropriately sized compute and memory resources.
  • Deploy secure React and Python ML applications across AWS and Azure using private networking, MFA, RBAC, encryption, and least-privilege access.
  • Collaborate with Data Scientists and Product stakeholders to operationalize models, improve performance, and address reliability gaps.

Technical Environment

  • Languages & Frameworks: Python (pandas, Polars, boto3, joblib, LightGBM/XGBoost), SQL, JavaScript/React
  • Data Engineering: AWS DMS, Athena, Snowflake, Oracle
  • Pipeline & Orchestration: Kedro, EventBridge, AWS Batch, Amazon EKS, Karpenter, Kueue, Fargate
  • MLOps: MLflow, Docker, ECR, Azure DevOps, GitHub Actions
  • Infrastructure & Observability: Terraform/OpenTofu, CloudWatch, Prometheus, Grafana, structured logging, drift monitoring
  • Cloud & Deployment: AWS (EC2, S3, RDS, Batch, EKS, Fargate, ECR, EventBridge, Lambda), Azure integration and parallel deployment
  • Security: AWS IAM, Cognito, RBAC, MFA, Secrets Manager, PrivateLink, encryption, and network access controls

Qualifications

Education & Professional Experience

  • Bachelor’s or Master’s degree in Computer Science, Machine Learning, or a related field.
  • Minimum of 5 years of full-time professional experience, excluding internships and academic training, in ML Engineering, MLOps, or data-pipeline development.
  • 7+ years of relevant professional experience is preferred.
  • Proven ability to design, build, and automate production-scale, end-to-end ML pipelines in cloud environments.

Technical Expertise

  • Strong Python and SQL skills, including complex querying against large datasets.
  • Hands-on experience integrating Oracle and Snowflake with production ML systems.
  • Proficiency with Terraform or equivalent Infrastructure as Code (IaC) tooling.
  • Experience with containerized application deployment, CI/CD, and workflow orchestration.
  • Experience with Kubernetes/EKS and cloud-native infrastructure.
  • Experience with MLflow or equivalent MLOps tooling.
  • Understanding of model lifecycle practices, including versioning, lineage, quality validation, staged promotion, monitoring, and rollback.
  • Experience working with cloud platforms, preferably AWS and Azure.

Skills & Abilities

  • Strong ownership and hands-on engineering mindset, from architecture through production.
  • Analytical and performance-focused approach to solving complex technical and operational problems.
  • Ability to balance scalability, cost, security, reliability, and maintainability when designing solutions.<