Principal AI Cloud Engineer

2 weeks ago


Toronto, Ontario, Canada BMO Full time $103,000 - $192,000 per year

The Team
We accelerate BMO's AI journey by building enterprise-grade, cloud-native AI solutions. Our team combines engineering excellence with cutting-edge AI to deliver scalable, secure, and responsible solutions that power business innovation across the bank. We enable and accelerate our partners on their AI journeys across the enterprise, helping teams across BMO unlock value at scale. We support one another in times of need and take pride in our work. We are engineers, AI practitioners, platform builders, thought leaders, multipliers, and coders. Above all, we are a global team of diverse individuals who enjoy working together to create smart, secure, and scalable solutions that make an impact across the enterprise. Our ambition is bold: deploy our capital and resources to their highest and most profitable use through a digital-first operating model, powered by data and AI-driven decisions.

The Impact
As a Principal Cloud AI Engineer, you are a hands-on technical developer who designs, builds, and scales cloud-native AI solutions and products. You help set engineering standards, establish patterns, mentor senior engineers, and partner with multiple teams to deliver resilient, governed, and cost-efficient AI at enterprise scale. You'll help shape and evolve our AI cloud strategy from model serving and LLMOps to security, observability, and compliance so teams across the bank can innovate safely and rapidly.

You will advance BMO's Digital First strategy by:

  • Defining reference and production-grade solutions for AI/GenAI on cloud (Azure preferred; multi-cloud aware).
  • Building reusable, secure, and observable components (APIs, SDKs, microservices, pipelines).
  • Operationalizing LLMs and RAG with strong controls and Responsible AI guardrails.
  • Driving platform roadmaps that enable faster delivery, lower risk, and measurable business outcomes.

What's In It For You

  • Influence the technical direction of enterprise AI and the platform primitives others build on.
  • Ship high-impact systems used across many business lines and products.
  • Work across the full stack: cloud infra, data/feature pipelines, model serving, LLMOps, and DevSecOps.
  • Partner with a leadership team invested in your growth and thought leadership.

Responsibilities
Infrastructure & Platform Builder

  • Design, build, and operate cloud-native AI infrastructure for ML/GenAI workloads:

  • Compute: GPU/CPU clusters, autoscaling, spot instance strategies

  • Networking: Azure VNet, Private Link, peering, multi-region HA/DR
  • Storage & Databases: high-performance data lakes (e.g., Azure Data Lake Storage), relational DBs, vector DBs (FAISS, Milvus, Pinecone, pgvector)
  • Security: IAM, Key Vault-backed secrets management, encryption, policy-as-code

  • Implement observability and reliability for AI infra:

  • Metrics (latency, throughput, GPU utilization, cost)

  • Logging/tracing (OpenTelemetry), SLOs/SLIs for infra services

  • Build CI/CD and GitOps pipelines for infrastructure-as-code (Terraform/Bicep) and AI platform components

  • Drive FinOps for AI infra: GPU rightsizing, caching, inference optimization, cost governance

Application & Service Enablement

  • Enable frontend and backend services for AI platforms:

  • Secure APIs, microservices, and event-driven architectures

  • Integration with custom model runtimes (TensorRT-LLM, vLLM, Triton/KServe)

  • Provide infrastructure support for RAG systems: embeddings, chunking, retrieval pipelines

  • Ensure scalable serving infrastructure for LLMs and ML models with caching and token optimization

Strategy & Architecture

  • Define and evolve AI infrastructure reference architecture for cloud (Azure preferred):

  • Container orchestration (Kubernetes), service mesh, ingress

  • Serverless/event-driven patterns for AI pipelines
  • Multi-region, HA/DR, compliance-ready designs

  • Establish standards and best practices for containerization, IaC, and secure networking for AI systems

Security, Risk & Governance

  • Implement defense-in-depth for AI infra:

  • IAM least privilege, private networking, KMS/Key Vault, SBOM, image signing

  • Ensure compliance and Responsible AI controls at infra level:

  • Data residency, encryption, lineage, audit readiness

Delivery & Operations

  • Lead infrastructure discovery and solution design with stakeholders
  • Operate platforms with SRE principles: error budgets, incident response, chaos testing
  • Mentor engineers; create reusable IaC modules, templates, and golden paths

Must-Have Qualifications

  • Bachelor's/Master's/PhD in CS, Engineering, or related field
  • 7+ years building large-scale distributed cloud infrastructure
  • 5+ years hands-on with Azure (preferred); AWS/GCP nice to have
  • Proven experience with AI/ML infra: GPU clusters, Kubernetes, CI/CD, observability
  • Strong in IaC (Terraform/Bicep), Kubernetes, networking, security
  • Expertise in cloud-native patterns: containers, service mesh, serverless
  • Familiarity with MLOps/LLMOps infra: model serving, feature stores, vector DBs
  • Programming in Python (infra automation) and one of Go/TypeScript for tooling
  • Understanding of frontend/backend integration for AI services
  • Familiarity with MLOps/LLMOps infra: model serving, feature stores, vector DBs
  • Programming in Python (infra automation) and one of Go/TypeScript for tooling
  • Understanding of frontend/backend integration for AI services

Nice-to-Have

  • GPU optimization (CUDA/NCCL, TensorRT-LLM)
  • Observability tools (Prometheus, Grafana, OpenTelemetry)
  • Event streaming (Kafka/Azure Event Hubs), real-time systems
  • Experience with AI platform products (Azure ML, MLflow, KServe, Hugging Face)

Tech Stack

  • Cloud & Infra: Azure (AKS, Functions, Event Hubs, Key Vault), Terraform/Bicep, GitHub Actions/Azure DevOps
  • AI Infra: Kubernetes, KServe/Triton, vLLM, TensorRT-LLM, Ray, Spark
  • Ops: Prometheus, Grafana, OpenTelemetry, ArgoCD, OPA
  • Data: Feature stores (Feast), vector DBs (FAISS, Milvus, Pinecone), relational DBs
  • App Layer: APIs, microservices, frontend/backend integration for AI systems

Success Metrics

  • Reliability & Performance: SLOs met for infra services, GPU utilization optimized
  • Security & Compliance: Zero critical findings, auditable infra
  • Cost Efficiency: Reduced GPU/infra spend via FinOps strategies
  • Developer Velocity: Faster provisioning and deployment of AI infra
  • Technical Leadership: Influence on infra standards, mentorship, reusable patterns

Salary:
$103, $192,000.00

Pay Type:
Salaried

The above represents BMO Financial Group's pay range and type.

Salaries will vary based on factors such as location, skills, experience, education, and qualifications for the role, and may include a commission structure. Salaries for part-time roles will be pro-rated based on number of hours regularly worked. For commission roles, the salary listed above represents BMO Financial Group's expected target for the first year in this position.

BMO Financial Group's total compensation package will vary based on the pay type of the position and may include performance-based incentives, discretionary bonuses, as well as other perks and rewards. BMO also offers health insurance, tuition reimbursement, accident and life insurance, and retirement savings plans. To view more details of our benefits, please visit:

About Us
At BMO we are driven by a shared Purpose: Boldly Grow the Good in business and life. It calls on us to create lasting, positive change for our customers, our communities and our people. By working together, innovating and pushing boundaries, we transform lives and businesses, and power economic growth around the world.

As a member of the BMO team you are valued, respected and heard, and you have more ways to grow and make an impact. We strive to help you make an impact from day one – for yourself and our customers. We'll support you with the tools and resources you need to reach new milestones, as you help our customers reach theirs. From in-depth training and coaching, to manager support and network-building opportunities, we'll help you gain valuable experience, and broaden your skillset.

To find out more visit us at

BMO is committed to an inclusive, equitable and accessible workplace. By learning from each other's differences, we gain strength through our people and our perspectives. Accommodations are available on request for candidates taking part in all aspects of the selection process. To request accommodation, please contact your recruiter.

Note to Recruiters: BMO does not accept unsolicited resumes from any source other than directly from a candidate. Any unsolicited resumes sent to BMO, directly or indirectly, will be considered BMO property. BMO will not pay a fee for any placement resulting from the receipt of an unsolicited resume. A recruiting agency must first have a valid, written and fully executed agency agreement contract for service to submit resumes.



  • Toronto, Ontario, Canada BMO Full time US$103,000 - US$192,000 per year

    The TeamWe accelerate BMO's AI journey by building enterprise-grade, cloud-native AI solutions. Our team combines engineering excellence with cutting-edge AI to deliver scalable, secure, and responsible solutions that power business innovation across the bank. We enable and accelerate our partners on their AI journeys across the enterprise, helping teams...


  • Toronto, Ontario, Canada BMO Full time US$103,200 - US$192,000

    Application Deadline:11/29/2025Address:100 King Street West Job Family Group:Data Analytics & ReportingThe TeamWe accelerate BMO's AI journey by building enterprise-grade, cloud-native AI solutions. Our team combines engineering excellence with cutting-edge AI to deliver scalable, secure, and responsible solutions that power business innovation across...


  • Toronto, Ontario, Canada BMO Full time US$103,200 - US$192,000

    Application Deadline:11/29/2025Address:100 King Street West Job Family Group:Data Analytics & ReportingThe TeamWe accelerate BMO's AI journey by building enterprise-grade, cloud-native AI solutions. Our team combines engineering excellence with cutting-edge AI to deliver scalable, secure, and responsible solutions that power business innovation across...


  • Toronto, Ontario, Canada Bank of Montreal Full time US$103,200 - US$192,000

    Application Deadline:10/30/2025Address:100 King Street West Job Family Group:Data Analytics & ReportingWe are back in office 2-4 days/week This role is not remote/virtual.The TeamWe accelerate BMO's AI journey by building enterprise-grade, cloud-native AI solutions. Our team combines engineering excellence with cutting-edge AI to deliver scalable,...


  • Toronto, Ontario, Canada Menten AI, Inc. Full time $120,000 - $180,000 per year

    About Menten AI Menten AI is revolutionizing drug discovery by harnessing the power of generative AI to design cyclic peptide therapeutics. Our innovative platform targets challenging drug targets that lie beyond the reach of traditional small molecules and biologics. By creating potent macrocycles with optimal drug properties, including oral...


  • Toronto, Ontario, Canada Royal Bank of Canada Full time $130,000 - $220,000 per year

    Job DescriptionWhat's the opportunity?At RBC, you'll be joining a team of leading platform engineers and security specialists focused on implementing and optimizing our enterprise GenAI platform infrastructure. You will have access to cutting-edge GPU technologies, multi-cloud environments, and the computational resources to support novel AI/ML workload...


  • Toronto, Ontario, Canada KData AI Full time $120,000 - $180,000 per year

    Job Title: Agentic AI Platform EngineerClient: Banking ClientLocation:Downtown Toronto (On-site 4 days/week)Position Type:6-Month Contract (Renewable)Domain: Generative AI, Agentic Systems, AI Platform EngineeringRate: up to $90Role SummaryA leading Banking Client is building a next-generationAgentic AI Platformto enable scalable, enterprise-wide deployment...


  • Toronto, Ontario, Canada Revvity Full time $120,000 - $180,000 per year

    Revvity Signals makes market-leading software that empowers scientists across Life Science R&D, clinical research, and specialty chemicals to make better medicines and products, faster. Our flagship offering is the SaaS Signals Research Platform that provides knowledge capture, collaboration, analysis and visualization tools across the full depth and breadth...


  • Toronto, Ontario, Canada Boson AI Full time $120,000 - $180,000 per year

    About The Role We're seeking an experienced Network Engineer to design, build, and optimize the high-performance networking infrastructure powering our AI/ML operations in Toronto. You'll work at the cutting edge of network technology—managing InfiniBand and ultra-high-speed Ethernet fabrics that connect NVIDIA H100 and A100 GPUs, over 20PB of Ceph...


  • Toronto, Ontario, Canada Boson AI Full time US$150,000 - US$250,000

    About The RoleWe're seeking an experienced Network Engineer to design, build, and optimize the high-performance networking infrastructure powering our AI/ML operations in Toronto. You'll work at the cutting edge of network technology—managing InfiniBand and ultra-high-speed Ethernet fabrics that connect NVIDIA H100 and A100 GPUs, over 20PB of Ceph storage,...