Principal Engineer, AI Platform
3 hours ago
Toronto, Ontario, Canada
Bank of Montreal
Full-time
Free with email or Google
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
Free with email or Google
Application Deadline:10/29/2026Address:33 Dundas Street WestJob Family Group:TechnologyPrincipal Engineer, AI Platform & FabricsDescriptionBMO is building the platform capabilities that make enterprise AI safe, governed, and scalable. We are seeking experienced Principal/Senior engineers to build and operate the core infrastructure that governs how AI runs at BMO: the AI Gateway, Policy Engine, Identity Fabric, AI Registry, Guardrails Runtime, and AI Observability.
This is a build-and-run engineering role. You will own capabilities end to end; designing, shipping, and operating them in production, including on-call. You will not build the AI models or applications themselves (those are domain-owned); you build the governed platform they run on and the runtime evidence that proves they run within policy, across AWS, Azure, and Microsoft AI surfaces, under OSFI and OCC expectations.
You are a hands-on engineer who has built shared platform services at scale, cares deeply about operability, latency, and correctness, and understands that in a regulated bank the infrastructure must produce its own evidence . You are energized by taking real engineering assets that includes an existing developer portal, an AI registry, a body of policy-as-code, and gateway integrations, and hardening, scaling, and governing them into enterprise-grade platform capabilities. You raise the technical bar for those around you and mentor as you build.
What You'll Build & OperateDepending on your specialization, you will own one or more of the following capability areas:Enterprise Control PlanePortal & Registry — a federated AI Registry (agents, models, tools, channels, evaluations across 16+ asset types) with self-service onboarding and lifecycle workflows; federation with external registries (Agent 365, AgentCore, MLflow).
Policy Engine — policy-as-code infrastructure (Cedar/OPA), a policy compilation pipeline, GitOps-based domain-scoped bundle distribution, risk-tiered approval workflows, and a policy simulation sandbox.
Observability & Audit — a multi-pipeline telemetry architecture (operational + security + compliance), OpenTelemetry GenAI conventions, cross-pipeline trace correlation, lineage-stamped traces, and a 7-year tamper-evident audit lake producing regulator-ready evidence.
Governance & Lifecycle — certification workflows, automated compliance scoring, decommission governance, and evidence generation for architecture and model-risk review.
Domain OrchestrationGateway Runtime — domain-hub deployment across AWS and Azure; an inline enforcement engine performing request-time policy evaluation, routing, residency, budget/quota, and circuit breaking within strict tiered latency budgets (Fast <5ms / Standard <25ms / Full <50ms); cross-region failover.
Guardrails Runtime — a multi-stage safety pipeline (input moderation prompt-injection defense PII output validation hallucination detection policy enforcement) with bilingual EN/FR parity and behavioral guardrails for agentic workloads (goal hijacking, intent drift, excessive agency).
Identity Fabric — workload identity for AI (SPIFFE/SPIRE), token-exchange bridging, per-domain trust boundaries, Entra Agent ID integration, on-behalf-of identity propagation, and cross-cloud token federation with zero-trust attestation.
What You'll Do
Own capabilities end to end — design, implement, test, ship, and operate production infrastructure, including on-call ownership of what you build (no separate run team).
Engineer for operability and defensibility from day one — instrumentation, SLOs, latency budgets, failure modes, and runtime evidence built in, not bolted on.
Build the APIs, SD'able interfaces, and integrations through which domains, DevOps pipelines, and enterprise systems consume platform capabilities.
Implement policy enforcement, guardrails, identity attestation, and audit as first-class engineering concerns — correct, performant, and provable.
Ensure every capability produces runtime evidence connecting AI activity to policy enforcement, identity, and lineage for model-risk and regulatory review (OSFI E-23, OCC).
Assess emerging AI infrastructure, foundation-model access patterns, and standards; make deliberate, cost-aware engineering choices.
Mentor and raise the bar — set engineering standards, review designs and code, and grow depth across the team.
Partner closely with AI Developer Experience (so domains can consume what you build), AI Security and AI SDLC (embedded specializations), and the Senior AI Architect (architectural coherence).
Education & ExperienceBachelor's degree in Computer Science, Software Engineering, or a related technical discipline (Master's preferred).8+ years of software/platform engineering experience (Principal), or 5+ years (Senior), with substantial time building and operating shared platform services at enterprise scale.
Demonstrated experience operating production infrastructure with real SLOs and on-call ownership, ideally in a regulated industry (financial services
This is a build-and-run engineering role. You will own capabilities end to end; designing, shipping, and operating them in production, including on-call. You will not build the AI models or applications themselves (those are domain-owned); you build the governed platform they run on and the runtime evidence that proves they run within policy, across AWS, Azure, and Microsoft AI surfaces, under OSFI and OCC expectations.
You are a hands-on engineer who has built shared platform services at scale, cares deeply about operability, latency, and correctness, and understands that in a regulated bank the infrastructure must produce its own evidence . You are energized by taking real engineering assets that includes an existing developer portal, an AI registry, a body of policy-as-code, and gateway integrations, and hardening, scaling, and governing them into enterprise-grade platform capabilities. You raise the technical bar for those around you and mentor as you build.
What You'll Build & OperateDepending on your specialization, you will own one or more of the following capability areas:Enterprise Control PlanePortal & Registry — a federated AI Registry (agents, models, tools, channels, evaluations across 16+ asset types) with self-service onboarding and lifecycle workflows; federation with external registries (Agent 365, AgentCore, MLflow).
Policy Engine — policy-as-code infrastructure (Cedar/OPA), a policy compilation pipeline, GitOps-based domain-scoped bundle distribution, risk-tiered approval workflows, and a policy simulation sandbox.
Observability & Audit — a multi-pipeline telemetry architecture (operational + security + compliance), OpenTelemetry GenAI conventions, cross-pipeline trace correlation, lineage-stamped traces, and a 7-year tamper-evident audit lake producing regulator-ready evidence.
Governance & Lifecycle — certification workflows, automated compliance scoring, decommission governance, and evidence generation for architecture and model-risk review.
Domain OrchestrationGateway Runtime — domain-hub deployment across AWS and Azure; an inline enforcement engine performing request-time policy evaluation, routing, residency, budget/quota, and circuit breaking within strict tiered latency budgets (Fast <5ms / Standard <25ms / Full <50ms); cross-region failover.
Guardrails Runtime — a multi-stage safety pipeline (input moderation prompt-injection defense PII output validation hallucination detection policy enforcement) with bilingual EN/FR parity and behavioral guardrails for agentic workloads (goal hijacking, intent drift, excessive agency).
Identity Fabric — workload identity for AI (SPIFFE/SPIRE), token-exchange bridging, per-domain trust boundaries, Entra Agent ID integration, on-behalf-of identity propagation, and cross-cloud token federation with zero-trust attestation.
What You'll Do
Own capabilities end to end — design, implement, test, ship, and operate production infrastructure, including on-call ownership of what you build (no separate run team).
Engineer for operability and defensibility from day one — instrumentation, SLOs, latency budgets, failure modes, and runtime evidence built in, not bolted on.
Build the APIs, SD'able interfaces, and integrations through which domains, DevOps pipelines, and enterprise systems consume platform capabilities.
Implement policy enforcement, guardrails, identity attestation, and audit as first-class engineering concerns — correct, performant, and provable.
Ensure every capability produces runtime evidence connecting AI activity to policy enforcement, identity, and lineage for model-risk and regulatory review (OSFI E-23, OCC).
Assess emerging AI infrastructure, foundation-model access patterns, and standards; make deliberate, cost-aware engineering choices.
Mentor and raise the bar — set engineering standards, review designs and code, and grow depth across the team.
Partner closely with AI Developer Experience (so domains can consume what you build), AI Security and AI SDLC (embedded specializations), and the Senior AI Architect (architectural coherence).
Education & ExperienceBachelor's degree in Computer Science, Software Engineering, or a related technical discipline (Master's preferred).8+ years of software/platform engineering experience (Principal), or 5+ years (Senior), with substantial time building and operating shared platform services at enterprise scale.
Demonstrated experience operating production infrastructure with real SLOs and on-call ownership, ideally in a regulated industry (financial services