Lead AI Platform Engineer
14 hours ago
Toronto, ON, Canada
EQ Bank | Canada's Challenger Bank
Full-time
€120,000 - €190,000 Contract
Free with email or Google
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
Free with email or Google
By continuing, you agree to our Terms & Privacy Policy.
Purpose of the Job:
The Lead AI Platform Engineer is accountable for technical leadership, engineering excellence, reliability, operability, and the controlled enablement of the organization’s enterprise AI platforms.
This role provides hands‑on technical leadership across AI platform design, implementation, automation, observability, and production readiness. The incumbent ensures AI platform services and solutions are secure, resilient, observable, supportable, and compliant with enterprise standards for reliability, security, platform management, monitoring, incident coordination, and governance control enforcement.
The incumbent acts as a senior technical lead for platform engineering activities, guiding implementation decisions, establishing engineering patterns, mentoring team members, and partnering with cross‑functional stakeholders to enable the safe and scalable adoption of AI across the enterprise.
Main Activities:
AI Platform Engineering Leadership, Reliability and Operations
- Lead the engineering, configuration, and operation of enterprise AI platforms to ensure availability, performance, resilience, and scalability across environments, integrations, and supporting infrastructure.
- Define and implement platform engineering patterns, standards, reusable components, and operational guardrails that support secure and reliable AI solution delivery.
- Provide technical leadership for platform triage, incident resolution, escalation coordination, and post‑incident reviews to strengthen service stability and resilience.
- Track and report on service reliability indicators, incident trends, engineering risks, and operational performance improvements. AI Platform Enablement, Solution Readiness, Production Readiness and Technical Delivery
- Lead technical enablement of approved AI use cases into non‑production and production environments by ensuring environment readiness, dependency validation, release readiness, operational supportability, and service transition planning.
- Partner with architecture, security, cloud, infrastructure, delivery, and application teams to translate solution requirements into secure, supportable, and scalable platform implementations.
- Guide platform lifecycle management through release coordination, change readiness validation, maintenance planning, capacity planning, and technical risk mitigation.
- Ensure AI platform changes meet defined engineering, operational, security, and control readiness criteria prior to release. Observability, Automation and AI Ops Engineering
- Design, implement, and continuously improve observability capabilities, including telemetry, logging, metrics, traces, dashboards, and alerting required for enterprise AI operations.
- Lead automation initiatives using approved tools and practices to reduce manual effort, improve reliability, and standardize repeatable operational activities.
- Analyze operational data to identify anomalies, recurring issues, root‑cause patterns, performance bottlenecks, and opportunities for proactive service improvement.
- Implement AI Ops use cases such as alert correlation, anomaly detection, forecasting, root‑cause support, knowledge retrieval, and automation of repetitive operational tasks.
- Mentor engineers on observability, automation, troubleshooting, and service reliability practices. Governance, Risk and Control Engineering Execution
- Embed governance, security, privacy, auditability, traceability, and human oversight requirements into AI platform engineering and operational practices.
- Ensure platform implementations align with enterprise security policies, risk controls, compliance requirements, architecture standards, and operational readiness expectations.
- Partner with security, risk, compliance, architecture, and data teams to assess implementation risks, close control gaps, and enable responsible deployment of AI capabilities.
- Maintain documentation and evidence required for audit, governance reviews, production readiness checkpoints, and control validation.
- Identify technical and control risks, recommend mitigation options, and escalates appropriately to relevant governance and risk stakeholders. AI Asset Visibility, Platform Integrity, Operational Integrity and Engineering Standards
- Maintain engineering and operational visibility of AI platform assets required for monitoring, support, ownership, lifecycle management, and cost alignment.
- Validate asset ownership, relationships, configuration integrity, and lifecycle status in collaboration with application, platform, architecture, and infrastructure owners.
- Establish and promote engineering standards, reusable patterns, documentation, and technical practices that improve platform supportability and operational integrity. Knowledge/Skill
Requirements:
- University degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
- 7+ years of experienc
- Lead the engineering, configuration, and operation of enterprise AI platforms to ensure availability, performance, resilience, and scalability across environments, integrations, and supporting infrastructure.
- Define and implement platform engineering patterns, standards, reusable components, and operational guardrails that support secure and reliable AI solution delivery.
- Provide technical leadership for platform triage, incident resolution, escalation coordination, and post‑incident reviews to strengthen service stability and resilience.
- Track and report on service reliability indicators, incident trends, engineering risks, and operational performance improvements. AI Platform Enablement, Solution Readiness, Production Readiness and Technical Delivery
- Lead technical enablement of approved AI use cases into non‑production and production environments by ensuring environment readiness, dependency validation, release readiness, operational supportability, and service transition planning.
- Partner with architecture, security, cloud, infrastructure, delivery, and application teams to translate solution requirements into secure, supportable, and scalable platform implementations.
- Guide platform lifecycle management through release coordination, change readiness validation, maintenance planning, capacity planning, and technical risk mitigation.
- Ensure AI platform changes meet defined engineering, operational, security, and control readiness criteria prior to release. Observability, Automation and AI Ops Engineering
- Design, implement, and continuously improve observability capabilities, including telemetry, logging, metrics, traces, dashboards, and alerting required for enterprise AI operations.
- Lead automation initiatives using approved tools and practices to reduce manual effort, improve reliability, and standardize repeatable operational activities.
- Analyze operational data to identify anomalies, recurring issues, root‑cause patterns, performance bottlenecks, and opportunities for proactive service improvement.
- Implement AI Ops use cases such as alert correlation, anomaly detection, forecasting, root‑cause support, knowledge retrieval, and automation of repetitive operational tasks.
- Mentor engineers on observability, automation, troubleshooting, and service reliability practices. Governance, Risk and Control Engineering Execution
- Embed governance, security, privacy, auditability, traceability, and human oversight requirements into AI platform engineering and operational practices.
- Ensure platform implementations align with enterprise security policies, risk controls, compliance requirements, architecture standards, and operational readiness expectations.
- Partner with security, risk, compliance, architecture, and data teams to assess implementation risks, close control gaps, and enable responsible deployment of AI capabilities.
- Maintain documentation and evidence required for audit, governance reviews, production readiness checkpoints, and control validation.
- Identify technical and control risks, recommend mitigation options, and escalates appropriately to relevant governance and risk stakeholders. AI Asset Visibility, Platform Integrity, Operational Integrity and Engineering Standards
- Maintain engineering and operational visibility of AI platform assets required for monitoring, support, ownership, lifecycle management, and cost alignment.
- Validate asset ownership, relationships, configuration integrity, and lifecycle status in collaboration with application, platform, architecture, and infrastructure owners.
- Establish and promote engineering standards, reusable patterns, documentation, and technical practices that improve platform supportability and operational integrity. Knowledge/Skill
Requirements:
- University degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
- 7+ years of experienc