Lead Data Engineer
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
to apply
- email only, no card. You can also save this posting or score it againstyour profile with AI.##
About the role
The Lead Data Engineer will design and build a secure, scalable cloud-native data platform on AWS using medallion architecture. They will also provide technical leadership, mentor team members, and ensure the reliability and security of data pipelines.## RequirementsCandidates must have 7+ years of experience in data engineering and distributed systems on cloud platforms. Proficiency in AWS services, Snowflake, Spark, and infrastructure-as-code is required.## BenefitsContinuous learningCoaching and development opportunitiesCollaborative team environment## Full descriptionJob DescriptionGlobal Functions Technology (GFT) partners across RBC to deliver transformative platforms and solutions. In Anti-Money Laundering (AML), we are building a new Data Foundation Hub to ingest enterprise data and power analytics and controls using a medallion architecture. As a Lead Data Platform Engineer, you will be a senior individual contributor and technical lead, owning the design and build of our AWS-based data platform and mentoring other engineers.
You will work 70-80% hands-on across AWS (EKS, S3, RDS, EMR, Glue, Airflow), Snowflake, Spark, and dbt to deliver cloud-native, governed, and reliable data systems.
* Technical leadership and platform ownership
* Lead the technical direction for the AML Data Foundation Hub on AWS.
* Mentor and coach engineers (tech design reviews, pair programming, standards), influencing quality and delivery.
* Cloud-native data platform on AWS (hands-on)
* Design and build secure, scalable data platforms using AWS S3, Glue, EMR, RDS, and EKS.
* Define patterns for data lake and warehouse integration (e.g., S3 + Snowflake) including partitioning, storage classes, encryption, and cost optimization.
* Implement Infrastructure-as-Code (e.g., CloudFormation/Terraform) for repeatable environments, networking, IAM roles/policies, and security baselines.
* Data engineering and architecture (medallion)
* Design and build batch and incremental pipelines across Bronze/Silver/Gold layers using Snowflake (Streams, Tasks, Snowpark), Spark on EMR, and dbt.
* Implement schema evolution, SCD/CDC, partitioning, and performance tuning across both compute and storage (S3, EMR, Snowflake, RDS).
* Ingestion, orchestration, and observability
* Engineer resilient, observable ingestion patterns into S3/Snowflake/RDS.
* Orchestrate pipelines using Airflow (or equivalent) and/or AWS-native services (e.g., event triggers), enforcing SLAs, retries, idempotency, and alerting.
* Build operational dashboards and alerts for pipeline health, platform capacity, and cost.
* Reliability, DR, and security
* Design for high availability, resiliency, and disaster recovery (multi-AZ,/region backup/restore, RPO/RTO-aware architectures).
* Implement secrets management, encryption, IAM least-privilege, and network security in partnership with Security and Platform/SRE.
* Participate in incident response and postmortems; drive root-cause fixes and hardening of the platform.
* DevOps for data and platform enablement
* Own CI/CD for data and platform components: code review, environment promotion, automated tests (unit, integration, data contract), and versioned artifacts.
* Partner with Platform/SRE on SLIs/SLOs, capacity planning, and platform standardization across squads.
* Cross-functional collaboration
* Translate AML business and control objectives into technical roadmaps, platform capabilities, and reusable patterns.
Must-have
* Experience depth: 7+ years delivering production data pipelines and distributed systems at scale on cloud platforms; demonstrated ability to operate as a senior IC and technical lead influencing architecture and quality across a team.
* AWS platform depth: Hands-on with S3, Glue, EMR, EKS, and RDS; proficiency with IaC (CloudFormation or Terraform), IAM least-privilege design, VPC/networking, and security baselines.
* Snowflake expertise: Hands-on with Streams, Tasks, Snowpark, and Snowpipe; strong SQL and warehouse design; performance optimization across compute and storage.
* Distributed processing: Production experience with Spark (PySpark/Scala) for large-scale batch processing, optimization, and tuning.
* Data engineering and architecture: Medallion architecture patterns (Bronze/Silver/Gold), schema evolution, SCD/CDC, partitioning, and end-to-end pipeline performance tuning.
* Orchestration and automation: Airflow (or equivalent) for DAGs, SLAs, retries, idempotency, and observability; Git-based workflows and CI/CD for data pipelines (e.g., GitHub Actions/Jenkins).
* Reliability and security: Designing for HA/DR (multi-AZ, backup/restore, RPO/RTO); encryption, secrets management, and network security in partnership with Platform/SRE.
* DevOps for data: Ownership of automated testing (unit, integration, data contract), environment promotion, and versioned artifacts.
* Ways of working: Strong ownership, structured