Data Engineer

6 days ago

, Canada Soroc Technology Full-time

An experienced Data Engineer with strong expertise in Big Data technologies to design, develop, and support enterprise-scale data platforms. The ideal candidate should possess hands‑on experience in PySpark, Apache Spark, Kafka, Hadoop ecosystem components, and Apache NiFi, with a strong understanding of data ingestion, transformation, and real‑time processing frameworks.

Key Responsibilities

  • Design, develop, and optimize scalable data pipelines using PySpark, Spark, Hadoop, and Apache NiFi.
  • Build and maintain batch and real‑time data processing solutions.
  • Develop and support Kafka‑based streaming applications and event‑driven architectures.
  • Create and optimize ETL/ELT workflows for large‑scale structured and unstructured datasets.
  • Develop complex SQL queries for data extraction, transformation, validation, and troubleshooting.
  • Implement data ingestion solutions from databases, APIs, files, and streaming sources.
  • Monitor, troubleshoot, and enhance the performance of Spark jobs and data pipelines.
  • Collaborate with architects, business analysts, and development teams to deliver high‑quality data solutions.
  • Support platform upgrades, deployments, testing, certification, and production releases.
  • Ensure data quality, governance, security, and operational excellence across data platforms.

Mandatory Skills

  • PySpark
  • Apache Spark (Spark SQL, DataFrames)
  • Hadoop Ecosystem (HDFS, Hive, YARN)
  • SQL
  • Python

Preferred Skills

  • Spark Streaming
  • Hive
  • Scala
  • Jenkins, Bitbucket, Git
  • JIRA, Confluence
  • Data Warehousing concepts and Dimensional Modeling