Data Engineer

5 hours ago

Winnipeg, Manitoba, Canada Ampstek Full-time
You will architect ETL pipelines, manage vector databases, enforce governance, and build real-time data flows that continuously update embeddings and indexes. Your work ensures AI agents operate with fresh, trustworthy information. What You Will Do

Data Pipelines: Build ETL flows for structured/unstructured data, ensuring normalization, deduplication, and semantic consistency. Vector Infrastructure: Manage pgvector, Azure AI Search, Redis vector indexing, and hybrid search layers. Data Governance: Implement zero-trust access, privacy controls, and compliance within AI context pipelines. Real-time Processing: Build event-driven architectures that continuously refresh embeddings and indexes. Required Qualifications

Deep experience with distributed data systems, SQL, and orchestration tools. Experience tuning high-throughput database infrastructure. Knowledge of Google’s GECX is a plus. Familiarity with chunking strategies and embedding models. Skillset Requirements

ETL & Data Modeling: Designing pipelines for structured/unstructured data, normalization, deduplication, and semantic consistency. Vector Databases: pgvector, Redis, Azure AI Search, hybrid search, and index optimization. Distributed Data Systems: Kafka, Spark, Flink, or similar event-driven architectures. Data Governance: Zero-trust access, privacy controls, compliance, and auditability. Real-time Embedding Updates: Event-driven refresh pipelines for RAG and agent memory systems. Chunking & Embeddings: Semantic chunking, metadata tagging, and embedding model selection. Search Infrastructure: BM25, hybrid search, inverted indexes, and ranking algorithms. Performance Tuning: High-throughput read/write optimization. Data Quality & Lineage: Validation, schema enforcement, and lineage tracking (e.g., Great Expectations, OpenLineage).