Data Engineer
5 hours ago
Winnipeg, Manitoba, Canada
Ampstek
Full-time
Free with email or Google
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
Free with email or Google
You will architect ETL pipelines, manage vector databases, enforce governance, and build real-time data flows that continuously update embeddings and indexes. Your work ensures AI agents operate with fresh, trustworthy information.
What You Will Do
Data Pipelines: Build ETL flows for structured/unstructured data, ensuring normalization, deduplication, and semantic consistency. Vector Infrastructure: Manage pgvector, Azure AI Search, Redis vector indexing, and hybrid search layers. Data Governance: Implement zero-trust access, privacy controls, and compliance within AI context pipelines. Real-time Processing: Build event-driven architectures that continuously refresh embeddings and indexes. Required Qualifications
Deep experience with distributed data systems, SQL, and orchestration tools. Experience tuning high-throughput database infrastructure. Knowledge of Google’s GECX is a plus. Familiarity with chunking strategies and embedding models. Skillset Requirements
ETL & Data Modeling: Designing pipelines for structured/unstructured data, normalization, deduplication, and semantic consistency. Vector Databases: pgvector, Redis, Azure AI Search, hybrid search, and index optimization. Distributed Data Systems: Kafka, Spark, Flink, or similar event-driven architectures. Data Governance: Zero-trust access, privacy controls, compliance, and auditability. Real-time Embedding Updates: Event-driven refresh pipelines for RAG and agent memory systems. Chunking & Embeddings: Semantic chunking, metadata tagging, and embedding model selection. Search Infrastructure: BM25, hybrid search, inverted indexes, and ranking algorithms. Performance Tuning: High-throughput read/write optimization. Data Quality & Lineage: Validation, schema enforcement, and lineage tracking (e.g., Great Expectations, OpenLineage).
Data Pipelines: Build ETL flows for structured/unstructured data, ensuring normalization, deduplication, and semantic consistency. Vector Infrastructure: Manage pgvector, Azure AI Search, Redis vector indexing, and hybrid search layers. Data Governance: Implement zero-trust access, privacy controls, and compliance within AI context pipelines. Real-time Processing: Build event-driven architectures that continuously refresh embeddings and indexes. Required Qualifications
Deep experience with distributed data systems, SQL, and orchestration tools. Experience tuning high-throughput database infrastructure. Knowledge of Google’s GECX is a plus. Familiarity with chunking strategies and embedding models. Skillset Requirements
ETL & Data Modeling: Designing pipelines for structured/unstructured data, normalization, deduplication, and semantic consistency. Vector Databases: pgvector, Redis, Azure AI Search, hybrid search, and index optimization. Distributed Data Systems: Kafka, Spark, Flink, or similar event-driven architectures. Data Governance: Zero-trust access, privacy controls, compliance, and auditability. Real-time Embedding Updates: Event-driven refresh pipelines for RAG and agent memory systems. Chunking & Embeddings: Semantic chunking, metadata tagging, and embedding model selection. Search Infrastructure: BM25, hybrid search, inverted indexes, and ranking algorithms. Performance Tuning: High-throughput read/write optimization. Data Quality & Lineage: Validation, schema enforcement, and lineage tracking (e.g., Great Expectations, OpenLineage).