Co-Op Student

4 days ago

Ottawa, ON, Canada Canadian Medical Protective Association Full-time

CO-OP STUDENT — DATA ENGINEER, AI READINESS (MEDICAL-LEGAL DATA)

Fully Remote (Ontario or Quebec)

CONTRIBUTING TO THE CMPA

The Student, Data Engineer, AI Readiness will help transform raw, largely unstructured medical-legal records into governed, high-quality data assets that can be used safely and reliably for AI development and evaluation. This is not a one-time data-cleanup role: the student will help establish reusable data foundations that support CMPA’s broader AI roadmap. This placement offers a rare opportunity to see how disciplined data engineering—including data quality, traceability, privacy protection, and documentation—affects the reliability and auditability of AI systems in a regulated, high-stakes environment.

POSITION OVERVIEW

CMPA's AI & Advanced Analytics team builds AI solutions that process sensitive medical-legal content. We are looking for a co-op student to focus on transforming raw medical-legal data into AI-ready datasets: cleaned, standardized, labeled, and governed so they can reliably train, fine-tune, and evaluate AI models. This role works entirely with internally hosted data and infrastructure (on-prem GPU servers) to maintain strict privacy and security for PHI/PII, and supports real production pipelines already in active development.

You will gain hands‑on experience designing data pipelines and quality controls for sensitive, unstructured information; applying privacy and governance principles in practice; and contributing to AI capabilities that are being developed for real organizational use. You will work closely with experienced AI, data, privacy, and domain professionals.

POSITION ACTIVITIES

  • Clean, normalize, and structure raw text from call transcripts, internal documents, and case files into standardized, machine-consumable formats.
  • Help define and apply AI-readiness criteria for medical-legal datasets — completeness, consistency, labeling accuracy, lineage, and auditability — so data quality becomes measurable rather than assumed.
  • Develop Python scripts and lightweight data pipelines for ingesting, de‑identifying, chunking, formatting, and validating documents at scale, with traceability suitable for audit review.
  • Create and maintain evaluation/test sets (e.g., scripted advisor-member conversation samples) to benchmark transcription, summarization, and classification performance.
  • Document data preparation standards, labeling guidelines, and known data quality issues so the process is repeatable, governed, and auditable across teams.
  • Work closely with the AI program lead to align data preparation work with governance, privacy, and security requirements for regulated healthcare data.
  • Support document classification and medical‑legal coding workflows by preparing labeled datasets and validating tagging accuracy against existing coding schemes.

EDUCATION AND EXPERIENCE

  • Currently enrolled in a co‑op program in Computer Science, Data Science, Health Informatics, Software Engineering, or a related technical discipline.
  • Foundational Python programming skills, including experience working with structured or unstructured data. Experience with libraries such as pandas, regular expressions, spaCy, or similar is an asset.
  • Familiarity with core NLP or document‑processing concepts—such as text normalization, tokenization, named‑entity recognition, classification, or information extraction—through coursework, personal projects, research, or prior work.
  • Ability to work methodically with ambiguous, incomplete, and messy real‑world data.

SKILLS AND ABILITIES

  • Strong attention to detail, sound judgment, and comfort handling sensitive or high‑stakes information in accordance with established procedures.
  • Ability to document data transformations, assumptions, validation steps, and known limitations clearly and consistently.
  • Strong written communication and an ability to collaborate remotely with technical and non‑technical stakeholders.
  • Exposure to healthcare, legal, insurance, public‑sector, or other regulated‑domain data is an asset.
  • Familiarity with privacy‑preserving data practices, including de‑identification or anonymization of PHI/PII is an asset.
  • Exposure to data quality, data lineage, metadata, data governance, audit, privacy, security, or compliance concepts is an asset.
  • Experience with data annotation, labeling tools, building labeling workflows, or evaluating label quality is an asset.<