Co-Op Student
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
By continuing, you agree to our Terms & Privacy Policy.
CO-OP STUDENT — DATA ENGINEER, AI READINESS (MEDICAL-LEGAL DATA)
Fully Remote (Ontario or Quebec)
CONTRIBUTING TO THE CMPA
The Student, Data Engineer, AI Readiness will help transform raw, largely unstructured medical-legal records into governed, high-quality data assets that can be used safely and reliably for AI development and evaluation. This is not a one-time data-cleanup role: the student will help establish reusable data foundations that support CMPA’s broader AI roadmap. This placement offers a rare opportunity to see how disciplined data engineering—including data quality, traceability, privacy protection, and documentation—affects the reliability and auditability of AI systems in a regulated, high-stakes environment.
POSITION OVERVIEW
CMPA's AI & Advanced Analytics team builds AI solutions that process sensitive medical-legal content. We are looking for a co-op student to focus on transforming raw medical-legal data into AI-ready datasets: cleaned, standardized, labeled, and governed so they can reliably train, fine-tune, and evaluate AI models. This role works entirely with internally hosted data and infrastructure (on-prem GPU servers) to maintain strict privacy and security for PHI/PII, and supports real production pipelines already in active development.
You will gain hands‑on experience designing data pipelines and quality controls for sensitive, unstructured information; applying privacy and governance principles in practice; and contributing to AI capabilities that are being developed for real organizational use. You will work closely with experienced AI, data, privacy, and domain professionals.
POSITION ACTIVITIES
- Clean, normalize, and structure raw text from call transcripts, internal documents, and case files into standardized, machine-consumable formats.
- Help define and apply AI-readiness criteria for medical-legal datasets — completeness, consistency, labeling accuracy, lineage, and auditability — so data quality becomes measurable rather than assumed.
- Develop Python scripts and lightweight data pipelines for ingesting, de‑identifying, chunking, formatting, and validating documents at scale, with traceability suitable for audit review.
- Create and maintain evaluation/test sets (e.g., scripted advisor-member conversation samples) to benchmark transcription, summarization, and classification performance.
- Document data preparation standards, labeling guidelines, and known data quality issues so the process is repeatable, governed, and auditable across teams.
- Work closely with the AI program lead to align data preparation work with governance, privacy, and security requirements for regulated healthcare data.
- Support document classification and medical‑legal coding workflows by preparing labeled datasets and validating tagging accuracy against existing coding schemes.
EDUCATION AND EXPERIENCE
- Currently enrolled in a co‑op program in Computer Science, Data Science, Health Informatics, Software Engineering, or a related technical discipline.
- Foundational Python programming skills, including experience working with structured or unstructured data. Experience with libraries such as pandas, regular expressions, spaCy, or similar is an asset.
- Familiarity with core NLP or document‑processing concepts—such as text normalization, tokenization, named‑entity recognition, classification, or information extraction—through coursework, personal projects, research, or prior work.
- Ability to work methodically with ambiguous, incomplete, and messy real‑world data.
SKILLS AND ABILITIES
- Strong attention to detail, sound judgment, and comfort handling sensitive or high‑stakes information in accordance with established procedures.
- Ability to document data transformations, assumptions, validation steps, and known limitations clearly and consistently.
- Strong written communication and an ability to collaborate remotely with technical and non‑technical stakeholders.
- Exposure to healthcare, legal, insurance, public‑sector, or other regulated‑domain data is an asset.
- Familiarity with privacy‑preserving data practices, including de‑identification or anonymization of PHI/PII is an asset.
- Exposure to data quality, data lineage, metadata, data governance, audit, privacy, security, or compliance concepts is an asset.
- Experience with data annotation, labeling tools, building labeling workflows, or evaluating label quality is an asset.<