AI/ML Scientist

3 days ago

Calgary, Alberta, Canada Socket.dev Full-time

Who We Are & What We Do
At Helix, our mission is simple: to help everyone improve their lives through their DNA. Ready to make a real-world impact with your skills? At Helix, we're transforming healthcare by making genomics a standard of care. We partner with health systems and life science companies to accelerate the integration of genomic data into clinical practice. Join us in building a future where healthcare is personalized, proactive, and powered by genomics.

What is special about this role?
Helix maintains one of the largest linked genomic and clinical datasets in the United States. Helix's Data Science & AI team innovates and develops state-of-the-art models that leverage this clinicogenomic dataset, for example, germline genomes linked to longitudinal clinical records, to better understand and predict a patient's journey, taking precision medicine to a new level. You will take ownership of impactful work early on; your results will be influential, and you will collaborate closely with domain experts. You will have direct access to data that most professionals in this field only read about, along with the opportunity to develop a deep understanding of genomic and clinical data alongside senior experts who have dedicated their careers to this work.

Work you would touch in the first year:

Training and evaluating generative models on longitudinal clinical records fused with variant-level germline genomics, to forecast disease onset and patient trajectory

Building and measuring LLM agentic solutions that pull structured findings out of clinical text and the published literature

Contributing to our multi-agent pipeline for clinical variant interpretation, which is in production today and used by our clinical genomics team on every uncurated variant

Feature and representation work on genomic data: turning variant and gene-level information into something a model can actually use

Running experiments carefully enough to ensure we can trust the results, which primarily means designing the comparison correctly before launching the job.

As an AI/ML Scientist, you will:

Own well-scoped pieces of a larger modeling effort: implement, train, evaluate, and report

Write experiment code that colleagues on the team can read, rerun, and trust

Build and maintain evaluation pipelines, and be honest about what they do and do not measure

Investigate data quality issues. In healthcare data, these are not a distraction from modeling work; they are a necessity

Present your results to the team and leadership, including the ones that did not work, and have opportunities to represent the work externally over time

Collaborate as the AI/ML voice across bioinformatics, clinical, and engineering stakeholders

Draft scientific writeups and internal summaries of the team's findings, and carry results through to papers or conference presentations

About you:

MS or PhD in machine learning, computer science, statistics, computational biology, bioinformatics, or a related quantitative field, or equivalent experience

2 to 5 years of applied machine learning experience beyond your degree. Strong PhD work counts

Solid Python skills. You can write a training loop, read someone else's, and debug it when the loss goes flat

You get real leverage out of AI tools and coding assistants, and you know when to trust their output and when to check it

Comfortable with the practical parts: git, containers, running jobs on cloud GPUs, tracking experiments

Working knowledge of statistics: you know what a train/test leak is, why a baseline matters, and when a difference between two numbers is not a difference

Genuine interest in biology and the clinical problem. You do not need to arrive knowing genomics, but you do need to want to learn it

You are comfortable saying "I do not know yet"

Pluses:

Any exposure to healthcare, genomics, or proteomics data: EHRs, claims, sequencing, imaging, registries

Coursework or projects in computational biology, statistical genetics, or biomedical NLP

Experience with LLM APIs, fine-tuning, or agent frameworks, especially where you had to measure whether the thing worked

Experience with large-scale data tooling (Spark, Dask, Ray, or similar) or with SQL on genuinely large tables

Familiarity with healthcare data standards (OMOP/CDM)

Familiarity with variant interpretation and classification guidelines (ACMG/AMP) or clinical genetics more broadly

Public code, a paper, or a technical writeup: something we can read that shows how you think

PyTorch, Lightning, AWS SageMaker

Expected Interview Process:
1) Recruiter Screen2) Manager Screen 3) Tech Screens4)Final Loop5)Offer

Expected Pay For This Role: There are 3 distinct parts to your Helix offer:1)Base Salary2)Annual Bonus3)Equity

Expected Helix Base: $97