Staff Machine Learning Engineer

5 days ago

Toronto ON, Toronto Census Division, ON; Ontario, Canada OpenTable Full-time

This hybrid role requires working in the office two days per week.

With millions of diners, 70,000+ restaurant partners and 25+ years of experience, OpenTable, part of Booking Holdings, Inc. (NASDAQ: BKNG), is an industry leader with a passion for helping restaurants thrive. Our world-class technology empowers restaurants to focus on what matters most – their team, their guests, and their bottom line – while enabling diners to discover and book the perfect restaurant for every occasion.

Every employee at OpenTable has a tangible impact on what we do and how we do it. You’ll also be part of a global team and its portfolio of metasearch brands. Hospitality is all about taking care of others, and it defines our culture.

About the Role

The Data Science team at OpenTable supports a wide range of initiatives targeting diners, restaurants, and internal stakeholders. The team is expanding its capabilities across multiple areas, including building AI Agents to power restaurant search and discovery as well as AI-augmented products for restaurant management.

As a Staff Machine Learning Engineer ,you will set the technical direction for how OpenTable builds, serves, and operates machine learning systems in production. You will partner with Machine Learning Scientists and engineers across the company to take models from experimentation to reliable, monitored production services --- and you will define the standards and patterns the rest of the team builds on.

This is a deliberately engineering-forward role. We are looking for an engineer who builds and operates production systems, rather than a modeller who deploys occasionally. The strongest candidates will bring hard-won judgment from more than one organization about how mature ML teams actually work, and the ability to apply it here. This posting is for an existing vacancy.

Key Initiatives

  • Personalized recommendations for diners
  • Developing and serving high-throughput predictive models for strategic marketplace optimization initiatives
  • Building and integrating tools into our agentic platform via LLM tool calls, MCP, and Agent-to-Agent protocols
  • Multimodal understanding of restaurant content (text, images, geospatial)
  • Creating an AI-powered platform for restaurant partners to gain insights into their business performance and diner demand

What You’ll Own

  • Technical direction for how models are served, deployed, and monitored; the architecture, the patterns, and the tradeoffs behind them.
  • Production ML services that are high-throughput, low-latency, and observable, from design through operation.
  • Engineering standards for ML systems: testing, CI/CD, observability, alerting, rollback, and on-call practice.
  • Ambiguous, cross-team problems: scoping them with Product Managers and stakeholders, then ruthlessly prioritizing what the team actually builds.

Requirements

  • 7+ years of professional software engineering experience, with a substantial portion spent building and operating machine learning systems in production.
  • Breadth of industry perspective. You have seen how ML systems are built and operated at more than one organization, and can speak to industry-standard practices, common reference architectures, and where the real tradeoffs lie.
  • Hands-on experience with a major cloud platform (AWS, GCP, or Azure) as a primary model-serving environment, including its managed services for deployment, scaling, and observability.
  • Deep engineering fundamentals: distributed systems, service and API design, concurrency, latency and throughput tradeoffs, testing discipline, and genuine production ownership including on-call.
  • Strong command of Python and proficiency in at least one strongly typed language (Java preferred).
  • Demonstrated experience training, serving, and deploying ML models at production scale.
  • Production MLOps ownership: model and feature monitoring, drift and data-quality detection, retraining and promotion workflows, versioning, safe rollout and rollback, and incident response when a model misbehaves.
  • A track record of technical leadership: leading multi-quarter projects, influencing engineering decisions beyond your immediate team, and coordinating with Product Managers and other stakeholders.

Strong Preference

  • Serving LLMs in production: inference infrastructure, GPU utilization, batching and caching strategies, quantization, and managing the latency/cost frontier (vLLM, TGI, TensorRT-LLM, or similar).
  • Applied ML depth in ranking, recommendations, classification, NLP, RAG, and/or agentic systems.
  • Kubernetes in production at meaningful scale.
  • Experience developing ETL jobs (especially Spark) or data warehouse in