Platform Engineer

2 hours ago

Toronto, Ontario, Canada ThinkOn Inc. Full-time

Salary Range: $110,000.00 To $115,000.00 Annually

As aPlatform Engineerat ThinkOn, you'll build and maintain the monitoring, observability, and infrastructure automation that keeps our cloud platform reliable. Your day-to-day will span Zabbix and Prometheus/Grafana for monitoring and dashboards, Opsgenie for alert routing and on-call management, and infrastructure-as-code tooling (Ansible, Terraform, GitLab CI/CD) to deploy and manage it all. You'll work within a VMware Cloud Foundation (VCF) and Kubernetes environment, collaborating with infrastructure, network, security, and DevOps teams.

ThinkOn is a remote-first organization. At this time, we are welcoming candidates from Canada for this position.

Please note that this listing is for a current vacancy at ThinkOn.

We are looking for qualified candidates who are eager to contribute and grow with us.

You Will:

Monitoring & Observability

  • Deploy, configure, and maintain Zabbix for system and network monitoring across the platform.
  • Build and maintain Prometheus exporters and Grafana dashboards for capacity planning, performance metrics, and operational visibility.
  • Configure and manage Ops genie for alert routing, escalation policies, and on-call schedules.
  • Analyze alert noise, tune thresholds, and reduce false positives to keep alerting actionable.
  • Integrate monitoring systems with ticketing, communication, and incident management tools.
  • Maintain and optimize monitoring infrastructure — database tuning, storage management, high availability.

Infrastructure Automation & CI/CD

  • Write and maintain Ansible playbooks for deploying and configuring monitoring infrastructure.
  • Use Terraform for provisioning infrastructure resources where applicable.
  • Build and maintain GitLab CI/CD pipelines for automated testing, linting, and deployment of monitoring and infrastructure code.
  • Follow infrastructure-as-code practices — version-controlled, peer-reviewed, reproducible.
  • Act as a first responder for monitoring-related incidents and alerts.
  • Investigate and resolve performance issues, outages, or anomalies detected by monitoring systems.
  • Escalate to the appropriate teams (network, security, infrastructure) when needed.
  • Document incidents, root causes, and resolutions. Contribute to post-incident reviews.
  • Provide technical support to internal teams on monitoring tools and dashboards.

Platform Operations

  • Work within VMware vCenter / VCF and Kubernetes environments to support monitoring and infrastructure needs.
  • Manage notification infrastructure (SMTP relay configuration, delivery troubleshooting).
  • Support compliance requirements (ISO 27001, SOC 2) by maintaining audit logging, access controls, and security configurations for monitoring systems.

You Have:

  • Diploma or degree in Computer Science, IT, or a related field (or equivalent practical experience).
  • Eligible to obtain Secret Level Clearance within your first 3 months of employment. This requires the successful candidate to be a Canadian Citizen and have 10+ years of verifiable police background history.
  • Relevant certifications are a plus but not required (e.g., Zabbix Certified Professional, CKA, CompTIA Linux+, ITIL v4 Foundations).
  • Hands-on experience with Zabbix(or comparable: Nagios, Icinga,Checkmk).
  • Working knowledge of Prometheus and Grafana— writing exporters, building dashboards, PromQL.
  • Experience with alert management and on-call tooling(Ops genie, PagerDuty, or similar).
  • Comfort with Linux systems administration (this is a Linux-heavy environment).
  • Proficiency in scripting and automation— Bash and Python at minimum.
  • Experience with at least one IaCtool (Ansible, Terraform).
  • Familiarity with CI/CD pipelines(GitLab CI, GitHub Actions, Jenkins, or similar).
  • Basic database administration (PostgreSQL or MySQL) for monitoring tool backends.
  • Strong diagnostic and troubleshooting skills — you can work through a problem methodically.
  • Clear written and verbal communication —you'll document your work and explain technical issues to varied audiences.
  • Attention to detail — monitoring generates a lot of data, and you need to separate signal from noise.
  • Comfort working independently in a remote environment while collaborating across teams.
  • Ability to stay composed during incidents and work under time pressure.

Nice-to-Have

  • Experience with Kubernetes operations and troubleshooting.
  • Familiarity with VMware vSphere / VCF environments.
  • Exposure to log aggregation tools (ELK/OpenSearch, Loki, Graylog).
  • Knowledge of email security standards (SPF, DKIM