Platform Engineer
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
By continuing, you agree to our Terms & Privacy Policy.
Salary Range: $110,000.00 To $115,000.00 Annually
As aPlatform Engineerat ThinkOn, you'll build and maintain the monitoring, observability, and infrastructure automation that keeps our cloud platform reliable. Your day-to-day will span Zabbix and Prometheus/Grafana for monitoring and dashboards, Opsgenie for alert routing and on-call management, and infrastructure-as-code tooling (Ansible, Terraform, GitLab CI/CD) to deploy and manage it all. You'll work within a VMware Cloud Foundation (VCF) and Kubernetes environment, collaborating with infrastructure, network, security, and DevOps teams.
ThinkOn is a remote-first organization. At this time, we are welcoming candidates from Canada for this position.
Please note that this listing is for a current vacancy at ThinkOn.
We are looking for qualified candidates who are eager to contribute and grow with us.
You Will:
Monitoring & Observability
- Deploy, configure, and maintain Zabbix for system and network monitoring across the platform.
- Build and maintain Prometheus exporters and Grafana dashboards for capacity planning, performance metrics, and operational visibility.
- Configure and manage Ops genie for alert routing, escalation policies, and on-call schedules.
- Analyze alert noise, tune thresholds, and reduce false positives to keep alerting actionable.
- Integrate monitoring systems with ticketing, communication, and incident management tools.
- Maintain and optimize monitoring infrastructure — database tuning, storage management, high availability.
Infrastructure Automation & CI/CD
- Write and maintain Ansible playbooks for deploying and configuring monitoring infrastructure.
- Use Terraform for provisioning infrastructure resources where applicable.
- Build and maintain GitLab CI/CD pipelines for automated testing, linting, and deployment of monitoring and infrastructure code.
- Follow infrastructure-as-code practices — version-controlled, peer-reviewed, reproducible.
- Act as a first responder for monitoring-related incidents and alerts.
- Investigate and resolve performance issues, outages, or anomalies detected by monitoring systems.
- Escalate to the appropriate teams (network, security, infrastructure) when needed.
- Document incidents, root causes, and resolutions. Contribute to post-incident reviews.
- Provide technical support to internal teams on monitoring tools and dashboards.
Platform Operations
- Work within VMware vCenter / VCF and Kubernetes environments to support monitoring and infrastructure needs.
- Manage notification infrastructure (SMTP relay configuration, delivery troubleshooting).
- Support compliance requirements (ISO 27001, SOC 2) by maintaining audit logging, access controls, and security configurations for monitoring systems.
You Have:
- Diploma or degree in Computer Science, IT, or a related field (or equivalent practical experience).
- Eligible to obtain Secret Level Clearance within your first 3 months of employment. This requires the successful candidate to be a Canadian Citizen and have 10+ years of verifiable police background history.
- Relevant certifications are a plus but not required (e.g., Zabbix Certified Professional, CKA, CompTIA Linux+, ITIL v4 Foundations).
- Hands-on experience with Zabbix(or comparable: Nagios, Icinga,Checkmk).
- Working knowledge of Prometheus and Grafana— writing exporters, building dashboards, PromQL.
- Experience with alert management and on-call tooling(Ops genie, PagerDuty, or similar).
- Comfort with Linux systems administration (this is a Linux-heavy environment).
- Proficiency in scripting and automation— Bash and Python at minimum.
- Experience with at least one IaCtool (Ansible, Terraform).
- Familiarity with CI/CD pipelines(GitLab CI, GitHub Actions, Jenkins, or similar).
- Basic database administration (PostgreSQL or MySQL) for monitoring tool backends.
- Strong diagnostic and troubleshooting skills — you can work through a problem methodically.
- Clear written and verbal communication —you'll document your work and explain technical issues to varied audiences.
- Attention to detail — monitoring generates a lot of data, and you need to separate signal from noise.
- Comfort working independently in a remote environment while collaborating across teams.
- Ability to stay composed during incidents and work under time pressure.
Nice-to-Have
- Experience with Kubernetes operations and troubleshooting.
- Familiarity with VMware vSphere / VCF environments.
- Exposure to log aggregation tools (ELK/OpenSearch, Loki, Graylog).
- Knowledge of email security standards (SPF, DKIM