Observability & Tooling Lead

Overview

We are looking for an Observability & Tooling Lead to lead Planet's observability transformation and operational tooling strategy. You will lead a team owning monitoring, alerting, CMDB integrations, automation, and the observability roadmap, with end-to-end accountability from design and architecture through implementation, engineering standards, automation, and continuous improvement. Working closely with Engineering, Infrastructure, Cloud, and Operations teams, you will ensure our observability capabilities provide actionable insights, accelerate incident response, and support the reliability of Planet's global payments platform.

Job Description

WHAT YOU WILL DO

  • Lead the Observability & Tooling team: three engineers across monitoring, alerting, dashboarding and tooling integration.
  • Own the operations tooling estate — monitoring, alerting, CMDB and ITSM integrations — including vendor and license management.
  • Drive the modernisation of observability tooling from on-prem-licensed platforms to modern cloud-native ones.
  • Define and enforce golden-signal monitoring standards (latency, traffic, errors, saturation) for critical services.
  • Own the pager-quality budget: test-mode onboarding for new alerts, acceptance criteria, and a measured reduction in non-actionable pages.
  • Partner with the Command Centre, SRE and engineering teams so observability serves operations, not the other way round.
  • Own the tooling roadmap and its alignment to SLO and AIOps ambitions.

WHO YOU ARE

  • 7+ years across monitoring / observability / operations tooling, with team leadership experience.
  • Deep hands-on history with enterprise monitoring and modern observability stacks (e.g. Splunk, Coralogix, Grafana, Datadog, Site24x7, Pingdom, Idera, Netreo, SolarWinds).
  • Strong opinions, loosely held, about alert quality — and a record of cutting pager noise.
  • Experience with CMDB and ITSM tooling integration (Device42, Jira ecosystems valuable).
  • Practical interest in applying AI / AIOps to alerting and operations.
  • Automation-first mindset; experience driving automation through Python, PowerShell, APIs, Infrastructure as Code, or similar technologies.
  • Comfortable operating at both strategic and technical levels, from defining roadmaps and architecture to reviewing implementations and guiding engineering teams.
  • Proven ability to influence and collaborate across Engineering, Infrastructure, Cloud, Security, and Operations teams.

Skills & Requirements

Monitoring, Observability, Operations Tooling, Team Leadership, Enterprise Monitoring, Splunk, Coralogix, Grafana, Datadog, Site24x7, Pingdom, Idera, Netreo, SolarWinds, Alerting, Golden Signals, Pager Quality, CMDB, ITSM, Device42, Jira, AI/AIOps, Python, PowerShell, APIs, Infrastructure as Code (IaC), Cloud-Native Technologies, SLOs, Observability Architecture, Automation, Incident Response, Engineering Standards, Roadmap Planning, Cross-functional Collaboration

Join Our Community

Let us know the skills you need and we'll find the best talent for you