Observability & Tooling Engineer

Overview

This role offers the opportunity to help shape our future monitoring architecture, introduce modern observability practices, expand OpenTelemetry adoption, improve SLO governance, and build the tooling that enables proactive detection and faster recovery across a global payments platform.

As part of the Observability and Tooling team, you will own the observability lifecycle end-to-end, from architecture and design through implementation, automation, optimization, and ongoing evolution. You will define monitoring standards, build scalable observability solutions, develop dashboards and alerting strategies, and continuously improve the platform through automation and engineering best practices.

Job Description

WHAT YOU WILL DO

  • Configure, tune and operate monitoring and observability across Planet’s broad, multi-vendor estate.
  • Own the alert lifecycle: test-mode onboarding, acceptance criteria, tuning and retirement of noise – make every page actionable.
  • Build dashboards and golden-signal views for the Command Centre floor and SLO reporting.
  • Manage monitoring, alerting and dashboards as code – version-controlled, peer-reviewed and repeatable.
  • Integrate tooling across the estate – observability, ITSM and CMDB (e.g. Jira / Jira Service Management, Device42) – so data flows and context is joined up.
  • Automate everything that gets done twice: scripting, integrations, self-healing and runbook automation.
  • Apply AI and AIOps where it adds value – correlation, noise reduction, anomaly detection and enrichment.
  • Instrument new and existing services to the golden-signal standard (latency, traffic, errors, saturation).
  • Help modernise and rationalise the toolset – reduce overlap, improve coverage and control cost.

WHO YOU ARE

Essential

  • 3+ years in observability / monitoring / tooling engineering or a tooling-heavy operations role.
  • Jira / Jira Service Management – our primary service management platform.
  • Hands-on with a modern log-analytics & observability platform – Coralogix (our primary monitoring & alerting platform) strongly preferred, or strong equivalent (Splunk, Datadog, Grafana).
  • Strong scripting & automation ability (Python, PowerShell or similar).
  • An eye for signal versus noise – you take pages personally.
  • English fluency required; based in Ireland or Poland (hybrid, ~3 days office).

Nice to have

  • Experience across the wider estate – Splunk, Grafana, Datadog, Site24x7, Pingdom, Idera, Netreo, SolarWinds, Device42.
  • Applying AI / AIOps to alerting and operations.
  • Monitoring & alerting as code (version-controlled config).

Skills & Requirements

Observability, Monitoring, Tooling Engineering, Jira, Jira Service Management, Coralogix, Splunk, Datadog, Grafana, Python, PowerShell, Scripting, Automation, OpenTelemetry, SLO Governance, Alerting, Dashboard Development, Log Analytics, ITSM, CMDB, AI, AIOps, Anomaly Detection, Monitoring as Code, Alerting as Code, Device42, Site24x7, Pingdom, Idera, Netreo, SolarWinds

Join Our Community

Let us know the skills you need and we'll find the best talent for you