SRE · Observability

Observability and on-call setup

Helix Payments — 2025

Structured logging, tracing, dashboards and alert routing for a payments platform.

Result
Mean time to detect cut from 40m to 3m
Timeline
8 weeks
Observability and on-call setup interface for Helix Payments — 2025

The challenge

A payments company learned about outages from customers. Logs were unstructured, there were no traces, and alerts were noise nobody trusted.

What we did

  • Instrumented services with OpenTelemetry traces and structured logs.
  • Defined service level objectives and alerted on symptoms, not causes.
  • Built dashboards per service plus a business-level payments view.
  • Set up on-call rotation, escalation and runbooks.

The outcome

  • Mean time to detect down from 40 minutes to 3.
  • Alert volume cut sharply while catching more real incidents.
  • Every incident now has a runbook and a written review.

Stack

  • Kubernetes
  • OpenTelemetry
  • Prometheus
  • Grafana
  • PagerDuty

Tell us about your project

Free consultation, no obligations. Send the details or schedule a call for the fastest reply.