SRE · Observability
Observability and on-call setup
Helix Payments — 2025
Structured logging, tracing, dashboards and alert routing for a payments platform.
- Result
- Mean time to detect cut from 40m to 3m
- Timeline
- 8 weeks

The challenge
A payments company learned about outages from customers. Logs were unstructured, there were no traces, and alerts were noise nobody trusted.
What we did
- Instrumented services with OpenTelemetry traces and structured logs.
- Defined service level objectives and alerted on symptoms, not causes.
- Built dashboards per service plus a business-level payments view.
- Set up on-call rotation, escalation and runbooks.
The outcome
- Mean time to detect down from 40 minutes to 3.
- Alert volume cut sharply while catching more real incidents.
- Every incident now has a runbook and a written review.
Stack
- Kubernetes
- OpenTelemetry
- Prometheus
- Grafana
- PagerDuty
Tell us about your project
Free consultation, no obligations. Send the details or schedule a call for the fastest reply.