Prod2 was intermittently unavailable

Incident Report for Harness

Postmortem

Summary

Between 12:34am PST and 12:38am PST on 5th September, the Delegate service manager experienced some elevated exceptions when attempting to write to the database. Consequently, delegate connections were dropped, causing them to disconnect. Delegate automatically re-attempts registration back to the delegate service manager  and majority of the delegates got connected back after the incident. For Docker and ECS delegates the automatic restart is not enabled unless these delegates have health monitoring enabled. For these delegates a manual restart is needed and was recommended.  Post restart the delegate would re-connect and the issue was resolved.

Root cause

On Prod2 cluster we identified a performance bottleneck in the delegate service that, under certain conditions, can increase database write latency and delay heartbeat processing which leads to delegates being disconnected.

Impact

All K8s delegates and `Docker/ECS` delegates got connected back immediately within 4 mins and started to function normally. The impact can be scoped to those specific types of delegates that didn’t have health monitoring enabled. 

Remediation

  • Immediate: We have added additional monitoring and increased resources for handling the influx of traffic. 
  • Permanent: We have identified a hotspot in the code that can cause high latency when writing to a database which we are actively working on resolving. 

Action Items

To prevent such issues from happening again, Harness will work on the following:

  1. Increased targeted monitoring and alerting to initiate timely mitigation and prevent this from happening again.
  2. Fix the identified delegate service managers database client reconnect failures
  3. Fix the hotpots that can cause query latency.
Posted Sep 09, 2026 - 11:22 PDT

Resolved

This incident has been resolved.
Posted Sep 05, 2026 - 02:40 PDT

Investigating

We are currently investigating this issue.
Posted Sep 05, 2026 - 02:40 PDT
This incident affected: Prod 2 (Platform).