Between 12:34am PST and 12:38am PST on 5th September, the Delegate service manager experienced some elevated exceptions when attempting to write to the database. Consequently, delegate connections were dropped, causing them to disconnect. Delegate automatically re-attempts registration back to the delegate service manager and majority of the delegates got connected back after the incident. For Docker and ECS delegates the automatic restart is not enabled unless these delegates have health monitoring enabled. For these delegates a manual restart is needed and was recommended. Post restart the delegate would re-connect and the issue was resolved.
On Prod2 cluster we identified a performance bottleneck in the delegate service that, under certain conditions, can increase database write latency and delay heartbeat processing which leads to delegates being disconnected.
All K8s delegates and `Docker/ECS` delegates got connected back immediately within 4 mins and started to function normally. The impact can be scoped to those specific types of delegates that didn’t have health monitoring enabled.
To prevent such issues from happening again, Harness will work on the following: