Customers on Prod1, Prod2, and Prod3 (US) clusters experienced failures when loading SEI 2.0 dashboards on August 6, 2026, from 7:22 AM PDT to 9:03 AM PDT. Customers calling the SEI 2.0 API also experienced similar failures.
No customer data was lost, and ingestion of all integration data continued to work uninterrupted. SEI customers using 1.0 were not impacted.
The incident was caused by resource exhaustion on the nodes serving queries. This resource degradation developed in a pattern that did not cross our existing alerting thresholds early enough to provide sufficient warning or allow mitigation before customer impact occurred.
Customers on Prod1, Prod2, and Prod3 (US) clusters were unable to load SEI 2.0 dashboards during the incident window.
Duration: August 6, 2026, 7:22 AM PDT – 9:03 AM PDT (~1 hour 41 minutes)
No customer data was lost.
Upon identifying the root cause, our team took immediate corrective action by adding capacity to restore the affected systems. Services were fully recovered, and all dashboards resumed normal operation at 9:03 AM PDT.
To prevent from such issues happening again, Harness is/has
Proactively added capacity updates have been applied to prevent this issue from recurring
Additional monitoring and alerting have been put in place to detect anomalies early, focused on a leading indicator, which in this case was thread pool exhaustion, before they can impact dashboard availability and data rendering.
We are working with our vendor to apply a patch to remediate this and similar issues completely.