Slowness in Prod1 and Prod2 environment

Incident Report for Harness

Postmortem

Summary

On September 8, 2026, customers in Prod 1 and Prod 2 experienced elevated platform latency and pipeline failures. The issue was caused by a regression in a newly released capability that triggered cascading failures under high load. Because the capability was behind a feature flag, it was quickly disabled, and service was restored after a brief monitoring period. 

Customer Impact

  • Customers encountered slowness and failures during pipeline execution and UI operations. Some API calls returned errors or timed out.
  • No data loss or corruption occurred.

Root Cause

The new capability introduced a regression that created contention on a shared backend resource used by multiple Harness components. This saturated the shared platform infrastructure and caused the cascading failures.

Mitigation

  • Disabled the capability across all environments
  • Temporarily increased platform capacity to restore stability

Next Steps

To prevent recurrence, Harness will:

  1. Permanently fix the capability by profiling and eliminating the sub-optimal code path and query
  2. Improve detection by enhancing alerting for resource-intensive queries on high-frequency platform paths
Posted Sep 09, 2026 - 16:23 PDT

Resolved

This incident has been resolved.
Posted Sep 08, 2026 - 21:30 PDT

Update

A fix has been implemented and we are monitoring the results.
Posted Sep 08, 2026 - 17:39 PDT

Monitoring

A fix has been implemented and we are monitoring the results.
Posted Sep 08, 2026 - 17:36 PDT

Update

We are continuing to work on a fix for this issue.
Posted Sep 08, 2026 - 17:29 PDT

Identified

The issue has been identified and a fix is being implemented.
Posted Sep 08, 2026 - 16:52 PDT

Update

We are continuing to investigate this issue.
Posted Sep 08, 2026 - 16:31 PDT

Update

We are continuing to investigate the issue.
Posted Sep 08, 2026 - 15:53 PDT

Update

We are continuing to investigate this issue.
Posted Sep 08, 2026 - 15:53 PDT

Investigating

We are currently investigating this issue.
Posted Sep 08, 2026 - 14:52 PDT
This incident affected: Prod 1 (Continuous Delivery (CD) - FirstGen - EOS, Continuous Delivery - Next Generation (CDNG), Cloud Cost Management (CCM), Continuous Error Tracking (CET), Chaos Engineering, Feature Flags (FF), FME) and Prod 2 (Continuous Delivery (CD) - FirstGen - EOS, Continuous Delivery - Next Generation (CDNG), Cloud Cost Management (CCM), Continuous Error Tracking (CET), Chaos Engineering, Continuous Integration Enterprise(CIE) - Self Hosted Runners, Continuous Integration Enterprise(CIE) - Mac Cloud Builds, Continuous Integration Enterprise(CIE) - Windows Cloud Builds, Continuous Integration Enterprise(CIE) - Linux Cloud Builds, Custom Dashboards, Feature Flags (FF), Security Testing Orchestration (STO), Service Reliability Management (SRM), Internal Developer Portal (IDP), Infrastructure as Code Management (IaCM), Software Supply Chain Assurance (SSCA), Software Engineering Insights (SEI) / AI DLC Insights, Code Repository, Artifact Registry, Platform, AI Test Automation, FME).