On August 6, 2026 (morning PDT), some customers running pipelines in the Prod2 production environment observed pipeline executions that stopped making progress — stages that did not advance and produced no further output or status updates. The issue was reported by affected customers. Harness engineers identified the cause, mitigated the impact, and pipeline executions returned to normal operation.
The issue was caused by a self-referential pipeline expression. A Git webhook triggered a pipeline that referenced the contents of the webhook payload, and the payload itself contained further copies of that same expression. Each round of expression resolution therefore produced more expressions to resolve, doubling the amount of work each time. This exhausted the resources of the service instance processing that execution, and other executions assigned to the same instance were unable to progress while it was in that state.
During the incident window (approximately 6:11 AM to 11:23 AM PDT on August 6, 2026):
There was no data loss. Pipeline definitions, execution history, and stored state were unaffected. The majority of pipelines on Prod2 continued to execute successfully throughout the incident; the primary impact was that some in-flight executions could not complete and needed to be re-run once the issue was mitigated.
Harness pipelines support expressions that are resolved at runtime — for example, an expression that inserts the contents of the Git webhook payload that triggered the pipeline.
In this case, a Git commit message contained the literal text of the payload expression itself, twice, and the pipeline referenced that same payload expression. Because the commit message is part of the webhook payload, resolving the expression inserted the entire payload — including the two literal copies of the expression carried in the commit message. Those newly inserted copies were then treated as expressions to be resolved, and each pass inserted two more full copies of the payload. The size of the value being processed, and the work required to process it, therefore doubled on every pass and grew exponentially rather than converging.
Harness has a safeguard intended to stop exactly this: expression resolution is bounded by a maximum nesting depth, beyond which resolution halts and the pipeline fails with an explicit error. A defect in that safeguard meant the limit was not applied in this specific self-referential case, so resolution continued unchecked.
Expression resolution runs inline on the threads that start pipeline steps. As each pass consumed progressively more memory and CPU without ever completing, the service instance performing that work stopped making progress, and every execution assigned to that instance stalled — which is what customers reported.
Harness completed the following immediate mitigation steps:
These actions restored normal pipeline execution behavior and resolved the customer-facing impact.
To reduce the risk of recurrence and improve detection, the following actions are in various stages of being implemented: