Entities in Harness are not loading on in Prod3

Incident Report for Harness

Postmortem

Summary

Between August 27 and August 28, 2026, customers experienced an issue where some pipelines, deployments, and related resources appeared as not found in the Harness UI and API, even though the underlying data remained intact.

The issue occurred during a planned internal infrastructure update that affected communication between internal platform services. As a result, requests that depended on account, organization and project scope resolution were unable to complete successfully, which led to incorrect not found responses being returned to customers for existing entities.

Engineering identified the issue, rolled back the change, and restored normal service. No customer data was lost or deleted during the incident.

Root Cause

The issue was caused by a configuration error introduced during a planned internal service routing update in Production.

An internal platform service responsible for resolving account, organization, and project context was unable to validate requests from other Harness services after the change was applied. Because that validation step is required before many entity reads and pipeline-related actions can proceed, the failed requests surfaced to customers as not found errors for resources that continued to exist normally.

The issue was limited to the affected production environment and was resolved by reverting the change and restoring the previous service communication path.

Impact

  • Some customers saw existing pipelines, deployments, and related entities appear as not found in the UI and API.
  • Some pipeline-related operations, including execution progression, webhook-triggered starts, scheduled trigger evaluation, and entity listing, were temporarily disrupted.
  • The issue affected availability and visibility of existing entities, but it did not remove data or change customer configurations.
  • No unauthorized access occurred, and no customer data loss was observed.

Remediation

  • Immediate: Reverted the infrastructure configuration update and restored the previously working service communication path.
  • Recovery validation: Verified that affected entity lookups, pipeline operations, and dependent APIs were functioning normally after rollback.
  • Permanent: Corrected the configuration handling associated with the update so similar issues do not interfere with service-to-service authentication in future rollouts.

Action Items

To prevent such issues from happening again, Harness will 

  1. Improve configuration validation by enhancing the pre-deployment tests to verify internal service communication before shifting production traffic.
  2. Enhance monitoring and alerting for internal authentication failures so issues can be detected earlier.
  3. Improve error handling so dependency failures are less likely to appear to customers as resource not found errors.
Posted Sep 02, 2026 - 15:16 PDT

Resolved

This incident has been resolved.
Posted Aug 28, 2026 - 00:35 PDT

Monitoring

A fix has been implemented and we are monitoring the results.
Posted Aug 28, 2026 - 00:29 PDT

Identified

The issue has been identified and a fix is being implemented.
Posted Aug 28, 2026 - 00:16 PDT

Investigating

We are currently investigating this issue.
Posted Aug 28, 2026 - 00:04 PDT
This incident affected: Prod 3 (Platform).