Login to Traceable Clusters Impacted

Incident Report for Harness

Postmortem

Summary

Between 06:50 UTC and 07:20 UTC on June 26, 2026, customers experienced intermittent login failures while accessing Traceable environments across multiple US clusters. The incident was traced to an issue with the external authentication provider (Auth0), where elevated socket timeouts caused increased login latency and authentication request failures. Login functionality gradually recovered as the upstream issue stabilized, and the incident was resolved after successful login validation across multiple impacted clusters.

Root Cause

The root cause was an issue with the external authentication provider (Auth0), which experienced elevated socket timeouts while processing authentication requests. These upstream timeouts increased login latency and caused intermittent authentication failures across multiple US-region clusters. Since authentication requests depended on the external provider, affected login attempts failed despite Traceable platform services remaining healthy. The incident was resolved once the upstream authentication service recovered and login requests consistently completed successfully.

Impact

Starting at approximately 06:50 UTC, customers experienced intermittent login failures when accessing Traceable environments across multiple US clusters. The issue affected user authentication, preventing some users from accessing the platform while underlying application services remained operational. Login functionality progressively recovered during the incident, and normal authentication was restored by 07:20 UTC.

Remediation

The engineering team worked with the external authentication provider while continuously monitoring authentication health across affected clusters. Login functionality was validated through platform metrics and manual verification across representative environments. After confirming consistent authentication success across impacted clusters, the incident was declared resolved.

Action Items

To prevent such issues going forward, Harness will,  

Increase authentication resilience: Evaluate improvements to authentication request handling, including timeout tuning, retry strategies where appropriate, and graceful degradation for transient upstream failures.

Posted Jun 30, 2026 - 10:13 PDT

Resolved

This incident has been resolved.
Posted Jun 26, 2026 - 00:31 PDT

Update

We are continuing to monitor for any further issues.
Posted Jun 26, 2026 - 00:29 PDT

Monitoring

Issue has been fixed and we are closely monitoring
Posted Jun 26, 2026 - 00:29 PDT

Identified

The issue has been identified and a fix is being implemented.
Posted Jun 26, 2026 - 00:20 PDT

Investigating

We are currently experiencing intermittent login issues with the Traceable cluster due to an issue with our login provider. We are actively working with the provider to resolve the issue.
This incident does not impact data ingestion, and all ingestion pipelines continue to function normally.
We will provide updates as we have more information.
Posted Jun 26, 2026 - 00:20 PDT
This incident affected: Traceable (US - app.traceable.ai / api.traceable.ai, APAC - app.apac.traceable.ai / api.apac.traceable.ai, APAC2 - app.apac2.traceable.ai / api.apac2.traceable.ai, US1 - app.us1.traceable.ai / api.us1.traceable.ai / api2.us1.traceable.ai, Canada - app.ca.traceable.ai / api.ca.traceable.ai).