Between 06:50 UTC and 07:20 UTC on June 26, 2026, customers experienced intermittent login failures while accessing Traceable environments across multiple US clusters. The incident was traced to an issue with the external authentication provider (Auth0), where elevated socket timeouts caused increased login latency and authentication request failures. Login functionality gradually recovered as the upstream issue stabilized, and the incident was resolved after successful login validation across multiple impacted clusters.
The root cause was an issue with the external authentication provider (Auth0), which experienced elevated socket timeouts while processing authentication requests. These upstream timeouts increased login latency and caused intermittent authentication failures across multiple US-region clusters. Since authentication requests depended on the external provider, affected login attempts failed despite Traceable platform services remaining healthy. The incident was resolved once the upstream authentication service recovered and login requests consistently completed successfully.
Starting at approximately 06:50 UTC, customers experienced intermittent login failures when accessing Traceable environments across multiple US clusters. The issue affected user authentication, preventing some users from accessing the platform while underlying application services remained operational. Login functionality progressively recovered during the incident, and normal authentication was restored by 07:20 UTC.
The engineering team worked with the external authentication provider while continuously monitoring authentication health across affected clusters. Login functionality was validated through platform metrics and manual verification across representative environments. After confirming consistent authentication success across impacted clusters, the incident was declared resolved.
To prevent such issues going forward, Harness will,
Increase authentication resilience: Evaluate improvements to authentication request handling, including timeout tuning, retry strategies where appropriate, and graceful degradation for transient upstream failures.