Data ingestion is delayed on Traceable US production

Incident Report for Harness

Postmortem

Summary

On 19 August 2026 between 12:35 and 17:29 UTC, the Harness Application Security service experienced a significant disruption affecting both the customer-facing console and the data ingestion pipeline in the SaaS Production and US1 regions.

‌

Root Cause

The internal configuration service that supplies runtime settings to nearly every other component became overloaded and entered a repeated restart cycle. Because so many services depend on it, the effects were broad: console pages such as protection policies, posture views, activity logs, API inventory, and custom policy failed to load or timed out, and downstream processing stalled while waiting for configuration it could not obtain.

Customer impact

Dimension Detail Console (UI) impact Multiple pages failed to load or timed out, including protection policies, posture event pages and posture views inside dashboards and insight pages, activity log queries, API inventory screens, custom policy, and sensitive-data views and widgets. Ingestion impact Security telemetry processing degraded severely and, in some paths, stopped entirely. Consumer lag grew across normalisation, grouping, anomaly detection, generation, and related processing stages. Data loss A subset of telemetry ingested during the disruption was permanently dropped.

‌

Mitigation

Several intermediate mitigations additional CPU and memory, relaxed health-check thresholds, a database restart, and a larger connection pool ameliorated the issue. Disabling the new feature in both affected regions restored throughput sharply and durably. The incident was resolved at 17:29 UTC.

‌

Preventive actions

The following actions are committed and tracked internally to completion. The feature that triggered this incident remains disabled and will not be re-enabled until the work below is complete and validated.

Action OPtimize the code by tuning parameters such as cache eviction and retention , evaluate cursor-based pagination for bulk rule retrieval as rule counts grow Add a purpose-built database index for the service-scoping access pattern Remediate pipeline recovery semantics so consumers replay safely after position-marker loss instead of skipping backlog Mandate staged rollout for configuration overrides that alter downstream request patterns: low-volume cluster, then mid-volume, then high-volume Add backpressure and concurrency protection to the configuration service: circuit breaking, bounded queues, and timeout isolation Enhance observability by Instrumenting more detailed metrics
Posted Aug 27, 2026 - 16:35 PDT

Resolved

This incident has been resolved.
Posted Aug 19, 2026 - 09:07 PDT

Update

We are continuing to monitor for any further issues.
Posted Aug 19, 2026 - 08:33 PDT

Monitoring

A fix has been implemented and we are monitoring the results.
Posted Aug 19, 2026 - 08:22 PDT

Identified

The issue has been identified and a fix is being implemented.
Posted Aug 19, 2026 - 08:13 PDT

Update

We are continuing to investigate this issue.
Posted Aug 19, 2026 - 07:55 PDT

Update

We are continuing to investigate this issue.
Posted Aug 19, 2026 - 07:03 PDT

Investigating

We are currently investigating this issue.
Posted Aug 19, 2026 - 06:22 PDT
This incident affected: Traceable (US - app.traceable.ai / api.traceable.ai).