Write-up published
Resolved
Service returned to normal performance at approximately 7:29 PM ET. Total impact lasted approximately nine minutes, including about five minutes of sustained request failures. The update that triggered the issue completed successfully, and no customer data was lost or affected.
Identified
We identified the cause as a structural database change included in a routine platform update. Completing that change required locking a core data table, and incoming requests queued behind the lock until the system could no longer accept new work. The operation completed at approximately 7:22 PM ET and error rates began recovering.
Investigating
Our monitoring detected elevated error rates affecting API requests and agent responses, beginning at approximately 7:14 PM ET. Affected requests returned errors rather than completing slowly. Engineers were engaged and began investigating immediately.