Root Cause Analysis
A routine platform update included a structural change to one of our largest and most frequently used database tables. Applying that change required locking the affected data for roughly a minute and a half, and because requests are handled in the order they arrive, normal traffic queued behind that lock until the system exhausted its capacity to take on new work and began returning errors.
Resolution
The issue cleared on its own as soon as the update finished: the lock was released, queued requests were processed, and service returned to normal with no rollback or restart required. Customer-visible impact lasted approximately nine minutes, including about five minutes of failed requests; the update itself completed successfully, and no customer data was lost or affected.
Remediation and Preventive Actions
- Separating database maintenance from platform updates. Changes of this kind are being moved out of the release process to run as independently scheduled work that does not interrupt live traffic.
- Automated review of high-risk changes. We are adding automated checks and specialist review to catch changes that could restrict access to high-volume data before they reach production.
- Faster detection and response. Operations like this will be time-limited so they stop automatically rather than hold up traffic, and we are expanding monitoring and on-call procedures so this type of contention is detected and cleared sooner.