From 11:08 AM – 11:25 AM EDT on July 10, 2026, Knock experienced a full outage of its workflow execution engine, the layer of our system that processes workflows and enqueues messages for delivery to recipients via downstream providers.
During this time frame, Knock stopped processing all queued workflow and outbound webhook jobs; all new job enqueue attempts failed permanently. We also observed elevated response times across multiple V1 API endpoints.
Then, at 12:25 PM EDT, during our efforts to replay failed work during the outage period, we unintentionally enqueued several thousand message send and delivery status check actions for messages that had previously failed delivery prior to the outage window. As a result, some customers later observed unexpected message deliveries to their recipients.
Despite the severity of this outage and the unintentional replay error, we were still able to successfully process all failed work from the outage window once our systems recovered.
To our customers, we sincerely apologize for any issues caused by this outage. At Knock we consider the uptime, stability, and consistency of our message delivery layer to be of paramount importance. In addition to providing a detailed root cause analysis below, we are taking several immediate actions to prevent these failures from happening again.
At Knock, workflow execution is currently backed by a Postgres-based job queue. This job queue runs on the AWS Aurora Postgres database cluster that serves as the primary database for the Knock notification service. This incident was caused by a compounding set of factors that put extreme pressure on the writer instance of that database cluster, thereby degrading workflow execution.
We recently shipped some changes that increased the average amount of data written per row added to the jobs table backing our workflow execution system. Specifically, this increase came in the form of larger payloads written to the jobs.args column, a JSONB column with a GIN index. As part of this change, we’ve been closely monitoring the size of this GIN index, performing nightly reindex actions to limit growth. However, what we were not aware of was degrading write performance to this index due to the increased frequency with which Postgres was needing to add new database pages to store indexed data.
This index write performance issue had gone unnoticed until the start of this incident, when Postgres kicked off an autovacuum analyze action against the jobs table. Then, in an instance of poor timing, a new code deploy to our notification service began to roll out about 20 seconds after the autovacuum kickoff. Together, these factors produced a sudden and critical increase in query and CPU pressure on this database. Specifically:
While running, autovacuum prevented writes to the GIN index from filling existing pages with data. Thus, all new job enqueues required multiple new pages allocated, creating a large backlog of relation extension locks that Postgres could not keep pace with.
This backlog of queued page allocations caused CPU pressure to increase rapidly and queries to slow. This caused autovacuum to slow down, prolonging the duration of the issue.
Lastly, as part of our notification engine service deployed, we spiked the number of connections to this database and the number of queries issued against the jobs table, further exacerbating both of the above factors.
During this period, the writer instance of this database cluster became virtually inaccessible, with CPU and query resources saturated. This brought workflow execution to a halt and caused all queries against the writer to slow. The degraded writer instance performance cascaded to several other layers of our service. It impacted message delivery and status checks running on AWS SQS queues. It also caused response times for many V1 API endpoints to increase.
The outage rapidly resolved when the autovacuum completed. At this point we observed page extension locks clear out, CPU pressure return to normal, and workflow execution resume as expected.
After resolution, we began to find and replay failed work. First, we identified and replayed workflow jobs that had failed execution during the outage period. Then, we identified all workflow job enqueues that had failed and reenqueued them for successful execution.
Finally, we initiated two full SQS DLQ redrive tasks to replay failed message deliveries and status checks from the outage. However, when performing these redrives we failed to account for each DLQ containing failed work from before the outage that was not viable for replay. At the time of replay, each DLQ contained failures as old as from 12:25 PM EDT on 2026-07-03.
Thus, for the redrive of our message delivery DLQ, we reenqueued messages that had previously failed delivery from 2026-07-03 thru 2026-07-10. For some of these messages, Knock then succeeded in delivering the message to the customer recipient, which was not expected by the recipient nor the customer. However, no previously sent messages were included in this replay; no duplicate messages were ever delivered. The DLQ redrive for the delivery status checks led to delivery status checks becoming delayed for some customers, as queues become clogged with replay work.
July 10, 2026, 11:07:50 AM EDT — autovacuum analyze on the jobs table kicks off. The writer database instance for the Aurora database backing the Knock notification service sees a rapid increase in CPU pressure and active sessions associated with the Lock:extend event, indicative of queued database page allocations. Workflow executions quickly slow to a halt as the database struggles to serve queries and most new job enqueues fail. V1 API response times begin to climb.
11:08:10 AM EDT — A deployment to Knock’s notification service begins to rollout, spiking connections to the database and further exacerbating query pressure issues.
11:10 AM EDT — Knock primary on-call receives the first paging alert related to the outage. Knock’s Platform Team engineers quickly see degraded state across multiple systems and open an internal incident.
11:12 AM EDT — Knock engineers turn off a feature flag that had been rolled out last week, increasing the average size of rows written to the jobs table, under the hypothesis that increased write pressure might be a factor. This does not resolve the issue.
11:18 AM EDT — Knock publishes an incident to the public status page.
11:25 AM EDT — Autovacuum concludes and Knock engineers observe rapid recovery in the system.
11:40 AM EDT — Knock resolves the incident on the status page.
11:49 AM EDT — All workflow execution job failures from the outage window are identified and replayed.
12:25 PM EDT — Knock engineers initiate an SQS DLQ redrive to replay failed message delivery and status checks from the outage window. This redrive unintentionally covers messages that had failed delivery or status check from prior to the outage window, leading to unexpected message delivery to some customer recipients.
12:30 PM EDT — All failed workflow job enqueues from the outage window are successfully replayed.
Workflow and outbound webhook workflow job execution dropped to virtually zero.
All attempts to enqueue new workflow and outbound webhook workflow jobs failed.
Some message delivery and status checks failed unexpectedly.
V1 API response times increased for endpoints that executed queries against the primary service writer database instance.
We unintentionally replayed failed message delivery and status checks from 2026-07-03 – 2026-07-10 that predated the outage window. Some of these message deliveries succeeded, leading to unexpected messages sent to customer recipients. Delivery status checks became delayed for some customers. No previously sent messages were redelivered.
During the outage, we rolled back a feature flag that had led to the increased average row size and GIN index write pressure that were a root cause of this incident. We are leaving that change in place for the near future so that our system remains in a known stable state. Before rolling out this change again, Knock engineers will investigate what threshold of write volume this GIN index can absorb without risking another outage of this type.
Knock engineers will also investigate tuning autovacuum settings for the jobs table in the affected database to ensure autovacuum cannot cause such a critical backlog.
Knock’s Platform Team is already in-flight on a migration to move workflow execution off of a Postgres-backed job queue onto a different set of queueing technologies.
We are redesigning our procedure for SQS DLQ redrives for queues processing critical work, such as message delivery. We will no longer allow full DLQ redrives for these queues, rather requiring the redrive to scope to a specific timeframe and/or customer subset.