As a result, at 04:30 UTC the on-call engineer made the decision to mitigate the lack of capacity by deploying a new compute node within the region and manually culling certain high-memory workloads from the existing compute nodes to free up capacity. At 10:12 UTC we applied considerably more aggressive resource limits to the specific user workload and attempted a phased restart of the compute node’s user workloads, which succeeded in restoring the affected compute node.