None
EN
Incident Report: August 27th, 2024
[]
Railway Blog
[10:13 PM UTC] The on-call had identified that the pull request made at 10:04 PM UTC had recreated all instances of the new Railway proxy at once. Given this can be done with zero downtime/without draining the instances, the engineer elected to simply resize the boot disk on Google’s dashboard and create a pull request to make the required boot disk size increases to accommodate our retention window due to increasing traffic as we migrate from our old proxy to our new one. As a result, when the pull request to modify the number of instances of type “newproxy” was merged on Tuesday, the external vendor pulled the older configuration with the incorrect boot disk information and then proceeded to re-create machines, based on boot disk resizing, with live traffic causing the outage.