The Railway infrastructure engineering team were able to attribute the root cause of this incident to a newly deployed metrics collection agent that appeared to trigger a CPU core soft lock on these hosts, despite them being well below resource thresholds. At 17:05 UTC, Railway’s engineering team began manual restarts to recover affected hosts. Railway’s engineering team began engaging additional members of the Customer Success and Support teams to communicate the impact to affected customers across the Community Forum, Email, Twitter, and Discord.