I have an app with multiple machines that has been working fine for a couple of years.
These past weeks I’ve seen an issue with machines that won’t start and the only solution has been to manually delete those particular machines. The issue only seems to happens with machines created automatically in AMS during the past couple of weeks (I presume when migrating physical servers or something?). I just saw the issue again not one hour ago.
The machines that I created manually or were created automatically months ago don’t have any issues. These start and stop as expected. I’ve deleted all the unhealthy machines and right now all the machines seem to be working as expected.
As you can see all those machines are at least a month old. I can only assume there’s an issue with the scripting creating new machines in AMS or maybe other regions.
There’s really nothing in the logs that suggests it’s an issue with our code. This app has been running in production for a couple of years now with barely any changes.
The worst thing is that when this happens, the Fly router seems to stubbornly try to keep reaching these failing machines instead of sending requests to a new one. This basically kills our service until I manually delete the failing machines.
I would appreciate if you could look into this as it’s becoming an issue with our customers.
