Automatically generated machines won't start

I have an app with multiple machines that has been working fine for a couple of years.

These past weeks I’ve seen an issue with machines that won’t start and the only solution has been to manually delete those particular machines. The issue only seems to happens with machines created automatically in AMS during the past couple of weeks (I presume when migrating physical servers or something?). I just saw the issue again not one hour ago.

The machines that I created manually or were created automatically months ago don’t have any issues. These start and stop as expected. I’ve deleted all the unhealthy machines and right now all the machines seem to be working as expected.

As you can see all those machines are at least a month old. I can only assume there’s an issue with the scripting creating new machines in AMS or maybe other regions.

There’s really nothing in the logs that suggests it’s an issue with our code. This app has been running in production for a couple of years now with barely any changes.

The worst thing is that when this happens, the Fly router seems to stubbornly try to keep reaching these failing machines instead of sending requests to a new one. This basically kills our service until I manually delete the failing machines.

I would appreciate if you could look into this as it’s becoming an issue with our customers.

this is odd, we do not ever create machines automatically (migrated ones keep the same id/name as the previous version that was created by you). could you share a machine-id where you’re seeing this happen so we can take a look?

I can’t share an id because I had to delete those machines. I’d be happy to share more details via email.

What I can tell you is that the faulty machines all were created less than a month ago (according to the dashboard) and I don’t think I have manually created a new machine in over a year.

sure, you can send an email with the app name to (email redacted).