Blue-green deploy seems to keep adding more machines

I’ve switched one of our projects to the bluegreen deploy strategy, but now it seems like it it never “cleans up” old machines. I want to have 1 web process that stops after there is no activity for a few minutes (this is for a testing environment that is used infrequently) and 1 worker process that is always running. I configured it to use bluegreen and added a health check. After a few deploys I noticed that I now had 8 web machines and 5 worker machines running… how do I prevent this from happening?

Here’s my current config:

[build]
  dockerfile = 'Dockerfile'

  [build.args]
    PORT = '8080'

[deploy]
  strategy = 'bluegreen'

[processes]
  web = 'node server/build/processes/server.js'
  worker = 'node server/build/processes/worker.js'

[http_service]
  internal_port = 8080
  force_https = true
  auto_stop_machines = 'stop'
  auto_start_machines = true
  min_machines_running = 0
  processes = ['web']

  [[http_service.checks]]
    grace_period = '15s'
    interval = '15s'
    timeout = '5s'
    method = 'GET'
    path = '/health'

[[vm]]
  memory = '1gb'
  cpu_kind = 'shared'
  cpus = 1
  memory_mb = 1024

Thanks in advance!

Hi… It would probably help to post the full output of fly deploy, since that usually has a lot of details that aid in troubleshooting.

(And also fly m list both before and after.)

My guess is that min_machines_running = 0 is confusing its count somehow…

Here’s everything from a deploy log after the image build:

Running ******************** release_command: node .yarn/releases/yarn-4.14.0.cjs --cwd server prisma migrate deploy
Starting machine

-------
 ✔ release_command 869592bee04108 completed successfully
-------
-------
Updating existing machines in '********************' with bluegreen strategy

Verifying if app can be safely deployed

[WARN] Machine 8ee054f7124618 [worker] doesn't have healthchecks setup. We won't check its health.
Creating green machines
  Created machine 801e36c62e5428 [worker]
  Created machine d8934e3b092768 [web]

Waiting for all green machines to start
  Machine 801e36c62e5428 [worker] - started
  Machine d8934e3b092768 [web] - started

Waiting for all green machines to be healthy
  Machine d8934e3b092768 [web] - 1/1 passing

Marking green machines as ready
  Machine d8934e3b092768 [web] now ready
  Machine 801e36c62e5428 [worker] now ready

Checkpointing deployment, this may take a few seconds...

Waiting before cordoning all blue machines
  Machine 843e11f47eee18 [web] cordoned
  Machine 8ee054f7124618 [worker] cordoned

Waiting before stopping all blue machines

Stopping all blue machines

Waiting for all blue machines to stop
  Machine 843e11f47eee18 [web] - stopped
  Machine 8ee054f7124618 [worker] - stopped

Destroying all blue machines
  Machine 843e11f47eee18 [web] destroyed
  Machine 8ee054f7124618 [worker] destroyed

Deployment Complete
Checking DNS configuration for ********************.fly.dev
✓ DNS configuration verified

Visit your newly deployed app at https://********************.fly.dev/

The fly list just showed 2 machines before and 2 machines after. It seems like maybe the issue is fixed after I explicitly set the machine count to 1 for both process groups?

I’m worried because I didn’t know it was happening (using git hook deploys so I’m not looking at the output). The client is a non-profit and 5+ machines running when we only need 0-1 is a significant difference for them.

I seem to recall someone saying in the past that min_machines_running = 0 doesn’t play nicely with blue-green, but I can’t find that reference just now. That would be the most likely culprit, I’d say.

Yeah, that’s not a small difference. This is generally the kind of situation that Fly.io’s accident forgiveness policy is intended for, but it would be better to avoid it in the first place.

I would use the default rolling deploys instead of bluegreen unless there was a really compelling advantage to the latter. Blue-green is more of a high-end feature, in a sense, and it has a bit of a reputation for surprises…

Yea, I know that a blue-green deploy doesn’t really make sense for a server that is scaling down to 0, but I wanted to try it out on our test server before potentially enabling it on the production server (which doesn’t scale down). Maybe it’s just not worth the trouble though.

Thanks for your help!

Hi, a failed bluegreen deploy will not clean up created machines (there’s a bunch of reasons - mainly that at some point in the past it did, but the logic was tricky to get right and it resulted in “all machines are gone” scenarios, so instead of exposing people to that, now the decision as to how to clean up is left to the human.

This happens only if the deploy failed, so if it does, you can then do fly machine list to see how many machines you have and clean those up.

Rerunning the deploy usually (but not always) will tell you there are two sets of deployed images and exit, instead of retrying the deploy. However, if the deploy does not change the deployed image (e.g. if you’re redeploying the same image for some reason) this detection won’t work, and it’s in those scenarios that you can get extra machines (if you bg-deploy one machine and it fails, you now have two machines. If you rerun the deploy and it fails again, you’ll end up with 4 machines, and so on).

BG deploys are safe to run with min_machines_running=0; if there’s any evidence to the contrary I’d like to know!