Buildkite Certification

Learn · Failure handling

Retries worth having

Some failures are worth retrying and some are not. Buildkite makes you say which.

Automatic retries happen without anyone watching:

steps:
  - key: integration
    command: make integration
    retry:
      automatic:
        - exit_status: "*"
          limit: 2

You can be narrower, which is usually better — retry the failures you know are infrastructural, and let genuine test failures fail:

    retry:
      automatic:
        - exit_status: -1   # agent lost
          limit: 3
        - exit_status: 255  # forced termination
          limit: 2

Manual retries put a button in the UI and leave the decision to a person:

    retry:
      manual:
        allowed: true

Choosing between them

Automatic retries are right when the failure mode is environmental — a lost agent, a spot instance reclaimed, a registry timing out. The retry costs you compute and gets you a correct answer.

They are wrong when the failure mode is a flaky test. A retry there does not fix the test; it hides it, and it hides it in a way that gets steadily more expensive as the suite grows. exit_status: "*" with a generous limit is how a team stops noticing that a quarter of their builds are failing on the first attempt.

If you retry everything automatically, at least look at how often the retry is what saved the build. That number is a measure of how much flakiness you are paying to ignore.

Check

Write a pipeline with one command step, keyed `integration`, running `make integration`. It should retry automatically, at most twice, on any exit status.

Sign in to answer and record your progress.