Buildkite Certification

Learn · Concurrency

A gate is not a capacity limit

Two different things make a job wait, and they look identical on a dashboard.

Waiting for capacity. Every agent that could run this job is busy. The fix is more agents.

Waiting for a gate. An agent is free and idle, but the job is not allowed to start because its concurrency group is full. The fix is not more agents — adding them changes nothing at all.

Both show up as queue wait time. Only one of them is a capacity problem.

Why this matters

The instinct on seeing high wait times is to scale up. If the wait is coming from a concurrency gate, scaling up spends money and moves nothing, because the constraint was never how many agents existed.

Before recommending capacity, check what the waiting jobs have in common. If they share a concurrency_group, they are queued behind each other by design and the queue is doing its job.

The same trap has a third form worth knowing: a job that appears to be waiting may be sitting behind a block step, waiting for a person. That is not platform latency at all — it is someone at lunch. Measuring it as queue wait and then buying agents to fix it is a well-trodden mistake.

The question to ask of any wait metric: is this time the platform’s, or a human’s?

Check

Jobs in a pipeline are spending a long time waiting before they start. The step has `concurrency: 1` with a concurrency group. What is the most likely explanation, and what would adding agents do?

Sign in to answer and record your progress.