Skip to main content

Your browser is older than this site supports, so this page is showing without its design. To see it properly, update to Safari 15.4 or later, or to a current version of Chrome, Edge or Firefox.

Back to Blog
·Updated ·19 min read·By Emanuele Pugliese

We Collapsed Our CI Gates and Bought a Server. Here Is What Got Faster.

Listen to this article
Illustration. On the left, a swarm of small CI jobs, each tagged with a stopwatch reading a few seconds and a coin marked 1 minute billed. They funnel into one wide job card labelled Org checks: one job per push. On the right, a server labelled ci-runner-1, 8 cores, 6 job slots, with check jobs flowing into it, while a separate lane labelled Deploys stays on GitHub-hosted runners. Headline: Stop paying for the round-up. Subline: One gate job per push, one server for every check, two dials to route them.

GitHub bills every job rounded up to a whole minute. A five-second gate costs the same as a fifty-nine-second one, and on every push we were paying for dozens of them.

A month ago we published a breakdown of our GitHub Actions bill, with a section called What we deliberately did not do. The first item on that list was we did not collapse jobs. Granularity bought parallelism, precise failure names and reuse; the answer to round-up was “cheaper compute, not a worse design.”

Four weeks later the organisation had used 88,405 Linux minutes in 27 days, spent $504 and hit 90% of its Actions budget. So we reopened both questions. We collapsed the jobs that were only small because of history, and we bought the cheaper compute.

This is what we changed, what we measured, and, as last time, what we decided not to do.

Measure first, again

Same method as before. List every job of every run through the API, sum ceil(job_seconds / 60) per job, and group by workflow and job name. Our model reproduced 98% of the minutes GitHub actually billed, which is the test that tells you the model is complete enough to plan with.

What the list showed was not one big offender but a long tail. Across 60 repositories, every push ran the same handful of tiny gates:

  • branch hygiene: refuse pushes to a branch whose pull request is already merged;
  • caller contract: every call to a shared workflow passes the inputs it must pass;
  • docs standard: architecture decision records and changelog entries are well formed;
  • isolation: no workflow reaches into another environment’s account.

Each one did five to fifteen seconds of work, and each one was billed a full minute, on every push to every branch in every repository.

Five small gate jobs of 5 to 15 seconds each, every one rounded up to a billed minute (5 billed minutes), next to one job running the same five checks as steps in 15 seconds (1 billed minute).

Learning 1: one gate job per push

The four gates became steps of one job, in one reusable workflow that every repository calls the same way:

# .github/workflows/org-checks.yml in each repository: the whole caller
name: Org checks
on:
  push:
    branches-ignore: [main]
jobs:
  check:
    uses: your-org/shared-cicd/.github/workflows/org-checks.yml@main
    secrets: inherit

Inside the reusable (simplified), the order and the conditions matter more than the YAML suggests:

jobs:
  org-checks:
    name: Org checks
    runs-on: ${{ inputs.runs-on || vars.RUNS_ON || 'ubuntu-latest' }}
    timeout-minutes: 10
    steps:
      - name: Branch hygiene          # first: the one that must never be skipped
        run: ./hygiene.sh
      - name: Caller contract
        if: ${{ !cancelled() }}      # run even when an earlier step failed
        run: python3 contract/check.py .
      - name: Docs standard
        if: ${{ !cancelled() }}
        run: python3 docs/check.py
      - name: Isolation
        if: ${{ !cancelled() }}
        run: python3 isolation/check.py

Three details made this safe rather than merely cheap:

  • if: !cancelled() on every step. A failed step does not hide the others, so one run still reports every problem, just as four jobs did. What you lose is four separate red crosses; what you keep is every finding in the log.
  • It stays push-triggered. Branch hygiene is a push guard. On a pull_request event it is vacuous, because the open pull request is itself the proof the branch is live. The other checks are cheap enough to ride along with it.
  • Checks that were advisory became blocking. Folding them into a required job made them required. We made that an explicit decision, not a side effect, and fixed the repositories that failed before switching.

Saving: about 14,800 billed minutes in 27 days, the largest single item, from a change that removed no check at all.

Learning 2: a skipped job passes a required check, and renames have an order

Last time we wrote that GitHub treats a skipped check run as passing. We hit the same rule again, from the other side.

A required check is matched by name, for example check / Org checks. Renaming the job therefore means editing the organisation rulesets, and the order is not obvious. The repository that hosts the reusable workflow calls it by a local path (./), so its own pull request reports the new name immediately. Merge the workflow first and that pull request can never satisfy the old rule. Change the ruleset first and every other open pull request is briefly waiting on a name it does not report yet.

We did it ruleset first, then the workflow, then one push to each of the open pull requests. The general rule: never give a job that carries a required name an if: of its own, because skipping it passes the gate. Put the condition on the steps inside it.

Learning 3: bots edit pull requests, and edited re-runs the gate

Our release gates listened for pull_request types including edited. Our release tooling rewrites the release pull request’s description every time it refreshes the list of changes. Every rewrite re-ran the gate on an unchanged tree.

Dropping edited from all 57 release gates saved about 2,410 minutes. A retarget of the base branch still reports “Expected — waiting”, so nothing can slip through that way.

Learning 4: cost the replacement before you build it

Two ideas from our own plan died on the spreadsheet, and we are glad they did.

  • Replacing polling with events. A gate that polls every 30 seconds for up to 25 minutes looks like obvious waste. But the event-driven version needs a listener job per deploy, and each listener is billed a minute. For our Dependabot auto-merge poller, events would have cost about 170 minutes against 92 for the poller. We kept the poller.
  • Release-gate waits to events. Worth about $3–4 a month once re-costed, and it needed a write permission in every preprod deploy that reusable workflows cannot grant themselves. Parked.

The habit: every optimisation gets costed with the same model that found the problem, before a line is written. A pleasing architecture is not a saving.

Learning 5: the budget is a kill switch for the whole organisation

Halfway through the rollout, every workflow run in the organisation started failing in two seconds with startup_failure. The cause was not our change: the organisation had hit its Actions spending limit. Then, after the limit was raised, a card payment failed:

The job was not started because recent account payments have failed or your spending limit needs to be increased.

Two things we did not know until that afternoon:

  • A budget stop refuses every job, including jobs routed to self-hosted runners, which cost GitHub nothing. The billing check happens before runner selection.
  • It fails fast and quietly. Nothing pages you, and a red pipeline looks like a code problem until you read the job’s annotation.

Set a budget alert well below the stop, not only the stop itself, and know which account holds the card.

Learning 6: a self-hosted runner is a design, not a VM

Even after the savings, most of what remained was check jobs: tests, builds and browser suites. That is the textbook case for a self-hosted runner, and the case against one is mostly about security and maintenance. So we designed for those first.

We rented one dedicated server: a Hetzner CCX33 with 8 dedicated vCPUs and 32 GB of RAM, at €139 a month. At our effective rate of about $0.006 per Linux minute, that breaks even at roughly 27,000 minutes a month, and our check jobs alone run well above that.

Architecture. GitHub dispatches check jobs to the runner group self-hosted-ci. On the server, a systemd unit mints a one-hour registration token from a GitHub App whose only permission is Self-hosted runners; the App key never leaves the host. Six ephemeral containers each take one job and are destroyed, with a read-through npm cache on a private network. Deploy jobs follow a separate dial and stay on GitHub-hosted runners.

Security, because a runner is handed credentials

  • Ephemeral. Each of the six slots is a container that registers with --ephemeral, takes exactly one job and is destroyed. The next job starts on a filesystem that has never seen the previous job’s checkout, dependencies or tokens.
  • A GitHub App with one permission. Registration uses an App whose only permission is Self-hosted runners. Its private key stays on the host, readable only by root. A container receives a one-hour registration token, never the key.
  • No Docker socket. Mounting the socket would make every job root on the host. Jobs that genuinely need Docker (sam build --use-container, image builds) stay on GitHub-hosted runners.
  • Private repositories only. The runner group refuses public repositories, so a fork’s pull request can never land on our hardware.
  • SSH by key, from one address, password login off, automatic security upgrades on.

Speed, because otherwise why bother

  • The image starts from Microsoft’s Playwright image, so browsers and their ~40 system packages are already there. Downloading them on every run was the single largest slice of our E2E minutes.
  • A pre-warmed tool cache holds Python 3.11–3.14 and Node 20/22/24 in the exact layout actions/setup-python and actions/setup-node look in first, so those steps take a second instead of a download.
  • A read-through npm cache (Verdaccio) sits on a private Docker network. Our own @your-org/* scope is refused there, so private packages always come straight from GitHub Packages and can never be shadowed by a public name.

Control, because the next decision should be one line

Every job in the organisation now chooses its runner from one of two organisation variables:

# check jobs: tests, builds, E2E, release gates
runs-on: ${{ vars.RUNS_ON || 'ubuntu-latest' }}

# deploys, promotes, publishes, teardowns
runs-on: ${{ vars.RUNS_ON_DEPLOY || 'ubuntu-latest' }}

A script rewrote 514 jobs across 58 repositories in one pass: 248 checks onto the first dial, 266 deploy-type jobs onto the second. RUNS_ON is set to the server’s label. RUNS_ON_DEPLOY is deliberately unset, so every deploy still runs on GitHub-hosted runners. Moving production deploys onto hardware we operate means that machine holds production credentials, and that is a separate decision with its own review. When we take it, it is one command and no workflow edits.

Two practical notes from the rollout:

  • Widen the variable repository by repository. We scoped RUNS_ON to three canary repositories first. Four repositories were held back because the version of their workflows on main still sent production deploys through RUNS_ON; they joined once their releases carried the new dials.
  • Some checks need the host itself. Our own unit-file syntax check runs systemd-analyze verify, which needs a real systemd. That one job is pinned to GitHub-hosted, with a comment saying why.

Learning 7: what got faster, and what did not

We re-ran five of the day’s heaviest pull-request runs, originally on GitHub-hosted runners, on the server. All 15 jobs passed.

Bar chart comparing GitHub-hosted and self-hosted times. Chat widget build: 264 s hosted, 117 s self-hosted. Frontend tests: 67 s and 58 s. Small gate job: 15 s and 15 s. Python backend tests, both on the self-hosted runner: 210 s on one core, 62 s on four cores with pytest -n 4.

Job GitHub-hosted Self-hosted
Chat widget build (npm install + Playwright) 264 s 117 s
Frontend tests 67 s 58 s
Small gate (Org checks) 15 s 15 s
Python unit-test suites 66–239 s 5–20% slower

The npm- and browser-heavy jobs won, because the expensive part of them was downloading things the image already has. The small gates did not move: they are dominated by job start-up and checkout, and a minute is billed either way on hosted runners. That is exactly why learning 1 mattered more.

The Python suites were the surprise. Our first guess was package downloads, so we measured a PyPI caching proxy on the server before building it. Installs got slower: 13.3 seconds through the warm proxy, against 9.9 seconds direct. The install steps already took the same time on both kinds of runner. The cost was running the tests, and pytest runs on one core by default.

On one 210-second suite:

pytest workers Test step
1 (default) 210 s
4 (pytest -n 4) 62 s
8 (pytest -n auto) 68 s

Two lessons came out of that table:

  • The faster machine did nothing for a serial workload. Parallelism did. Four workers was the sweet spot; beyond that, starting each worker costs more than it saves on a suite this size.
  • Cap the container, not the job’s ambition. We first capped each slot at 4 of the 8 cores. That is a ceiling a lone job cannot exceed even when the machine is idle, and tools that size themselves from nproc still see 8 cores and oversubscribe. We now allow each slot all 8 cores. When several jobs run at once, the kernel’s equal CPU shares split the machine fairly between them.

Learning 8: add slots before servers, and use GitHub’s runners as the failover

Added 28 September 2026, the day after the switch.

Once all 60 repositories sent their checks to the server, the obvious next questions were: should we also use the minutes GitHub includes in the plan, and should we add more servers? We measured before answering.

Runner Jobs sampled Median wait to start 90th percentile Worst
Our server (6 slots) 151 2 s 122 s 486 s
GitHub-hosted 177 2 s 4 s 109 s

In the bursts, jobs waited up to eight minutes for a free slot, while the CPU sat 30–55% idle and 27 GB of memory was free. The slots were the bottleneck, not the machine. Most check jobs are short gates that spend their time on start-up and the network, not on the CPU.

So the first step cost nothing:

  • Ten slots instead of six.
  • An 8 GB memory cap per job. It is a ceiling, not a reservation: a runaway job is killed on its own before it can take the host down.
  • Every job now logs its own peak memory. A runner hook prints it at the end of each job’s log. The first jobs on the new slots peaked at 105–171 MiB, so the next sizing decision will come from data, not a guess.

Scale slots before servers. Each step is taken on a measurement: queue waits, CPU idle, logged memory peaks. Step 1, done: more slots on the one server, 6 to 10 slots, an 8 GB cap per job, and every job logs its peak memory. Step 2, next: automatic failover; when no runner is online twice in a row, RUNS_ON becomes ubuntu-latest and the owner is emailed. Step 3, when the CPU saturates: a second server, same runner group, same label, no workflow change. Step 4, only if load is large and spiky: servers per job, created when jobs queue and destroyed after. Beside the steps, Deploys: their own track, on GitHub-hosted runners today; if they move, a separate, small machine, never beside check jobs; RUNS_ON_DEPLOY. Footer: GitHub’s runners are the failover, not an overflow: a job’s runner is fixed when it is queued.

Why not use GitHub’s included minutes as overflow? It sounds free, but three things argue against it:

  • GitHub cannot do it. A job’s runner is fixed when the job is queued, and there is no fallback from a busy self-hosted runner to GitHub’s own. You would need a router: an extra job, billed a minute, to choose the runner for every real job.
  • The prize is small. The most a router could ever save is the included allowance, about $18 a month on our plan.
  • The allowance is already spent. Our deploys, Docker jobs and release machinery still run on GitHub, and they use it up.

GitHub’s runners are more valuable as the failover. One server is a single point of failure: if it goes offline, every required check in 60 repositories just waits in the queue. Because the runner choice is one organisation variable, the fix is already built in: set RUNS_ON back to ubuntu-latest and every check runs on GitHub again. We are automating that flip:

  • A small scheduled function on AWS checks every five minutes whether any of our runners is online.
  • If none is, twice in a row, it flips the variable and emails us. It flips it back after fifteen healthy minutes.
  • It uses its own GitHub App, with just two permissions: read the runners, write the variables.
  • It deliberately does not run as a scheduled GitHub workflow. Every five minutes, that would bill about 8,600 minutes a month to guard against an outage.
  • One thing no failover fixes: a budget stop or a failed payment refuses self-hosted jobs too (learning 5).

More servers come later, and in order. A second always-on server joins the same runner group with the same label: GitHub hands each job to any idle runner, and no workflow changes. It is justified when the CPU, not the slot count, is saturated at peaks, and it also removes the single point of failure. Servers created and destroyed per job are cheapest at scale, but they add a webhook receiver, cleanup of stray machines, and a registration key on every one. That only pays when the load is large and spiky.

And the deploys? They are now the biggest thing still billed: 18% of September’s minutes, about 17,400 a month, roughly $85–105. It is tempting to point RUNS_ON_DEPLOY at the same server. We won’t:

  • Check jobs run code from every branch and every dependency update.
  • Deploy jobs hold credentials for production.
  • Containers on one machine share a kernel, so a check job that escaped its container could read a deploy job’s credentials. A GitHub-hosted runner gives every deploy a fresh virtual machine.

If deploys move, they go to a separate, small machine with its own runner group, open only to the repositories that deploy. Deploys mostly wait on CloudFormation, so the machine can be small. Preprod deploys move first, as a canary, and the owner still approves every production release. We wrote this down as an architecture decision before building any of it.

What we deliberately did not do

  • Deploys stay on GitHub-hosted runners. Moving them is worth about $85–105 a month, but only onto a machine that runs nothing else (learning 8), and only after its own security review.
  • No overflow router onto GitHub’s runners. GitHub cannot fall back from a busy self-hosted runner, and the most it could save is about $18 a month.
  • No second server yet. The CPU was idle while jobs queued, so more slots came first. They cost nothing, and the next measurement decides whether a second machine is needed.
  • No pip cache. Measured, slower, not built.
  • No event-driven Dependabot merges. Measured, more expensive, not built.
  • No weaker gates. Every check that ran before still runs, and three advisory ones now block.
  • No bet on ubuntu-latest staying still. It moves to Ubuntu 26.04 on 19 October 2026. We ran our suites against the ubuntu-26.04 label first. The one real risk was Playwright: versions older than 1.62 do not support 26.04, and every suite we run locks 1.62.1 or newer.

The results

  • About 21,600 billed minutes saved per 27 days by design alone: one gate job per push, no re-runs on edited, and the security checks folded from four jobs into one. That is roughly a quarter of the bill, with every check still running.
  • Every check job in all 60 of our repositories now runs on one flat-rate server, with deploys still on GitHub, and a one-line switch for either dial.
  • Capacity grows from measurements: ten slots on the same server came first, because jobs queued while the CPU was idle, and GitHub’s runners become the automatic failover.
  • Browser- and npm-heavy builds up to 2.3× faster, and a clear, measured path for the Python suites, which is parallelism rather than more hardware.

Last month we wrote that the right answer to round-up was cheaper compute, not a worse design. Half of that held up. The other half turned out to be design after all: four tiny jobs were never really four jobs, they were four steps we had never had a reason to put together.

A checklist you can run this week

  1. Sum ceil(job_seconds / 60) over your jobs for a month, grouped by job name. Check that it reproduces your bill.
  2. Find the jobs under 20 seconds that run on every push. They are candidates to become steps of one job, if nothing requires them by separate name.
  3. Remove edited from pull_request types on any gate a bot can trigger by rewriting a description.
  4. Cost every replacement with the same model before building it.
  5. Set a budget alert below your stop, and know that the stop halts self-hosted jobs too.
  6. If you self-host, route with organisation variables, keep deploys on their own dial, make runners ephemeral and register them through an App.
  7. Profile a slow test job before buying hardware for it. If it is one process on one core, the fix is -n 4, not a bigger server.
  8. Measure queue waits and CPU before buying a second server. If jobs wait while the CPU is idle, you need more slots, not more machines.

We help teams design CI/CD that stays fast and affordable as the estate grows: shared workflow libraries, release lanes, runners you can trust with credentials, and the measurements that say which change is worth making. Get in touch if your Actions bill has started to look like ours did.

Related Articles

Comments

Comments are powered by GitHub Discussions. Sign in with GitHub to leave a comment.

Basket is empty