How to benchmark GitHub Actions runners fairly
Compare CI runners on the same YAML, warm caches and repeated runs, measure cost per job instead of rate per minute, and see what our benchmark shows.
Guzman Pintos
10 min read
Every runner provider publishes a chart where its bar is shorter than GitHub's, and so do we. Those charts are hard to compare with each other, because a runner benchmark has a lot of knobs: which workflow, which machine size on each side, whether the cache was warm, how many runs, and whether "cheaper" means the per-minute rate or the bill for the job.
This post covers the choices that make a runner comparison fair, how we ran ours, what the numbers show (including the rows that don't flatter us), and a workflow you can copy to test your own jobs on both runners.
What a fair runner benchmark controls for
Run the same YAML on both sides and change only runs-on. If one side gets a custom cache action, a prebuilt image or a trimmed test suite, you're comparing configurations instead of runners.
Then warm the caches. A first run on a new runner downloads everything and fills an empty cache, and since your normal CI doesn't run like that, a cold-cache number overstates the cost of every job. Tenki hosts its own drop-in Actions cache (caching), so a cache warmed on GitHub is empty on Tenki and the reverse. Run each side at least once before you start recording.
Repeat the runs, too, because a single run tells you little. For short jobs the noise can be bigger than the difference you're measuring: a job that takes 5 seconds on one runner and 3 on another differs by two seconds, and a slow package download or a busy host can move it by that much.
State the machine size on each side plainly. Two vCPUs on one provider and four on another is a legitimate comparison as long as you say so, but it answers a different question than an equal-core test does (more on that below).
Finally, compare cost per job rather than rate per minute. The number that ends up on your bill is duration × per-minute rate, with each provider's rounding rules applied. A runner with a higher rate can still cost less per job if it finishes fast enough, and a lower rate doesn't help if the job runs longer.
How we ran ours
The benchmarks page has the full setup. In short:
- The Tenki side ran on
tenki-standard-medium-4c-8g, 4 vCPU and 8 GB on bare-metal x64 Linux, with one fresh VM per job. - The GitHub side ran on
ubuntu-latestwith 2 vCPU. - Both used identical YAML with only
runs-onchanged, and no caching plugins, prebuilt images or substituted actions. - Each workload ran three consecutive times per runner, on a warm cache, in April 2026.
- The workloads are real builds from public open-source projects: Rust, Docker, Node.js, Go, Android and the n8n monorepo.
If you reproduce this, decide up front which number you report: the last warm run or the median of the warm runs. The first of three runs mostly seeds the cache, and the two rules can give different results, so say which one you used.
"No caching plugins" doesn't mean Tenki ran without caching. Tenki's dependency mirrors for npm, APT, Go modules, Cargo, Maven, Gradle and Docker Hub are on for every job with nothing to configure (caching). They're part of what you get on Tenki, so they're in the Tenki numbers.
Equal cores or equal price
This benchmark puts Tenki on twice GitHub's cores: 4 vCPU against 2. The two are close in price, but not equal. Tenki charges $0.002 per core-minute (pricing), so the 4-core runner is $0.008/min. GitHub's 2-core Linux runner is $0.006/min (GitHub's runner pricing).
So the Tenki side costs a third more per minute. At per-second rates, it only comes out cheaper per job if it finishes in less than 75% of GitHub's time, meaning at least 25% faster. For jobs under a few minutes, GitHub's rounding up to the whole minute lowers that bar, as the cost table below shows.
The other fair comparison is equal cores: Tenki's 2-core runner at $0.004/min against GitHub's 2-core at $0.006/min, or Tenki's 4-core at $0.008/min against GitHub's 4-core larger runner at $0.012/min. At equal cores Tenki's rate is 33% lower, so it's cheaper even if the job runs at the same speed. Our published benchmark doesn't include an equal-core run, so it can't tell you how much of the speedup comes from faster hardware and how much from the extra cores. If that matters for your decision, run both comparisons.
The results
Durations below are from the benchmarks page, plus the chainwayxyz/citrea workflow from the Runners product page:
| Workload | GitHub (2 vCPU) | Tenki (4 vCPU) | Time |
|---|---|---|---|
Rust cargo build | 5s | 3s | 40% faster |
| Docker build | 27s | 19s | 30% faster |
Node.js npm install | 10s | 8s | 20% faster |
| Go build | 10s | 4s | 60% faster |
Android assembleDebug | 1m 38s | 1m 2s | 37% faster |
| n8n monorepo (full CI) | 55m 58s | 29m 15s | 48% faster |
| chainwayxyz/citrea | 37m 3s | 12m 12s | 67% faster |
Speedups vary a lot by workload, from 20% to 67%, so there is no single "Tenki is X% faster" number that applies to your pipeline.
What each run costs
The two providers round differently. GitHub rounds each job up to the next whole minute (pricing reference), so a 5-second job is billed as 60 seconds. Tenki bills per second of job runtime (limits and billing). The table shows both views: GitHub prorated to the second, which isolates the rate and speed difference, and GitHub as billed, which is what the invoice shows. Tenki's cost is the same in both.
| Workload | GitHub, prorated | GitHub, billed | Tenki | Difference, prorated | Difference, billed |
|---|---|---|---|---|---|
Rust cargo build | $0.0005 | $0.006 | $0.0004 | 20% less | 93% less |
| Docker build | $0.0027 | $0.006 | $0.0025 | 6% less | 58% less |
Node.js npm install | $0.0010 | $0.006 | $0.0011 | 7% more | 82% less |
| Go build | $0.0010 | $0.006 | $0.0005 | 47% less | 91% less |
Android assembleDebug | $0.0098 | $0.012 | $0.0083 | 16% less | 31% less |
| n8n monorepo (full CI) | $0.3358 | $0.336 | $0.2340 | 30% less | 30% less |
| chainwayxyz/citrea | $0.2223 | $0.228 | $0.0976 | 56% less | 57% less |
Read the sub-minute rows as an illustration of the rounding rule rather than a measured bill. The benchmarks page doesn't say whether they time the whole job or only the build step, and a whole job also includes runner setup and checkout. If they're step times, Tenki's real cost per job is higher than shown. GitHub's stays at $0.006 as long as the job finishes within a minute. The billed column also treats each workload as a single job. GitHub rounds every job in a workflow separately, so a workflow with many jobs, such as a full CI run, is billed more than one rounded total.
Rounding matters most for short jobs. On the long builds it adds a fraction of a minute to an hour-long bill, and the billed and prorated columns match. On a job under a minute, GitHub charges the full minute however quickly the job finishes, and those jobs add up: lint, typecheck and each leg of a matrix are billed separately.
Prorated, the npm install row costs more on Tenki: at 20% faster it misses the 25% break-even. Billed, it costs less, because GitHub charges a full minute for a 10-second job. On a longer network-bound job, where rounding stops mattering, the break-even applies again, and Tenki's 2-core runner at $0.004/min is the one to test.
Per run, the long builds still move the bill the most. A 60% gain on a 10-second Go build saves about half a cent, while the n8n and citrea workflows save 10 to 13 cents per run, multiplied by every push that triggers them.
What drives the differences
This benchmark changes the hardware, the core count and the download path at once, and it doesn't separate them.
On hardware, Tenki's x64 runners are bare-metal AMD EPYC servers, where each pair of vCPUs maps to one physical core (x64 runners), with one VM per job.
The extra cores help work that runs in parallel, like compiling many packages or modules, and do little for work that waits on the network or runs on one thread (which runner to use). The compile-heavy rows gained the most and npm install the least, which fits that pattern, though the benchmark doesn't break down where each job spent its time.
Downloads are the third difference. Package installs and FROM pulls go through Tenki's dependency mirrors, so repeat downloads don't cross the public internet, and Docker pulls avoid Docker Hub's anonymous rate limits (caching). How much that contributes depends on how much of your job is downloading.
If you want to isolate one factor, change one thing at a time: an equal-core run shows the hardware effect, and switching off the Actions cache for the workspace in the dashboard shows how much your jobs lean on it.
Benchmark your own workflow
You can run this comparison on your own job in an afternoon. Copy the workflow you want to test into a new file on your default branch (workflow_dispatch only works from there, per GitHub's docs), and turn runs-on into a matrix so both runners run the same steps:
name: runner-benchmark
on: workflow_dispatch
jobs:
build:
strategy:
fail-fast: false
matrix:
runner: [ubuntu-latest, tenki-standard-medium-4c-8g]
runs-on: ${{ matrix.runner }}
steps:
- uses: actions/checkout@v4
# Your existing setup, cache and build steps, unchanged
The matrix value resolves to a single label, which is what Tenki needs: a Tenki label must be the whole runs-on value (sizes). Your repository has to be connected to Tenki first (quickstart).
Run it three times, one after the other, so the later runs find a warm cache on both sides:
for i in 1 2 3; do
gh workflow run runner-benchmark.yml
sleep 10
gh run watch "$(gh run list --workflow runner-benchmark.yml --limit 1 --json databaseId --jq '.[0].databaseId')" --exit-status
done
Then pull each job's duration from the GitHub API and apply each provider's billing rules:
gh run list --workflow runner-benchmark.yml --limit 3 --json databaseId --jq '.[].databaseId' |
while read -r id; do
gh api "repos/{owner}/{repo}/actions/runs/$id/jobs" --jq '
.jobs[]
| ((.completed_at | fromdate) - (.started_at | fromdate)) as $s
| if .labels[0] == "ubuntu-latest"
then [.labels[0], $s, (($s / 60 | ceil) * 0.006)]
else [.labels[0], $s, ($s / 60 * 4 * 0.002)]
end
| @tsv'
done
Each line prints the runner label, the job's duration in seconds, and its cost: GitHub's rounded up to whole minutes at $0.006/min, Tenki's per second at $0.002 per core-minute × 4 cores. Drop the oldest run, since it filled the caches, and compare the two remaining runs on each side.
Adjust the numbers to what you're testing:
- For other sizes, change the
4to the Tenki runner's core count, and use GitHub's rate for the size you're comparing against: $0.012/min for its 4-core Linux runner, $0.022/min for 8-core (pricing reference). - In a public repository, standard GitHub-hosted runners are free (Actions billing), and
ubuntu-latesthas 4 vCPU and 16 GB instead of 2 vCPU (runner specs). The cost comparison above doesn't apply there, and the fair speed comparison is against a 4-core Tenki runner. - For jobs under a minute, run more than three times before trusting a small difference.
Job duration excludes queue time, so these numbers measure the work itself. If time-to-first-feedback matters to you, also compare how long each run takes from dispatch to completion.
Once you have per-job costs, GitHub Actions cost optimization covers turning them into a monthly bill, including Tenki's plan fees and credits. For the full methodology and current results, see the benchmarks page.


