
Sandbox
How to benchmark inference providers across coding-agent harnesses
Test any model or inference endpoint across coding-agent harnesses like Claude Code and OpenCode on Terminal-Bench, with every trial in a Tenki sandbox.
6 min read
Guides and research on GitHub Actions, CI/CD, and AI code review
from the team building Tenki's runners, code reviewer, and agent sandboxes.

Sandbox
Test any model or inference endpoint across coding-agent harnesses like Claude Code and OpenCode on Terminal-Bench, with every trial in a Tenki sandbox.
6 min read

Sandbox
Omnara now supports Tenki Sandboxes as a machine pool provider. Add a Tenki API key and your Omnara agents get their own Linux VMs on demand.
4 min read

Sandbox
Naïve's AI employees incorporate companies, launch campaigns and write code. Every task they run now gets its own fresh Tenki sandbox.
2 min read

Sandbox
Tenki swaps a placeholder for your real API token outside the sandbox, so agent code can call an API without ever holding the credential.
10 min read

CI/CD
Tenki sticky disks keep dependency and compiler caches on a disk that follows the cache key, so jobs start warm without restoring a tarball.
7 min read

Sandbox
Run your agent in a Tenki Sandbox and drive an Anchor-managed browser with Playwright, with the API key kept out of images, templates and snapshots.
8 min read

CI/CD
Compare CI runners on the same YAML, warm caches and repeated runs, measure cost per job instead of rate per minute, and see what our benchmark shows.
10 min read

Sandbox
Give an AI agent or computer-use model its own Chrome in a throwaway VM, driven over CDP with Playwright or Puppeteer and watchable through noVNC.
8 min read

Sandbox
Sync your uncommitted working tree to a warm Linux sandbox with Crabbox and run your test suite there, so your laptop and your coding agent stay free.
9 min read

Sandbox
Your laptop, a container, a CI runner or a disposable VM? Where to run agent-generated code and untrusted MCP servers so a bad result stays contained.
9 min read

CI/CD
Import a .p12 into a temporary keychain, install profiles, notarize with notarytool or use fastlane match, without leaving signing keys on the runner.
9 min read

AI Code Review
How Claude Code, Codex, Copilot and Cursor read AGENTS.md and CLAUDE.md, what to put in the file, and the anti-patterns that make it useless.
9 min read

CI/CD
Reduce GitHub Actions costs with runner right-sizing, caching, concurrency controls, workflow audits, spend alerts and a lower per-minute rate.
14 min read

CI/CD
How self-hosted GitHub Actions runners register and autoscale with ARC, an 8-point hardening checklist, and the operational work you take on.
12 min read

CI/CD
Copilot code review on private repos bills AI credits and GitHub Actions minutes. What each review costs, how to measure it, and which runners it can use.
8 min read

AI Code Review
How Tenki's code review benchmark scores seven AI reviewers bug by bug on 50 real pull requests, and what per-bug scoring does and doesn't show.
7 min read

CI/CD
AI coding agents open more PRs and push more often, and every push runs CI. How to model the extra cost, measure the agent share, and cut it.
10 min read

AI Code Review
Seven ways code from AI coding agents goes wrong while looking right, with an example of each and what to check before you approve the PR.
12 min read

CI/CD
Install actionlint, read its errors, set custom runner labels in .github/actionlint.yaml, and run it in CI with problem matchers, reviewdog or pre-commit.
12 min read

CI/CD
Run only the CI jobs a change affects, using path filters, Turborepo and Nx affected builds, dynamic matrices, required checks and runner sizing.
17 min read

CI/CD
When to share GitHub Actions config as a reusable workflow or a composite action, covering secrets, runners, logs, nesting limits and versioning.
14 min read

CI/CD
How the GitHub Actions cache matches keys, scopes and evicts entries, and how to cache npm, pnpm, Go, Python, Rust, Turborepo and Docker layers.
16 min read