Browse by type

Your AI coding agent bills by the word and writes like it knows that. Caveman make it stop.
▶️ ThePrimeagen reacts: "No way this actually works"
🏆 #1 on GitHub Trending · July 2026 · 🥇 #1 Repository of the Day on Trendshift · April 2026
#1 on Hacker News · 904 points · 366 comments · #8 Product of the Day on Product Hunt
📄 Cited in CAVEWOMAN, an Adobe Research paper that measured caveman-style output cutting cost 1.4 to 2.4×, up to 3× · 🧪 Tested by JetBrains on 86 real coding tasks: "costs you nothing measurable in quality"
⚡ One command, no account, no API key. npx skills add JuliusBrussee/caveman -g → Quick Start
See it · Quick Start · The Numbers · In the Wild · The Skill · The Proxy · Wrap · When to Skip · Docs
| 🗣️ Normal agent · 69 tokens | |
|---|---|
| > The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I'd recommend using useMemo to memoize the object. | > New object ref each render. Inline object prop = new ref = re-render. Wrap in `useMemo`. |
Same diagnosis. Same fix. Same useMemo. The only thing that died was the throat-clearing.
Code, commands, file paths, and exact error messages never get cavemanned. Only the prose around them does. Security warnings and "are you sure?" confirmations come back in full sentences on their own, then caveman resumes.
Caveman no make brain smaller. Caveman make mouth smaller.
Half the fun is that your agent talks like it just discovered fire. The other half is that it is still right.
A token is what AI billing counts, roughly three quarters of a word. Your agent pays for every token it writes and every token it reads. Most agents write like a cover letter and read like a firehose.
Caveman attacks both ends:
Started as a joke on a Friday in April 2026. Hit 4,000 stars in a week. Now past 100,000, with a research paper, a JetBrains lab test, and a Primeagen reaction video. The joke got serious. The voice did not.
Caveman come in two sizes. Start small.
A rule file that makes your agent answer in caveman. MIT, free forever, works in 30+ agents (Claude Code, Codex, Gemini, Cursor, Windsurf, Cline, Copilot, more). One command:
npx skills add JuliusBrussee/caveman -g
Type /caveman if your agent doesn't wake up on its own. That the whole install. One rock.
Runs on your machine, between your agent and the AI provider, and shrinks what the agent reads before every call. MIT CLI, BSL-1.1 runtime:
npm install -g @caveman-ai/cli && caveman setup --install
caveman claude # or codex · gemini · aider · kilo · qwen · opencode · hermes · openclaw · pi
They stack. Most people start with the small rock and graduate.
More doors into the cave · full installer, Windows, single agents, uninstall
The full installer wires up Claude Code hooks and the statusline badge, finds every supported agent on your machine, and skips agents you no have. Safe to re-run. Needs Node.js 22.13+.
curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/v2.7.0/install.sh | bash
Windows, PowerShell 5.1+:
irm https://raw.githubusercontent.com/JuliusBrussee/caveman/v2.7.0/install.ps1 | iex
Just one agent:
# Claude Code
claude plugin marketplace add JuliusBrussee/caveman && claude plugin install caveman@caveman
# Gemini CLI
gemini extensions install https://github.com/JuliusBrussee/caveman
# Qwen Code CLI, then its Caveman wrapper
npm i -g @qwen-code/qwen-code
caveman qwen
# Codex, Cursor, Windsurf, Cline, and other skills-compatible agents
npx skills add JuliusBrussee/caveman --skill '*' -a codex --yes -g # replace codex with your agent profile
Install broke? Open your agent in this repo and say: "Read CLAUDE.md and INSTALL.md, install caveman for me." Agent read repo, agent fix own brain. Snake eat tail.
Changed your mind: npx -y github:JuliusBrussee/caveman -- --uninstall
The full 30+ agent matrix, dry runs, flags, and verification live in INSTALL.md.
Small rock. The skill, right after npx skills add:
/caveman lite for tight-but-polite. /caveman ultra for grunts. /caveman wenyan for classical Chinese, because someone asked./caveman-commit writes a Conventional Commit in one line./caveman-review gives one finding per line: L42: 🔴 null deref. Guard it./caveman-compress CLAUDE.md cuts the prose, keeps every heading, path, and command, and backs up the original.stop caveman. Normal prose returns. No hard feelings.Big rock. The proxy, right after npm install -g @caveman-ai/cli:
caveman learn reads months of agent history already on your disk, locally, and ranks your token sinks worst-first with a one-line fix behind each. Do this before anything else. It is the most useful five minutes in this README.caveman learn implement hands each fix to Claude Code or Codex one diff at a time, applied only on your yes, and reverts anything that did not lower tokens per turn.caveman claude (or codex, gemini, aider, opencode, pi, …) puts the proxy in front of it. Logs, test output, JSON, and diffs get shrunk before the provider sees them. Originals stay on disk, and the agent can pull any of them back.caveman shrink -- pnpm test compresses command output. caveman browse <url> gives the agent a compressed view of a web page instead of a 15,000-token accessibility dump.caveman trial -- claude runs a real session with and without caveman, then caveman trial report shows the difference. That A/B outranks every number on this page.caveman convert --dry-run shows which installed skills get cheaper as PNG pages the model reads as an image. Convert the profitable ones, revert byte-for-byte any time.caveman stats for history and estimates. /caveman-stats inside Claude Code for that session.Every number below is either from a committed run in this repo or from a named third party. Nothing rounded up. Where a number is small, it says so. Where a row is red, it stays red.
| Who measured | What they measured | Result |
|---|---|---|
| Adobe Research (CAVEWOMAN, arXiv 2606.24083) | Eight models, five datasets, five compression levels | Output-side caveman style cuts realized cost 1.4 to 2.4× per model, up to 3× in the best case |
| JetBrains | 86 real coding tasks, paired A/B, Claude Code 2.1.200. Skill only, no proxy (July 2026, before the proxy existed) | 8.5% fewer output tokens, about 10% cost. No detectable quality change (sign test p = 0.82) |
| This repo (committed eval snapshot) | Ten dev questions, skill vs a plain Answer concisely. control, claude-opus-4-6 |
50% fewer output tokens at the median on top of the terse control. Length only, not correctness |
Read those three together and you get the honest picture. Chat-style Q&A: big cut. Agentic coding sessions, where most tokens are code and tool calls that the skill never touches: high single digits on output, quality flat.
The JetBrains number is why the proxy exists. They measured the skill alone, in July 2026, before the proxy shipped. Their finding was that an agent's bill is mostly reading, not writing, and no talking style fixes that. So we built the thing that shrinks the reading. The table below is what that changed.
The Adobe paper's other finding matters too: compressing the human's prompt into caveman-speak makes models answer longer and worse. Caveman never rewrites your prompts. Only the agent's mouth.
The rules add input tokens on every call, and whether shorter output pays for them depends on your agent, caching, and billing. Full accounting: docs/HONEST-NUMBERS.md.
No reviewed API benchmark result is published here yet. Run
uv run python benchmarks/run.py to generate a new result, then review its raw
response pairs and quality before publishing the generated table.
Your agent rereads logs, test output, diffs, and half your repo all day. The proxy shrinks that stream before it reaches the provider. Pinned 54-run Claude Code benchmark, provider-reported input tokens, three runs per case, every answer checked against an exact oracle:
| Case | Direct Claude Code | Through caveman | Change |
|---|---|---|---|
| CSV outlier hunt | 165,823 | 74,484 | -55.1% |
| Log needle in haystack | 148,807 | 74,068 | -50.2% |
| YAML config drift | 132,124 | 71,027 | -46.2% |
| Test output failure | 150,377 | 108,514 | -27.8% |
| Deployment JSON drift | 147,975 | 108,939 | -26.4% |
| Dashboard HTML alert | 140,687 | 154,641 | +9.9% |
| Total | 885,793 | 591,673 | -33.2% |
18 of 18 answer checks passed. Case-clustered 95% interval: 14.6% to 48.5%. In the same suite, Headroom's wrap saved 6.7% and failed 3 of 18 checks. Method, provenance hashes, and limits: docs/WRAP-BENCHMARK.md. Raw harness artifacts are not in this checkout, so treat it as a pinned report, not a public reproduction.
Maintainer note. The HTML row is red and it stays red. That case had no compression transform, so caveman paid its own overhead and won nothing back. The day I hide a red row is the day you should stop trusting the green ones.
| Surface | Measured | Number |
|---|---|---|
| Browser pages | Focused question against a 200-row table, vs the Playwright ARIA snapshot | 121 tokens vs 15,704. 129.8× smaller. Tiny f |
browse all types & interfaces →
$ claude mcp add caveman \
-- python -m otcore.mcp_server <graph>