3 repos that stop
Claude Code eating
your limits alive.
You commented REPOS. Here are the three links plus everything the reel didn't have time to say.
The limits problem is real. But most people are solving the wrong version of it. They hit the wall and immediately start rationing prompts or splitting into shorter sessions. That's the wrong fix. The actual problem is waste. Claude Code is generating tokens you never needed, burning through context on overhead you never measured, and running skills that haven't been active since day one.
These three repos attack the problem at the source. One compresses what Claude says back to you. One shows you exactly where your tokens are disappearing. One pulls the design system of any site on the internet so Claude builds something worth the token spend in the first place.
Run them in that order and your weekly limit becomes functionally irrelevant for most builds.
Caveman
The reel covered the 75% compression stat. Here's what it didn't cover: where the savings actually come from. Claude's default output is padded with "Let me think through this," "Based on what you've shared," "Here's a summary of what I did." You scroll past all of it. Caveman strips every token of that narrative while keeping every line of code, every file path, every technical detail exactly intact.
Real benchmark from the repo: 294 tokens average response vs 1,214 in normal mode. That's a 65% drop in output tokens per turn. Across a full day of builds, the difference compounds fast.
What the reel couldn't fit: Caveman has five modes, not three.
| Mode | Command | What it does | Best for |
|---|---|---|---|
| lite | /caveman lite |
Light filler removal. Grammar intact. | Client demos, readable output |
| full | /caveman |
~65% reduction. Default mode. | Daily builds, most workflows |
| ultra | /caveman ultra |
Maximum compression. Terse to the point of blunt. | Sprint sessions, known codebase |
| wenyan | /caveman wenyan |
Classical Chinese compression. Statistically the most token-dense written language ever developed. | Absolute token emergency |
| wenyan-ultra | /caveman wenyan ultra |
Peak. Ancient scholar on a budget. | You'll know when you need it |
/caveman:compress CLAUDE.md and it rewrites the file into compressed caveman format, saving a separate human-readable backup. Tested at 46% input token reduction per session. Your CLAUDE.md loads on every session start. Make it small forever.
CodeBurn
The reel said CodeBurn shows you where your tokens go. Here's the level of detail it actually gives you. This isn't a summary. It's a breakdown by project, by activity type, by MCP server, by tool, by shell command, and by session. It classifies every Claude Code turn into one of 13 activity categories and shows you the one-shot success rate per category. That last metric is the one that actually changes how you build.
One-shot rate = what percentage of turns Claude got right without a retry. If your "implement feature" category is at 40% one-shot, you're burning roughly 2.5x the tokens you should be on every feature. The fix is almost always a tighter prompt or a missing context file.
The hidden money: most power users start at 50-70K tokens of overhead per session before they type anything. System prompt, tool definitions, loaded skills, active MCPs, CLAUDE.md. CodeBurn makes this visible so you know which MCPs to kill.
codeburn report and look at the MCP servers panel first. Any MCP you haven't actively used in the last 5 sessions is burning context on every turn for nothing. Kill it. Then look at your CLAUDE.md token count. If it's over 2,000 tokens, it got bloated. Trim it or run caveman-compress on it. Those two fixes alone typically recover 30-40% of your session budget before you change anything else.
codeburn report --format json | jq '.projects' pulls your cost breakdown per project. Run it after optimizations to prove the savings are real.
Design Extract
The reel covered brand voice, responsive behaviour, hover states, and motion language. Here's what's actually in the full output. One command against any live URL generates 19 output files covering every layer of the design system. It runs a headless browser against the live DOM and computes everything from rendered styles, not source CSS.
Section 16 is the one nobody else extracts. The reel mentioned it but didn't have time to explain what motion language actually means in this context.
That feel fingerprint is the actual unlock. Instead of describing an animation style in vague terms and hoping Claude interprets it right, you paste one token from the report and Claude knows precisely what easing family, what duration bucket, and what keyframe kind to use.
--emit-agent-rules to the command. It auto-generates a CLAUDE.md.fragment from the extracted design system. Paste that fragment into your project CLAUDE.md and Claude now has the full design language loaded as context every session. No more re-describing the aesthetic on every new component. The system knows.
The Limit
Stack
These three repos are designed to compound. Here's the order that extracts maximum value from minimum setup time.
codeburn today to see the immediate drop.
npx skills add juliusbrussee/caveman, then /caveman:compress CLAUDE.md. Set full as default in your CLAUDE.md. Your next session will feel noticeably faster.
npx designlang [url] --emit-agent-rules, paste the fragment, and start the build. Your Claude interactions now arrive with full design context pre-loaded. Less back-and-forth. Fewer correction loops. Lower token spend per feature.