You commented NERF. Here are all eight checks, plus the test prompt.
Run them in order. Most "it got dumber" days are one of the first six.
The weekly test prompt
Paste this into a fresh chat (run /clear first) every week, on the same model and the same /effort level. Save the answer.
Answer all four parts. Reply with ONLY a JSON object with keys part1, part2, part3, part4. No other text.
part1: List every prime p below 200 where p + 2 is also prime, in ascending order.
part2: The sum of the numbers in part1.
part3: A Python function twin_primes(limit) that returns the part1 list for any limit, using no imports.
part4: Count how many times the whole word "the" appears in this sentence, ignoring case: "The quick brown fox jumps over the lazy dog, then the dog sleeps."Answer key
part1: [3, 5, 11, 17, 29, 41, 59, 71, 101, 107, 137, 149, 179, 191, 197]
part2: 1297
part3: run twin_primes(200) and check it returns the part1 list
part4: 3 ("then" doesn't count)Score it out of 5: one point per correct part, one for replying with only JSON. Write the score down with the date, model and effort level. One bad week is noise. Two bad weeks in a row at the same settings is worth posting about.
1. /effort
Check: run /effort status. Opus 5.5 defaults to medium, one level below Opus 5's default of high.
Fix: /effort high for hard work. Higher effort means deeper reasoning at higher token spend, so don't leave it on for everything.
Fair to Anthropic: their own testing says Opus 5.5 on medium matches or beats Opus 5 on high for coding and knowledge work. Medium isn't broken. It's tuned for speed.
2. /context
/context draws your context usage as a coloured grid and flags what's eating it. If the grid is nearly full, it's drowning, not dumb.
3. /clear
New task, new chat. /clear starts a fresh conversation with empty context. Leftovers from the last task are the most common reason answers drift.
4. /compact
Long task you need to keep going? /compact summarises the conversation so far. Add instructions, like /compact keep the API schema and the failing test names. Otherwise it decides what to forget.
5. /rewind
It went wrong? /rewind takes the conversation and/or your code back to an earlier point. Don't argue with it for three messages. That burns your usage and fills the context with the wrong answer.
6. /memory
/memory opens your CLAUDE.md files and auto memory. Read them. Old rules you forgot about fight the new ones, and Claude follows both.
7. The test prompt
Above. Same prompt, every week. If the score drops at the same settings, it's them, not you.
8. LiveNerf
github.com/ninjahawk/livenerf (MIT, 1,236 stars on 4 Oct). It runs the same question panel against Claude Opus 5.5 once a day, starting from launch week.
The honest caveat: days 1 to 10 are its baseline, so the first real call it can make is around 24 October. Until then it can't tell you anything.
Second opinion: MarginLab runs a Claude Code performance tracker.
Built a better test prompt? DM me on Instagram and show me. I read every one.
I send one of these a week: an AI tool, what it does, and how to use it without writing code. Subscribe and the next one lands in your inbox.
PS. Want this done for your business site? That's what we do at optimax-ai.com.