Know in 10 seconds if your prompts are good enough.
Prompt Janitor scans every AGENTS.md and CLAUDE.md on your Mac, grades them A–F against the industry's own standards, and flags what's rotting before your agents trip on it.
What a clean prompt file gives you back
Prompt health isn't cosmetic. Every defect taxes tokens, turns, and your own time, on every single run.
Stop paying the defect tax
In our first controlled runs, one missing example cost an average of 36,759 extra tokens per task. Same task, same agent, one defect apart. Every prompt defect is a recurring bill: it charges you again on every run, in every repo. Finding and fixing them is the cheapest optimization you haven't done.
Less babysitting after the run
In the same runs, the fixed prompt left 0.4 fewer major issues per task for a human to clean up. The point of an agent is that you review once, not line by line.
Agents converge faster
Ambiguous context makes agents wander: 0.8 extra turns per task with the defective prompt. Clear roles, examples, and output contracts get to done sooner.
Watch the waste shrink
Health history and per-project trends ship in the app. Next up: an impact view that translates every fixed defect into tokens and turns saved.
Impact view on the roadmapLearn from your real sessions
Mining your actual agent transcripts to show where a prompt file cost you money, and what changed after the fix. Grounded in your usage, not our lab.
On the roadmapBenchmark figures come from our first N=5 controlled runs and are not yet statistically significant. Read the methodology.
We measure what bad prompts actually cost.
We run controlled benchmarks: the same coding task, the same agent, one prompt defect apart. Then we count the damage.
with one defective prompt
to finish the same task
after the prompt was fixed
Early numbers from our first controlled runs (N=5), not yet statistically significant, and we say so. The full powered benchmark runs next, and we're publishing everything, methodology included.
Read the methodology →Visibility you've never had
Until now there was no single place to see how good your prompts actually are. Prompt Janitor is that layer.
Every prompt, graded A to F
One health score per file, rolled up per project. Watch grades rise and fall as you edit, so you never have to guess whether a prompt is actually any good.
Catches what's rotting
Stale model names, contradictory instructions, missing examples, walls of text. Every issue is cited to its source (Anthropic, OpenAI, or the practitioners who wrote the playbook) and comes with a suggested fix.
Your standards, enforced
Start from trusted rule packs, then write your own in plain English. “Never name a specific model version.” Done, and checked on every scan, every file.
Quiet by default
Scans in the background on your schedule. A glance from the menu bar, a calm weekly digest, and alerts only when something regresses. Never naggy.
Scan. Grade. Treat.
Diagnosis is free forever. Treatment is what you pay for.
Scan
Point it at your projects. It finds every prompt file (CLAUDE.md, AGENTS.md, .cursorrules) and rescans on a schedule.
Grade
Each file gets an A–F health grade against source-cited standards from Anthropic, OpenAI, and the practitioners who wrote the playbook.
Treat
Pro rewrites the weak parts with AI: apply with a backup, one-click undo, and an optional git branch so changes stay reviewable.
Custom rules, plain English
House rules live right alongside the built-ins: type the intent, pick a severity, done.
Grouped by project & source
Filter your whole prompt estate by repo, file type, or which guidebook flagged the issue.
Built for everyone with a CLAUDE.md
If you write prompts for agents, you need to know they're holding up.
Solo developers
Keep your one CLAUDE.md sharp without thinking about it.
Eng teams & leads
One view of every repo's prompt health across the org.
AI engineers
Hold your agent prompts to a measurable standard.
Agencies
Audit client prompt files and prove the quality.
Indie hackers
Ship faster knowing your prompts won't quietly rot.
Anyone, really
Got a prompt file on disk? Now you can see how good it is.
“Prompt files are infrastructure.
Nobody inspects them.”
Your agents read these files on every single run, yet there's no linter, no review, no grade. We think diagnosis should be free, for everyone, forever. Treatment is what you pay for.
Read the manifesto →Diagnosis free. Treatment paid.
We grade your prompt files against the industry's own standards: free, unlimited, on your machine, with your compute. Grading against YOUR standards, and fixing anything: that's Pro.
- Unlimited scanning: scheduling, watch mode, notifications, history & trends. No scan caps, ever
- All deterministic fact rules, with source-cited findings, never hidden or blurred
- Built-in 25-standard AI catalog evaluation, free when you bring compute (local Ollama or your own API key)
- Standards updates keep flowing to free users
Launching soon: the waitlist gets the download first. No payment, no account.
Everything in Free, plus:
- AI rewrites + Apply fix / Auto-fix: backup, undo, optional git branch
- Custom rules in plain English: your standards, enforced on every scan
- Starter template packs: A-grade CLAUDE.md, AGENTS.md & .cursorrules per stack
- Bonus: Prompt-File Field Guide (PDF): the 25 standards, explained
- 12 months of feature updates
$19 is the launch sale price; it goes to $30 afterwards. Waitlist members lock in the sale price.
This pricing isn't definitive and may still change before launch.
One-time purchase: perpetual license + 12 months of updates · $29/yr optional renewal, never required.
Frequently asked questions
Privacy, compatibility, and how grading actually works.