AI & Productivity
Prompt Testing Tracker
Log prompt versions, the model used, test results and notes side by side, so prompt iteration has a record.
- Delivery
- Instant download
- Purchase
- One-time payment
- Format
- Google Sheets
- Excel
- Updated
- August 2026
Best pass rate
v6 at $0.0215/run
Cheapest that passes
v5 at $0.0043/run
- v3 — output schema87.5% (+10.0%)
- v4 — shortened to cut cost72.5% (-15.0%)
- v6 — larger model97.5% (+7.5%)
- v9 — drafted, not run— not run
Sample data preview of the Prompt Testing Tracker spreadsheet, shown for illustration only.
What this spreadsheet does
Prompt iterations tested across chat windows or docs are easy to lose track of, and it's hard to tell which version actually performed best.
What you can do with it
- Keep a record of every prompt version tested
- Tell whether a change actually helped or quietly regressed
- Decide between the best score and the version you can afford to run
What's included
Sheet tabs
- Instructions
- Comparison
- Prompt Log
Features
- A row per prompt version with model, what changed, cases run and cases passed
- Pass rate against a bar you set, with a verdict per version
- Each version compared against the one before it, so regressions are visible
- A drafted-but-untested version excluded from every statistic instead of scoring 0%
- Cost per run and latency logged alongside the score
- The best-scoring version and the cheapest version that still passes, side by side
How it works
- 1Log each version with the model, what changed, and cases run versus passed
- 2Record cost per run and latency too
- 3Read the Comparison tab: change vs previous, and the cost of the best version
Who it's for
Builders iterating on prompts for a product feature or workflow.
Frequently asked questions
Does it run the prompts for me?
No. This logs and organises results you have already tested; it does not call any model.
Why keep versions that made things worse?
Because that is what a log is for. A deleted regression gets proposed again in six weeks. Regressions are counted, and the worst one is reported.
How do I record a version I haven't tested yet?
Log it with zero cases run. The pass rate stays blank and the verdict reads "Not run" — reporting 0% would be a claim about a prompt nobody has tried.
You might also need
Picked by hand for people who bought this one.
Similar tools
The same kind of tool, for a different job.
Book Reading Journal
Track every book you read — pages, dates and rating — with totals calculated automatically.
- Google Sheets
- Excel
Tenant Screening & Application Log
Track every rental applicant's screening status and income-to-rent ratio in one log.
- Google Sheets
- Excel
Change Order Tracker
Log every change order with cost impact and approval status.
- Google Sheets
- Excel
Flock Health Log
Record health checks, treatments and vet visits per bird.
- Google Sheets
- Excel