Skip to content
Spryhand

AI & Productivity

Prompt Testing Tracker

Log prompt versions, the model used, test results and notes side by side, so prompt iteration has a record.

Delivery
Instant download
Purchase
One-time payment
Format
  • Google Sheets
  • Excel
Updated
August 2026
Sample preview — illustrative data

Best pass rate

v6 at $0.0215/run

Cheapest that passes

v5 at $0.0043/run

VersionChange
  • v3 — output schema87.5% (+10.0%)
  • v4 — shortened to cut cost72.5% (-15.0%)
  • v6 — larger model97.5% (+7.5%)
  • v9 — drafted, not run— not run

Sample data preview of the Prompt Testing Tracker spreadsheet, shown for illustration only.

What this spreadsheet does

Prompt iterations tested across chat windows or docs are easy to lose track of, and it's hard to tell which version actually performed best.

What you can do with it

  • Keep a record of every prompt version tested
  • Tell whether a change actually helped or quietly regressed
  • Decide between the best score and the version you can afford to run

What's included

Sheet tabs

  • Instructions
  • Comparison
  • Prompt Log

Features

  • A row per prompt version with model, what changed, cases run and cases passed
  • Pass rate against a bar you set, with a verdict per version
  • Each version compared against the one before it, so regressions are visible
  • A drafted-but-untested version excluded from every statistic instead of scoring 0%
  • Cost per run and latency logged alongside the score
  • The best-scoring version and the cheapest version that still passes, side by side

How it works

  1. 1Log each version with the model, what changed, and cases run versus passed
  2. 2Record cost per run and latency too
  3. 3Read the Comparison tab: change vs previous, and the cost of the best version

Who it's for

Builders iterating on prompts for a product feature or workflow.

Frequently asked questions

Does it run the prompts for me?

No. This logs and organises results you have already tested; it does not call any model.

Why keep versions that made things worse?

Because that is what a log is for. A deleted regression gets proposed again in six weeks. Regressions are counted, and the worst one is reported.

How do I record a version I haven't tested yet?

Log it with zero cases run. The pass rate stays blank and the verdict reads "Not run" — reporting 0% would be a claim about a prompt nobody has tried.