CPA The GTM eval · by Charlotte Han

How the GTM eval works

This page explains how an eval is produced: what gets measured, where the benchmarks come from, what the machine does, what a human does, and the rules the instrument follows. It also says plainly what we keep private and what this instrument does not do.

The five areas

Every eval looks at the same five areas. Each one starts with a simple question:

The questions are easy to ask. Answering them with evidence is the work, and each area breaks into dozens of checks.

Take the Clay case, built from public sources alone. That meant twelve of the main words Clay uses to describe itself, each checked against how real buyers talk. Four possible buyers, each scored on four separate factors. A price tested against three independent reference points. Every named competitor found and counted. Each fact carries the date it was captured.

A paid eval goes further, because it also looks at the founder's own evidence. The intake alone covers 38 fields across six parts of the business, from what the product actually does to what customers say in their own words.

The full checklist stays private. Every check that finds something in your eval shows its evidence.

Each area's checks add up to a score out of 10, where 10 means we found nothing wrong. The five areas do not count equally: who you sell to counts about twice as much as how you sound. Together they make one score out of 100.

The bands

The score places a company in a band. Every scorecard gives a verdict and a degree:

ScoreVerdictDegree
81 to 100Benchmark positionaligned
61 to 80Sound positionminor drift
41 to 60Exposed positionmaterial drift
21 to 40Misaligned positionserious drift
1 to 20Misaligned positionsevere drift

Above 60 means the position is sound: what needs fixing is the wording, not the strategy. We will not say how rare any band is until we have scored enough companies to make that claim true.

How a score is produced

The analysis runs five times, independently, and we keep the middle result. Any analytical system can give a slightly different answer each time it runs. The middle of five runs is stable. And when the five runs disagree with no clear majority, we claim the lower confidence.

Measuring and writing are two separate jobs, by design. The scoring engine has its own version number, and it does not change when we improve how reports read. Better wording can never move a score. Every report is stamped with the engine version that produced its numbers and the version that wrote its words.

Where the benchmarks come from

A score is not our opinion. It is measured against a benchmark database: what buyers like yours can actually spend, what your category actually charges, how your buyers actually talk, and which words they have learned to distrust.

The database keeps facts and judgment apart. Facts the machine can check, such as a price on a live pricing page, are collected automatically every month and stamped machine-verified with their source and date. Judgment, such as whether a buyer type is underserved, is never automated. It is drafted, reviewed, and signed by a human before it grades anyone.

The rules

Every figure carries its date and source. This market moves monthly. A number that does not say when it was true is not evidence.

Confidence is stated on everything. HIGH means evidenced. MEDIUM means supported but not fully verifiable. LOW means the instrument is telling you it is partly guessing, and it shows you where more evidence would change the answer.

Blanks are never filled with guesses. Anything only a founder could answer is left blank rather than estimated, and the blank limits the confidence we claim.

Self-reported numbers are labeled. A number that arrives without the record behind it, such as a win rate with no CRM report to show, is marked founder-reported in the eval. It is never presented as verified, and it is never dropped. A number far outside the range we can verify against is flagged as such, with the innocent explanation offered first.

Third-party data is attributed. When we calculate using someone else's research, the eval says whose research it is: their estimates, our arithmetic.

Nothing ships unsigned. The machine collects and scores. A human reviews every finding, challenges weak evidence, and signs the verdict. Every eval carries a name, not a model number.

What we keep private

The scoring rules, the weightings in full, the benchmark ranges, and the contents of the database stay private, the way any exam worth taking keeps its answer key private. What you see is your score, the evidence behind it, and the date it was true. The questions are public. The answers to your company are yours. The grading key is ours.

What this instrument does not do

The eval measures the strategy layer of go-to-market: who you sell to, what you charge, what you say. It does not audit the execution layer: how your teams are structured, how the sales motion runs, how retention is operated. When strategy leaves its mark on execution, the eval says so. It will not pretend to measure what it cannot see, and public case studies built from public sources say exactly which figures they could not reach.

Versions

Engine and wording versions are stamped on every eval. Benchmark snapshots are dated by quarter. When the instrument changes in a way that could move a score, the change is recorded, and a fixed reference company is scored again to measure exactly what moved. Corrections to published figures are made in place and noted.