Skip to content
WhittleOSWhittleOS
← All guidesValidation tools8 min read

Validating a startup idea without talking to customers

A conversation gives you an opinion you grade afterwards. A pre-registered numeric bar gives you evidence you cannot re-grade. Here is the method, the five things it can measure, and the much shorter list of what a pass actually proves.

By Boris Binyaminov · ·

Method
Pre-registered async test
Gradable metrics
5, a closed set
What a pass proves
55% of the scorecard

The standard answer to this question is: talk to customers. The useful version is narrower. You are not skipping the customer — you are replacing what you collect from them. A conversation gives you an opinion you grade after you have heard it. A number checked against a bar you wrote down first gives you something you cannot re-grade once you have seen it.

55%

of the weighted scorecard a clean market test can evidence. The rest — support load, platform exposure, cash flow — it cannot reach.

0

categories that traffic on its own evidences. A channel is proven when it delivers a person, not when it delivers attention.

Both numbers are read from the grader, not from us — the rest of this page is how they come out that way.

The problem with the conversation is not the conversation

Interviews are not useless. They are unfalsifiable in the hands of the person who wants the answer.

You ask ten people whether they would pay. Seven say something encouraging. You now hold ten data points and the freedom to decide, afterwards, which of them counted — and you will decide in the direction you were already leaning, because you wrote the questions, you heard the tone, and you are the one who has to abandon eight months of an idea if the answer is no.

That is not a discipline problem you can fix by being more rigorous in the room. The bar moved because nothing was holding it in place. So put something in place.

Pre-registration, which is the whole method

Before the test runs, write down what would count. Not "good traction" — a number, a metric and a deadline. Then run the test and compare. The order is the entire trick, and it is the only reason a founder's own report about their own idea is admissible evidence at all.

This is not a rhetorical device here. Our validation plans emit machine-gradable criteria alongside the English ones, and the grader that scores them is pure code with no model call in it. It can read exactly 5 things:

  • visitors
  • signups
  • paid_conversions
  • revenue_usd
  • replies

A closed set, on purpose. A grader that cannot read a metric cannot grade it, and a criterion that exists only as English prose is marked ungradable rather than quietly interpreted. An ungradable outcome never unlocks a build decision — the honest failure mode is "we cannot tell", not a guess dressed as a pass.

Three rules that do the actual work

A blank is not a zero. If you do not report visitors, that criterion grades as not reported, never as not met. The distinction sounds pedantic until you notice what collapsing it would buy you: an empty form would read as evidence the idea failed, and "I did not measure it" is not that. When every criterion comes back blank, the whole outcome is ungradable and nothing moves.

The window only runs one way. A criterion says reach N within D days. Reaching it faster passes — five pre-orders in three days does not fail a fourteen-day bar. Taking longer does not pass. A founder who needed thirty days to clear a fourteen-day bar produced a different result from the one the plan asked for, and letting that through deletes the only time-box in the loop.

You cannot correct upward. Within one plan, the report that counts is the weakest one filed. Reporting a worse number takes effect immediately; reporting a better one leaves the earlier grade standing. The asymmetry is deliberate — the direction that can only cost you evidence stays open, and the direction that can only gain it is closed, because a bar written in advance is worth nothing if the answer can be rewritten after you have seen how it graded.

Copy this before you publish the page or send the message:

Promise being tested:
Named buyer:
Distribution channel:
Observable action:
Minimum result:
Measurement window:
If the bar is missed, I will:

If any line is blank, the test is not pre-registered. Filling it after traffic arrives turns a decision rule back into a story about why the result was encouraging.

What a pass actually proves, which is less than you think

Here is where most validation advice quietly overreaches. A landing page test that goes well gets described as proof the idea works. It is not. It is proof of some specific things and silence on others, and the difference is worth drawing:

visitorsevidences nothing
— no category is evidenced by traffic alone
signups35% of the scorecard
Pain + Demand Evidence · Distribution Fit
paid_conversions55% of the scorecard
Economics / Pricing / Margin · Pain + Demand Evidence · Distribution Fit
revenue_usd55% of the scorecard
Economics / Pricing / Margin · Pain + Demand Evidence · Distribution Fit
replies35% of the scorecard
Pain + Demand Evidence · Distribution Fit

Notice the top row. Bars are relative to the widest, and both the categories and the percentages come from the same function the product calls when it decides what a reported outcome is allowed to affect — so this chart cannot disagree with the grader.

Traffic evidences nothing at all. Visitors is the one metric in the set that unlocks no category, and that is not an oversight — it was a bug once. An early version credited the channel whenever any criterion passed, so a plan whose only bar was "at least fifty visitors" confirmed that distribution worked, off fifty page views that one forum post produces. A channel is evidenced when it delivers someone — a buyer, a signup, a reply — not when it delivers attention.

The other thing to read off that chart is the ceiling. Even the strongest single result covers 55% of the weighted scorecard, which leaves the rest untouched:

A passing market test evidences at most 55% of the weighted scorecard and leaves 45% untouched.Evidenced by a passing test — 55%Untouched — 45%: Sales & Trust Fit, Operations / Support Load, Risk / Platform / Regulation, Cash-Flow Sustainability + Transferability

Widths are the category weights, read from the scorecard itself. The grey half is what a market test is structurally unable to answer — how much support the thing will need, what it is exposed to, and whether the cash works.

A test that goes well tells you people want it, will pay for it, and can be reached. It tells you nothing about whether serving them fits in your week, or what happens when the platform you sit on changes its terms. Those stay open, and a method that pretended otherwise would be doing the same thing the encouraging interview does.

What it looks like on one plan

Three criteria, frozen before the test. Two weeks later, the numbers come back:

at least 40 waitlist signups within 14 days52 · met
at least 5 pre-orders within 14 days3 · below the bar
at least 12 replies to the outreach sequence within 14 days · not reported

The criteria and the reported numbers are an example; the verdict beside each one is not — the grader produced them. Read the third row: a criterion left blank is marked unanswered, not failed.

The grade that comes back is partially met, and the summary the founder sees says so in plain terms: 1 of 3 measurable criteria met within the planned window. 1 further criterion/criteria in this plan could not be measured numerically and were NOT checked.

That last sentence is doing more than it looks. The plan published 4 criteria in prose and only 3 of them could be expressed as a number, so one bar in this plan was never checked at all. Without that pair of counts on the report, "1 of 3 met" would read as a verdict on the whole plan, while a criterion nobody could measure went unexamined. Showing the reach of a grade is part of the grade.

What this method costs you

It is slower to set up than a conversation and it produces less texture. An interview tells you why — the phrasing people use, the workaround they already have, the thing they assumed you meant. A number tells you only that the bar was cleared. If you have never spoken to anyone in this market, the criteria you pre-register will be uninformed, and pre-registering an uninformed bar is a well-documented way to be precisely wrong.

The honest scope, then. This is not a claim that async testing beats interviews at understanding a market — we have run no comparison and would not publish one without it. It is a claim about admissibility: a result you committed to in advance survives your own motivated reading of it, and a conversation you grade afterwards does not.

The short version

  • The bar goes on the record before the test, or the result is an opinion with a number attached.
  • 5 things can be graded mechanically; anything else stays prose and is marked ungradable rather than interpreted.
  • Not reported is not zero, faster than planned passes, slower does not, and the weakest report filed is the one that counts.
  • Traffic on its own evidences nothing — a channel is proven when it delivers a person, not attention.
  • Even a clean pass covers 55% of what decides an idea. Support load, platform exposure and cash flow are not on the table, and a test that claimed them would be overreaching.