Skip to content
WhittleOSWhittleOS
← All guidesValidation tools4 min read

Do startup idea validators work? We tested ours and published the misses

Vendors say yes, forums say no, and nobody has run a test. We ran 8 real companies with known endings through our own gate under two founder profiles — it caught 4 of 5 flops, and the honest conclusion is still not an accuracy percentage.

By Boris Binyaminov · ·

Companies tested
8, real outcomes, anonymized
Flops caught
4 of 5, solo lens
Accuracy claimed
None, and that is the point
Run on
2026-06-23

Testimonials and forum jokes do not answer this question. A row-level test can, so we ran one: 8 real companies with known endings, anonymized into one-liners, put through our own gate. It caught 4 of the 5 flops. We still will not tell you an accuracy percentage, and that refusal is the most useful thing on this page.

What we did

Eight companies whose stories are over — five that failed and three that worked. Each was reduced to the one-liner it could have been described by at founding, with the name stripped, and fed to the gate under two founder profiles: a solo founder, and a funded full-time team. The known outcome and its public source sit in the table beside our verdict.

Only KILL counts as catching something. A HOLD or a PILOT_FIRST is the gate declining to decide, and scoring those as saves is precisely the arithmetic that lets every tool in this category call itself accurate.

ReallyOutcomeSolo lensFunded lens
QuibiFAILEDKILL 0KILL 0
JuiceroFAILEDKILL 0HOLD 39
Google Glass (consumer)FAILEDKILL 0HOLD 40
WebvanFAILEDKILL 0KILL 0
single-platform-API research toolsFAILEDHOLD 58PILOT_FIRST 58
StripeSUCCESSKILL 0PILOT_FIRST 55
SlackSUCCESSHOLD 46PILOT_FIRST 52
FigmaSUCCESSHOLD 55PILOT_FIRST 64

Run 2026-06-23 on Claude Opus 4.8. Every row, including the ones that went against us — the tallies below are computed from this table, so a summary cannot drift from the evidence it summarises.

The tallies, including the ones that hurt

Solo lens

4 / 5

flops killed — and 1 of the 3 successes killed too.

Funded lens

2 / 5

flops killed, 0 successes killed. Deliberately less kill-biased, and it shows.

Counted from the rows above, with KILL as the only verdict that counts as catching something.

Three rows were never killed on either lens: single-platform-API research tools, Slack, Figma. Two of those are successes, which is the outcome you want. The first is not — it is a flop that walked through both lenses, and it stays in the table.

The solo lens also killed a genuine success. That one is not an error we are hiding: a payments company is a correct KILL for a person working alone, because the verdict is bound to the founder rather than to the idea. The same one-liner gets a different answer for a funded team, which is the whole design and also the reason a single accuracy number for a founder-bound tool is incoherent.

Why there is no percentage

Every instinct in marketing says to divide 4 by 5, print the result, and move on. Here is why that number would be a lie in four separate ways.

The sample is eight, and hand-picked. Chosen because their outcomes are known and famous, which is a selection rule that has nothing to do with the population of ideas a real user brings.

Hindsight leaks. The cases are anonymized, but a model may still recognise a famous shape. We cannot prove it did not, so we cannot treat a hit as clean.

The cells move. Each deal-breaker verdict is a five-sample supermajority vote, which makes the gate materially steadier than one pass — but the weighted score is still a single pass, so a borderline row can drift a point or two between runs. This is a dated snapshot you can re-run, not a statistic.

We catch duds without always naming the cause. An adversarial audit found our stated reason is sometimes adjacent to the real one. The claim is catch, not diagnosis.

Any of those alone would make a percentage misleading. Together they make it marketing.

What to ask any validator, including this one

Will it show you its misses? Not a case study — the full table, with the rows that went the wrong way still in it. A tool that only publishes wins has published nothing.

Is the verdict bound to you or to the idea? If the same one-liner gets the same answer whoever asks, the tool is rating ideas in the abstract, which is not the question you have.

Can you check a claim? A score with no source under it is a confident number, and a confident number is the cheapest thing an AI can produce.

Does it ever say no? A tool that has never returned a verdict its user did not want is not an instrument. Ours is kill-biased by design, and the cost of that is on this page: it killed a success.

The short version

  • 4 of 5 known flops killed on the solo lens, 2 of 5 on the funded lens. Every row published, including 3 nobody killed at all.
  • No accuracy figure, because n=8, the cases were hand-picked, hindsight may leak, and the cells move between runs.
  • A founder-bound verdict cannot have one accuracy number: the same idea is correctly a KILL for one person and a PILOT_FIRST for a team.
  • The test to apply to any tool in this category, ours included: make it show you the rows where it was wrong.