Skip to content
WhittleOSWhittleOS
← All guidesValidation tools6 min read

Preuve AI and WhittleOS: two published answers to a weak result

The lazy comparison says only we show evidence, and both products advertise it. The real difference is the default action on a weak result: rework the angle, or end it on a failed check. Both are defensible.

By Boris Binyaminov ·

Read on
2026-09-19
The axis
What to do with a weak result
Not run
We have not used their product

Both products advertise source-linked evidence — theirs on their pricing page, ours on every run we publish — so "only we show evidence" is not a difference we can claim, and this page will not pretend otherwise. What genuinely differs is published, in both directions: what each one tells you to do when the result is weak. Theirs, on their own blog: A low score is a signal, not a verdict. Ours is a list of checks that can fail outright, and a shortlist that comes back empty when they do.

The same situation, two published answers

The situation

What their pages say

What ours does
The score comes back lowRead the risk flags, try a sharper angle, narrow the target, find the competitor gapThe verdict stands and is kept; the checks that failed are named
When to abandon the ideaTheir post names a threshold: three or more tested angles all stallingA failed deal-breaker check ends it immediately, whatever the score
The evidence is thinTheir pricing page offers pivot directions to raise the scoreA market with nothing behind it returns an empty shortlist

Their column comes from preuve.ai/blog/low-validation-score (updated 12 June 2026, read 2026-09-19) and their pricing page, read on 2026-09-19. Ours is what the product does. Neither column is a quality judgement — we have not run their tool, and this compares stated behaviour rather than output.

The specifics in that middle column, with the page each was read from:

  • Four levers they name for raising a low score: read the risk flags, try a sharper angle, narrow the target customer, find the competitor gap (blog/low-validation-score, key takeaways)
  • Their stated condition for moving on: three or more tested angles all stalling under 50 (blog/low-validation-score, key takeaways)
  • Three pivot directions included in the paid report (pricing page feature list)
  • Every key claim linked to its source (pricing page feature list)

Each one recorded in our own competitor module with the page it came from, read on 2026-09-19, so a reader can check the source rather than take the paraphrase. None of it is a quality judgement, and we have run none of it.

Read the middle column charitably, because it deserves it. A weak score often is a positioning problem, their post gives a concrete threshold for when to stop rather than waving at one, and a founder who abandons an idea on a first bad number will abandon a lot of workable ideas. That is a defensible product stance and it is not the one we took.

Ours starts from the opposite default. 10 published checks decide first, and a deal-breaker failing is not a score to be improved by rewording the pitch — it ends the candidate. On the run we publish in full, 36 of 69 candidates were rejected and 5 survived. The reasons are attached to each one.

Which default costs you more when it is wrong

This is the honest way to choose between them, because both defaults have a failure mode and neither vendor's marketing will tell you theirs.

  • A rework-first default fails by keeping you. Four levers and three angles is weeks of work on an idea whose problem was never the angle. The cost is time, and it is invisible while you are spending it, because each rewrite feels like progress.
  • A kill-first default fails by losing you a good idea. A gate that rejects on a missing document rejects real problems that nobody wrote down. The cost is an opportunity you never hear about again — which is why our benchmark publishes the 1 documented success our own solo lens rejected, rather than only the cases where it was right.

Neither is the safe choice. The question is which error you can afford this quarter, and a founder on a deadline with one idea usually answers it differently from a founder with a portfolio and a year.

What each one actually sells

Their planPrice

What their page states

Reality CheckFreeA score, a market overview and two competitor previews
Founder Report$29One 18-section report; packs of 5 for $95 and 10 for $159
Radar Pro$19/monthThe one subscription on their page
Investor-Ready Package$499A deck, memo and model, which their page says is built by hand

From

their pricing page

, read on 2026-09-19. They advertise Source-linked evidence and founder fit in every report. Ours starts at $19 for 2 credits, and a single idea check costs 1 credit against a signup grant of 2 .

The deliverables are different objects. Theirs is a document you read; ours is a decision with the reasoning attached and one recommended next move. If what you need is something to hand to a bank, a document is the right object and we do not produce one.

On evidence, since both of us claim it

Ours is countable, so here is the count rather than the adjective: the published run collected 56 documented problems, and 32 of them carry a link you can open. That gap is the honest part — not every extracted problem keeps a usable source, and a page that says "every claim is linked" while shipping that ratio is overclaiming.

We have not measured theirs, because doing so properly means buying and running their product and publishing the dated result. If we ever do, it goes on the benchmark with the method attached, the same way our own misses are.

When theirs is the better purchase

  • You want a long structured document — their page advertises an 18-section report, and ours is a decision with its reasons, not a document.
  • You want help repositioning an idea you are committed to. Their four levers are aimed exactly at that, and we do not do it.
  • You need an investor deck or a memo. They sell one; we sell nothing of the kind.
  • You would rather be told what to fix than what to drop. That is a real preference and their product is built around it.

Four cases where their product is the right one. The last is the honest summary of the whole page: these are two defensible defaults, not a right and a wrong one.

The short version

  • Both advertise claims linked to sources. That is not a difference we can sell on, and any page telling you it is has not read the other product's marketing.
  • The difference is the default when the result is weak: they publish a rework path with a stated abandon threshold, we publish checks that can end it outright.
  • A rework-first default fails by keeping you on a dead idea; a kill-first default fails by dropping a real one. Pick the error you can afford.
  • Their deliverable is a long document, ours is a kept decision. If you need something to hand to a bank, that settles it without any comparison of quality.