Skip to content
WhittleOSWhittleOS
← All guidesValidation tools6 min read

How to compare idea validators without fooling yourself

Most comparisons in this category put a free score next to a paid report and call the difference quality. Here is what has to be held constant, why a vendor cannot grade its own competitors, and our result on one idea in full.

By Boris Binyaminov ·

Read on
2026-09-19
Published
Our half, and the protocol
Not published
Their outputs — we ran none

We have not run our competitors' products, so this is not a head-to-head and it will not pretend to be one. What it is, is the half of such a comparison that can be published honestly: which modes are even comparable, what has to be held constant for the result to mean anything, and our own answer on one idea in full, so anybody can run the other side and check it against us.

The category error most comparisons start with

Put a free sixty-second score next to a paid long-form report, note that one is more thorough, and you have measured the price difference. It happens constantly in this category, usually by accident, and it is why most comparison articles are unfalsifiable.

Comparable means same job, not same cost. Here is the closest mode on each tool to the job of checking one idea you already have, taken from what each vendor publishes:

Tool

Closest paid mode

Its price

Why the mapping is arguable

DimeADozenStarterFrom $9Their mode is sold as: a long one-off research report on an idea you bring — line that up against a single-idea check and you are already making an editorial call
PainOnSocialStarter$19/monthTheir mode is sold as: pain points mined from reddit, scored and ranked — line that up against a single-idea check and you are already making an editorial call
IdeaProofStarter€19.99Their mode is sold as: validation plus business-plan, brand and ad-creative generation — line that up against a single-idea check and you are already making an editorial call
Preuve AIFounder Report$29Their mode is sold as: a long evidence-linked report on an idea you bring — line that up against a single-idea check and you are already making an editorial call
FounderPalFounder$99, shown at $49.50 with a discount codeTheir mode is sold as: an ai marketing assistant, not an idea validator — line that up against a single-idea check and you are already making an editorial call
IdeaBrowserStarter$499/yearTheir mode is sold as: a browsable database of researched ideas, plus agents and coaching — line that up against a single-idea check and you are already making an editorial call
ValidatorAIIdea validationFreeNo paid mode exists, so there is nothing here to put opposite a paid one

Modes and prices from each vendor's own pages on 2026-09-19. The last column is the honest one: every mapping in this table is arguable, because these products are not doing the same job and lining them up is already an editorial decision.

Read the fourth column before the third. Two of these tools are not idea checkers at all in the sense this comparison assumes — one sells a browsable library, one sells marketing output — and running them on your idea and scoring the result would be a test of the mapping rather than of the product.

What has to be held constant

Hold constantWhy
The wording of the idea, character for characterThese tools read a sentence. A reworded pitch is a different experiment, and rewording between tools is the most common way one of them is made to look better
The founder profile, where the tool accepts oneThe same idea is a business for a funded team and impossible on six hours a week. A tool that never asks is answering a different question, and that is a finding rather than a fault
The dateLive-search tools read a moving web. Two runs a month apart are not the same trial
The mode, by job rather than by priceA free score against a paid report is a category error dressed as a result
Who reads the outputs, and against what rubricWritten before the outputs arrive, or the rubric fits whichever result you preferred

Five controls, authored and published so they can be argued with. The second is the one this category ignores most: several of these tools never ask who is building the thing.

The last row is the one that decides whether the exercise is worth doing at all. Write the rubric before the outputs arrive, or it will quietly become a description of whichever answer you liked.

Our half, published

The idea, verbatim: InvoiceRescueAutomated late-payment follow-ups for freelancers and small consultancies — pulls invoice data from Stripe/FreshBooks, sends personalized reminders, escalates politely without burning the relationship. The buyer: Freelancers and small consultancies (1-10 people) who send invoices and currently chase payments manually via email.

What our single-idea check returned: PILOT_FIRST at 73, with medium confidence, across 12 deal-breaker filters. Its own weakest point, in its words: Pain + Demand Evidence (5/10): The pain narrative is credible and the kill filters support willingness to pay, but there are no public demand signals, customer quotes, or usage evidence included here.

Its recommended next move: Charge 3–5 named buyers for an async paid pilot before broader build. The whole thing, including the filter table and the rubric breakdown, is on the published report — which is the point. You cannot check a comparison against a result somebody describes but does not show.

What would make the whole exercise worthless

Including the ways we could rig it, since we are one of the tools in the table:

  • The person running it sells one of the tools, and grades the outputs themselves.
  • The idea is one of the vendors' own examples, which some tools will have seen.
  • Only the verdicts are compared, with no record of the reasons each tool gave.
  • A tool that refused to answer is dropped instead of being recorded as a refusal.
  • The sample is one idea, and the write-up uses the word 'better'.

The first one applies to us and is not solvable by good intentions — which is why our own comparison pages compare what each vendor publishes rather than what each product returns.

The first item is the reason this page exists in this shape. A vendor grading competitors' outputs against a rubric the vendor wrote is not an experiment, and we would not believe it from anybody else.

How to record it so the result survives a reader

The write-up is where most of these comparisons lose whatever validity the run had. Four rules, and they are the same ones we hold our own benchmark to.

  • Publish the input before the outputs. The exact sentence, the profile, the date. A reader who cannot reproduce the input cannot check anything downstream of it.
  • Record refusals as results. A tool that declined to answer, timed out, or demanded a subscription mid-flow has told you something real. Dropping it flatters whichever tools completed.
  • Quote the reasons, not just the verdicts. Two tools agreeing on "promising" for opposite reasons is the interesting finding, and a verdict column hides it completely.
  • State the sample size next to every conclusion. With one idea the honest verb is "did", not "is better at". If a sentence in your write-up would still read fine with a hundred ideas behind it, it is over-claiming with one.

What we do publish instead

One idea is not a sample, so the closest thing we have to evidence about our own behaviour is the benchmark: 8 anonymized cases with known real-world outcomes, run under 2 founder profiles, with our misses in the table rather than omitted from it. It carries 6 caveats and makes no accuracy claim, because with that sample size no honest one is available.

If you run the comparison above, the useful thing to publish is not which tool "won". It is the reasons each one gave, side by side, with the inputs. Reasons can be checked by a reader; a ranking cannot.

The short version

  • Comparable means same job, not same price. Most comparisons in this category compare a free score with a paid report and call the difference quality.
  • Hold the wording, the founder profile, the date, the mode and the rubric constant. Write the rubric first.
  • We publish our own half on one idea — verdict, score, filters, next move — so the other half can be run by somebody who is not us.
  • A vendor grading its competitors is not an experiment. That includes this vendor, which is why our comparison pages stick to what each one publishes.