Skip to content
WhittleOSWhittleOS
← All guidesValidation tools7 min read

An AI rejected your idea. How to tell whether it is right

Before you argue or give up, find out which check failed, whose constraints it was applied to, and what would overturn it. On our own published cases, the same one-liner came back with opposite verdicts under two different founder profiles.

By Boris Binyaminov ·

Cases
Eight, dated, with known outcomes
Split by founder
Three of the eight
Our own misses
Published, not hidden

A rejection is a claim, and a claim can be checked. Three questions settle it: which specific check failed, whose constraints it was applied to, and what evidence would overturn it. The middle one matters more than people expect — on our own published cases, 3 of 8 identical one-liners came back with opposite answers under two different founder profiles. One of those is a documented success our own solo profile rejected.

That table is a dated snapshot rather than a statistic. The ratings are sampled, so a re-run hands back a fresh draw and a borderline cell can move — which is a reason to read a verdict for its named check rather than its number, and the limits travel with it further down this page.

What a rejection is made of

A useful gate does not return a mood. It returns a named check that failed and a reason for the failure, and those are two different things to argue with.

Our discovery gate runs 10 of them, published in full, and the single-idea check runs a longer list. If the tool you used cannot tell you which one failed, you have not been given a rejection — you have been given a rating, and there is nothing in it to appeal. That alone is worth knowing before you spend a week rewriting a pitch.

So the first move is mechanical: find the check. Then ask whether the reason under it is a statement about the market, about the product, or about you.

The same idea, a different founder, the opposite answer

We run a small benchmark of eight anonymized one-liners with known real-world outcomes, and we run each one twice: once as a solo founder on nights and weekends, once as a funded team. Same text, different constraints. Three of the eight come back with opposite verdicts:

Case

What it really was

As a solo founder

As a funded team

PressPacksJuicero · failedKILL 0HOLD 39
DailyGlassesGoogle Glass (consumer) · failedKILL 0HOLD 40
DevPayStripe · successKILL 0PILOT_FIRST 55

The rows of our published benchmark where the two founder profiles disagree, from a run on 2026-06-23 using Claude Opus 4.8. The input text is byte-identical between the two columns — only the founder changes. The remaining 5 cases agreed and are on the benchmark page with their reasons.

The last row is the one to sit with. A payments company that became one of the largest software businesses in the world is rejected outright for a solo founder on nights and weekends, and held for a pilot when the same sentence is read for a funded team. Both readings are defensible, because licensing, compliance and banking partnerships are not a nights-and-weekends project — and neither reading is about whether the idea was good.

That is the most common way an AI rejection is right and useless at the same time. It is answering "can you build this" when you asked "is this worth building". Check which question yours answered before you throw anything away.

What our own gate got wrong

The same benchmark is where we publish our misses, so here they are in one sentence: of 5 documented flops, the solo lens rejected 4 and the funded-team lens 2; of 3 documented successes, the solo lens rejected 1 and the funded-team lens 0.

We attach no accuracy figure to that, and the limits are not a footnote — they are part of the result, so they travel with it:

  • No accuracy percentage. N = 8, hand-picked for known outcomes, anonymized (a model may still recognize a famous case), and hindsight is imperfect. The only claim is: matched the known outcome on these dated, anonymized cases.
  • Our verdict is founder-bound — that's the product, not a bug. The same idea gets a different read for a solo founder vs a funded team, so we show both profiles. A "Drop it" on a venture-scale success under the solo lens means "not for this founder," not "bad idea."
  • The funded-team lens is deliberately less kill-biased. It holds capital-heavy or under-specified ideas for evidence rather than killing them, so it catches fewer of the historical flops than the solo lens does. The solo lens is the strong dud-catcher.
  • Each deal-breaker check verdict is a 5-sample supermajority vote, so the gate is materially more stable run-to-run than a single sample — but the weighted score is still one pass, so borderline cells can drift a point or two. It's a dated snapshot you can re-run, not a statistic — the cells move between runs.
  • Kill My Idea returns a uniform "Drop it" on every idea by design — it's the adversarial "argue why not to build this" mode. Single Idea Check is the comparable, discriminating column shown here.
  • We catch the duds; we don't always name their exact cause. An adversarial audit found our cited reason is sometimes adjacent to the real one. We claim catch, not diagnosis.

All 6 caveats our benchmark carries, imported from the module that owns them rather than paraphrased — the same list the benchmark page prints under its own table. The fourth and the last are the ones that bound every number above.

What the set does establish, within those limits, is a direction: a kill-biased gate reads a hard build as a hard no, and a genuinely hard business that happened to work will sometimes land on the wrong side of it.

Knowing the error direction is what makes a rejection usable. If yours came from a tool that publishes no misses at all, you have no idea which way it leans, and no way to weigh it.

What actually overturns each kind of rejection

The reason attached to the check tells you what evidence would move it. Most appeals fail because they answer with something adjacent:

If the reason was

What overturns it

What does not, however much it feels like it should

No documented problem behind itA public post, thread or review from somebody who is not youYour own experience of the problem, however real
No repeatable way to reach the buyerOne channel where you have already reached the buyer twiceA list of channels that exist
Cannot charge enough to be worth one person's timeA comparable product with a published price, or a buyer who has paidA price you would be willing to charge
Needs hand-holding for every customerA first session that produces a result with no configurationA plan to write documentation later
Leans too hard on one platformA second route to the same buyer that already worksThe platform having been stable so far

A classification, published so it can be argued with row by row. The third column is the useful one: every entry in it is the thing a founder reaches for first, and none of it is evidence somebody outside your head can check.

There is a fourth possibility worth naming, which is that the check was applied to the wrong idea. Tools read the sentence you gave them. A one-liner that describes the category rather than the wedge gets rejected for the category, and rewriting the sentence honestly is not gaming the result — it is correcting the input.

If you still think it is wrong

Say so, on the record. Every Discovery run we publish takes a challenge: you pick the stage you think went wrong — the sources it read, the screening, the scoring — and write what it missed. The form is on the published run, and it is deliberately incapable of editing the verdict. A tool that lets a complaint change a score has stopped being a gate.

What it does do is create a record of where our rejections are contested, which is the only honest route to fixing the ones that deserve it.

The short version

  • Find the named check. A rating with no check behind it cannot be appealed and should not be weighted.
  • Ask whose constraints it was applied to. 3 of our 8 benchmark cases flip verdict on the founder alone, including a documented success our solo lens rejected.
  • Match your evidence to the reason. Experience of the problem does not answer "no documented problem", and a list of possible channels does not answer "no repeatable way to reach buyers".
  • If the input was wrong, fix the input. If the verdict is still wrong, contest it somewhere it gets recorded.