Honest benchmark
Does a kill-biased gate actually catch the flops?
We ran 8 ideas with KNOWN real-world outcomes through our own gate. On a solo-founder profile (our user) it killed 4 of 5 historical flops; the funded-team lens is deliberately softer. Here's the data — every miss, both founder profiles, no accuracy claim.
4/5
flops killed · solo lens
3/3
successes passed · funded lens
0
accuracy claims
How this works
Each idea is fed anonymized to our gate under two founder profiles — solo (our ICP) and funded team (pure idea-merit) — and the verdict is compared against the real-world outcome, which is already known.
WhittleOS columns from a run on Claude Opus 4.8 (2026-06-23), reproducible. Showing Single Idea Check — the comparable, discriminating verdict.
The results
Swipe the table sideways for the funded-team column →
| Idea | Solo founder | Funded team |
|---|---|---|
| QuibiFailed | Drop it · 0 | Drop it · 0 |
| JuiceroFailed | Drop it · 0 | Not yet · 39 |
| Google Glass (consumer)Failed | Drop it · 0 | Not yet · 40 |
| WebvanFailed | Drop it · 0 | Drop it · 0 |
| single-platform-API research toolsFailed | Not yet · 58 | Try a paid pilot · 58 |
| StripeSucceeded | Drop it · 0 | Try a paid pilot · 55 |
| SlackSucceeded | Not yet · 46 | Try a paid pilot · 52 |
| FigmaSucceeded | Not yet · 55 | Try a paid pilot · 64 |
Drop it don't build · Not yet get evidence first · Try a paid pilot one paid async pilot · Test it first run a no-call validation · Build it green light
How to read the match with history
On the solo lens the gate caught 4 of 5 documented flops and held the 5th (the platform-dependent tool) for evidence rather than green-lighting it. On the funded-team lens all 3 successes pass without a kill. But “matches history” means something specific and limited — here's exactly what:
- Nothing got history backwards. On the idea-merit (funded) lens, no success was killed and no documented flop was green-lit. The errors it can't afford, it didn't make.
- It's a cautious gate, not an oracle. It hard-kills the clear flops and holds the rest for evidence — and it never says “build”, not even to the winners. It stops you building a dud; it doesn't predict the next Stripe.
- A “Drop it” on a success is founder-fit, not idea-merit. Stripe is killed only on the solo lens — “not a nights-and-weekends build”, not “bad idea”. The funded lens passes it.
What counts as a match: “Drop it” on a flop = caught · “Not yet” / “Try a paid pilot” on a flop = flagged, not endorsed · “Drop it” on a success (solo) = founder-fit, not a miss.
Stripe: “Drop it” (solo) vs “Try a paid pilot” (funded)
A payments processor needs PCI-DSS, money-transmission licensing, and banking partnerships — not a nights-and-weekends solo build. Flip the profile to a funded team and the verdict moves to “Try a paid pilot”. Same idea, different honest answer. That spread is the product: we tell this founder what fits.
Limits & method
- No accuracy percentage. N = 8, hand-picked for known outcomes, anonymized (a model may still recognize a famous case), and hindsight is imperfect. The only claim is: matched the known outcome on these dated, anonymized cases.
- Our verdict is founder-bound — that's the product, not a bug. The same idea gets a different read for a solo founder vs a funded team, so we show both profiles. A "Drop it" on a venture-scale success under the solo lens means "not for this founder," not "bad idea."
- The funded-team lens is deliberately less kill-biased. It holds capital-heavy or under-specified ideas for evidence rather than killing them, so it catches fewer of the historical flops than the solo lens does. The solo lens is the strong dud-catcher.
- Each deal-breaker check verdict is a 5-sample supermajority vote, so the gate is materially more stable run-to-run than a single sample — but the weighted score is still one pass, so borderline cells can drift a point or two. It's a dated, reproducible snapshot, not a statistic.
- Kill My Idea returns a uniform "Drop it" on every idea by design — it's the adversarial "argue why not to build this" mode. Single Idea Check is the comparable, discriminating column shown here.
- We catch the duds; we don't always name their exact cause. An adversarial audit found our cited reason is sometimes adjacent to the real one. We claim catch, not diagnosis.
Where the other tools land
Competitor columns are being collected from each tool directly and will appear here — each tool's own verdict on these same 8 ideas, side by side, raw, vs history.
FAQ
- How is this different from a validator's accuracy score?
- We make no accuracy claim. This is a small, dated set of 8 ideas with known real-world outcomes; we show what our gate actually did against history, including our own non-kills and the cases we'd hold rather than kill. Independent reviewers report that validator scores skew high (around 78/100) (reviewers report; as of 2026-06-21) — so the contrast we're after is behavior, not a percentage.
- Why did WhittleOS kill Stripe?
- Only on the solo-founder profile, and the reason is honest: a payments processor needs PCI-DSS, money-transmission licensing and banking partnerships — not a nights/weekends solo build. On the funded-team profile the same idea passes ("Try a paid pilot"). That contrast is the founder-binding working: we tell THIS founder what fits, we don't rate ideas in a vacuum.
- Why are the ideas anonymized?
- If a tool sees 'Quibi' or 'Stripe' it may recite the known outcome instead of judging. We feed every tool the same anonymized, as-of-founding one-liner so it has to make a real judgment, then reveal the referent and outcome in the analysis. Residual recall risk is real and disclosed.
- Isn't 8 ideas too few to prove anything?
- Yes — and we don't claim it proves accuracy. It's a directional, reproducible snapshot of behavior on cases where the outcome is known. The point is to show a kill-biased tool catching documented flops — the kind a category that reviewers report skews ~78/100 'promising' (reviewers report; as of 2026-06-21) would tend to wave through — with our misses and limits stated plainly.
- Can I reproduce this?
- Yes. The WhittleOS columns come from a published test harness (tests/eval/benchmark-modes.test.ts) run on Claude Opus 4.8; the raw per-idea JSON is dumped locally. The 8 one-liners and the outcome key are published so you can run them through any tool yourself.
See what a kill-biased read looks like
Start with a founder First-Pass, then run a Single Idea Check on your own idea — both covered by the 2 free credits on signup, no card required.

