A chat can pressure-test an idea you already have, and for that it is genuinely useful. What it cannot do is go and look: it answers from what it already knows, gives a different answer each session, and cites nothing you can open. WhittleOS runs live searches, reads the pages it is allowed to read, and every problem it keeps carries the link it came from.
What a chat is actually good at
Start with the honest half, because a page that pretends otherwise is not worth reading. If you have an idea and you want it stress-tested — the obvious objections, the segments you have not considered, a first cut at who might pay — a general-purpose model does that well, immediately, and for free. Anyone selling you a subscription to do only that is selling you a wrapper.
The prompt that gets you most of the way there is not a secret either: describe the idea, describe yourself honestly, and ask it to argue the case for killing it. You will get a useful hour out of that. We would rather you spend that hour before paying us anything — an idea that dies to a free chat did not need a paid run.
Where it stops, and why the gap is structural
The limit is not that the model is not clever enough. It is that a chat answers from what it already knows, and what it knows is a compressed memory of the public internet with a cut-off date. Ask it whether people are complaining about a specific workflow this quarter and it will produce something plausible in the right shape. That shape is the problem: a fabricated complaint and a real one read identically, and the fabricated one costs you a month.
Three consequences follow, and none of them is fixable with a better prompt. It cannot show you where a claim came from, because there is no page it read. It gives a different answer on Tuesday than on Monday, because sampling varies and nothing anchors the standard between sessions. And it starts from zero every time — it cannot compare this idea against the problems it found last week, because there was no last week.
What going and looking changes
A Discovery run plans a market into sub-markets, issues live searches, and fetches the pages it is permitted to fetch. It considers up to 150 candidates and keeps at most 5. The throwing-away is the product: you are paying for the looking and the discarding, not for the writing.
Every problem it keeps is labelled with where it came from — a page we opened and read, a search result we saw but could not open, or something you pasted in yourself. When we cannot find a source, the label is "not verified" rather than a plausible sentence with no page behind it. That labelling is the entire difference, and it is the thing a chat structurally cannot offer.
The judgement is split the same way. A language model rates each area and supplies the evidence; the weighting, the total and the pass/fail gate are fixed in code, so the model never hands over its own overall number. An idea you bring passes 12 deal-breaker checks; a candidate we sourced for you passes a related list of 10, because a sourced idea has to prove a documented problem exists first and one you hand us does not have one to prove yet.
The part we will not claim
We do not claim accuracy, and neither should anything else in this category. There is no ground truth on whether an unbuilt idea would have sold, so nobody can be measured against it. The model's ratings vary between runs exactly as a chat's do — what does not vary is the standard they are judged against, or who does the judging.
So the honest framing is narrow: a rigorous, kill-biased second opinion that shows its work. Not an oracle, and not a machine that turns a weak idea into a good one. Most ideas should die early, and a tool that rarely says no is not protecting you from the three-month mistake.
How to decide which you need
Use a chat when you have one idea, you want it argued with, and you are willing to check its factual claims yourself. That is a real use and it costs nothing.
Use a run when the claims are load-bearing — when you are about to spend months, and "someone is probably complaining about this" needs to become a page you can open. And use it when you do not have the idea yet, which is the case a chat serves worst: one sweep across 30 markets produced 3,430 candidates and kept 93, measured 2026-06-04.
There is a complete real run published with no signup, and a full unedited report beside it. Read those before paying for anything — they are on this site precisely so the decision does not require trusting this page.
Common questions
Can I just paste a better prompt into ChatGPT?
For the reasoning half, largely yes, and we would not argue otherwise. A prompt cannot make it fetch a page it never read, cite a source that exists, or compare today's idea against what earlier runs found — those need a system that goes out and looks, not a better instruction.
Does WhittleOS use a language model too?
Yes, for what models are good at: reading pages, extracting problems, and rating each area with its reasoning attached. What it never does is set the overall score or the pass/fail gate — those are computed in code from a fixed weighted rubric, so the number is auditable rather than asserted.
Will I get the same answer if I run it twice?
Not exactly, and we will not pretend otherwise. The model's ratings sample and do move between runs. The rubric, the weights, the deal-breaker checks and the decision matrix do not move — so what changes is the reading, never the standard it is read against.
Is there a free way to see it work?
Yes. A complete Discovery run and a full idea report are published on this site, unedited and with no signup. Signup also includes a couple of free credits, enough for a founder read and a first idea check before paying anything.
See it rather than take our word
Everything this page describes is published somewhere on this site, unedited and without a signup — a complete run, a full report, and the checks that produced them.

