Skip to content
WhittleOSWhittleOS
← All answers

Answers

Can ChatGPT validate my startup idea?

A chat can pressure-test an idea you already have, and for that it is useful. With search or a research mode on it can also fetch pages and cite them. What no chat hands you is a fixed gate that drops candidates and says so, a label on every problem it could not verify, and a saved decision you can hold yourself to. WhittleOS is built around those three.

What a chat is actually good at

Start with the honest half, because a page that pretends otherwise is not worth reading. If you have an idea and you want it stress-tested — the obvious objections, the segments you have not considered, a first cut at who might pay — a general-purpose model does that well, immediately, and for free. Anyone selling you a subscription to do only that is selling you a wrapper.

The prompt that gets you most of the way there is not a secret either: describe the idea, describe yourself honestly, and ask it to argue the case for killing it. You will get a useful hour out of that. We would rather you spend that hour before paying us anything — an idea that dies to a free chat did not need a paid run.

Where each mode stops

There are three different things people mean by asking ChatGPT, and they stop in different places. A plain chat answers from what it already knows: a compressed memory of the public internet with a cut-off date. Ask it whether people are complaining about a specific workflow this quarter and it will produce something plausible in the right shape. That shape is the problem: a fabricated complaint and a real one read identically, and the fabricated one costs you a month.

Search mode fixes the first half. ChatGPT with search on, and Claude with web search, do fetch live pages and show you links, so "it cannot cite anything you can open" is not true of them and this page will not say it. Deep Research goes further: a multi-step report with citations that takes up to half an hour, and for desk research it is genuinely good.

What none of the three does is the part that makes a decision safe to act on. None runs the idea through a checklist whose weights stay put between sessions, so Tuesday's answer is graded like Monday's. None labels a claim it could not back as unverified instead of smoothing it into fluent prose. None drops a candidate for failing a gate and tells you it did, and none keeps a record that later runs, or a real outcome, can check against. Those are properties of a pipeline, not of a conversation, and they are what you pay for.

What going and looking changes

A Discovery run plans a market into sub-markets, issues live searches, and fetches the pages it is permitted to fetch. It considers up to 150 candidates and keeps at most 5. The throwing-away is the product: you are paying for the looking and the discarding, not for the writing.

Every problem it keeps is labelled with where it came from — a page we opened and read, a search result we saw but could not open, or something you pasted in yourself. When we cannot find a source, the label is "not verified" rather than a plausible sentence with no page behind it. That label is the difference that matters: a research mode cites what it found, but it does not attach a verified / not-verified label to each problem it relies on, and nothing in it moves a score when a source is missing.

The judgement is split the same way. A language model rates each area and supplies the evidence; the weighting, the total and the pass/fail gate are fixed in code, so the model never hands over its own overall number. An idea you bring passes 12 deal-breaker checks; a candidate we sourced for you passes a related list of 10, because a sourced idea has to prove a documented problem exists first and one you hand us does not have one to prove yet.

The part we will not claim

We do not claim accuracy, and neither should anything else in this category. There is no ground truth on whether an unbuilt idea would have sold, so nobody can be measured against it. The model's ratings vary between runs exactly as a chat's do — what does not vary is the standard they are judged against, or who does the judging.

So the honest framing is narrow: a rigorous, kill-biased second opinion that shows its work. Not an oracle, and not a machine that turns a weak idea into a good one. Most ideas should die early, and a tool that rarely says no is not protecting you from the three-month mistake.

How to decide which you need

Use a chat when you have one idea, you want it argued with, and you are willing to check its factual claims yourself. That is a real use and it costs nothing.

Use a run when the claims are load-bearing — when you are about to spend months, and "someone is probably complaining about this" needs to become a page you can open. And use it when you do not have the idea yet, which is the case a chat serves worst: one sweep across 30 markets produced 3,430 candidates and kept 93, measured 2026-06-04.

There is a complete real run published with no signup, and a full unedited report beside it. Read those before paying for anything — they are on this site precisely so the decision does not require trusting this page.

Common questions

Can I just paste a better prompt into ChatGPT?

For the reasoning half, largely yes, and we would not argue otherwise. Turn search or Deep Research on and it will fetch and cite pages as well. What a prompt cannot buy is a checklist and weights that stay fixed across sessions, a gate that returns an empty list when nothing qualifies, a not-verified label on every problem with no page behind it, and a saved decision that later runs check against.

Does WhittleOS use a language model too?

Yes, for what models are good at: reading pages, extracting problems, and rating each area with its reasoning attached. What it never does is set the overall score or the pass/fail gate — those are computed in code from a fixed weighted rubric, so the number is auditable rather than asserted.

Will I get the same answer if I run it twice?

Not exactly, and we will not pretend otherwise. The model's ratings sample and do move between runs. The rubric, the weights, the deal-breaker checks and the decision matrix do not move — so what changes is the reading, never the standard it is read against.

Is there a free way to see it work?

Yes. A complete Discovery run and a full idea report are published on this site, unedited and with no signup. Signup also includes a couple of free credits, enough for a founder read and a first idea check before paying anything.

See it rather than take our word

Everything this page describes is published somewhere on this site, unedited and without a signup — a complete run, a full report, and the checks that produced them.

Related