Ask a chatbot what it thinks of your startup idea and it will help you. That is the whole problem: helpfulness and judgement point in opposite directions. The flattery is not a measured property of every validator; it is the behaviour in the dated comparison below, and the incentive any validator has to guard against when the user wants encouragement.
Four reasons the answer comes back warm
You told it the idea is yours. A model completes the conversation it is in. "Here is my startup idea" sets a frame where the useful next turn is encouragement, and the same text pasted as "a competitor is doing this, what is wrong with it" reliably produces a different answer. Nothing about the business changed between those two prompts.
Agreeable answers were preferred during training. Models are tuned on human preference, and humans rate the response that engages with their plan above the one that dismisses it. That is a sensible objective for almost every task and a disastrous one for a go/no-go.
There is no cost to a false yes. A tool that encourages you loses nothing when you fail eight months later; a tool that kills your idea loses you today. The incentive is not hidden or disreputable — it is just pointing away from the thing you wanted.
Scores drift upward when nothing anchors them. If a number is generated rather than computed from a published rubric, it has no floor to fall to. There is no rule saying what a low rating looks like, so the model produces a plausible-sounding number, and plausible-sounding is high.
What that looked like when we tested it on ourselves
The fair way to make this concrete is to use our own idea, so nobody has to trust our judgement about someone else's. Same one-liner, 3 tools, one day.
One idea — ours — through 3 tools on 2026-06-10. Verdicts quoted verbatim from each tool's own output and imported from the repo rather than retyped here.
Read what the first row actually says. Not a cautious pass — "excellent potential", "among the top ideas we've seen", 72 / 100. Our own two modes returned 15 / 100 and 0 / 100 on the same sentence, and the second short-circuited: it stopped at a deal-breaker rather than producing a score at all.
Here is the limit of that evidence, stated before anyone quotes it at us. This is one idea, on one day, through each tool once. It supports exactly one sentence — these were the outputs — and it does not support any claim about how that tool behaves in general, any rate, or any adjective about the company that makes it. We publish the date so you can re-run it, which is the only thing that makes a sample of one worth publishing at all.
What actually distinguishes a warm answer from a judgement
Three properties, and none of them is tone. A tool can be blunt and still be flattering; what matters is whether anything constrains the answer.
| What to check | Flattering | An instrument |
|---|---|---|
| Whose decision it is | One verdict per idea, whoever is asking | Bound to the founder — the same idea can be a KILL for one person and a pilot for a team |
| Where the number comes from | A score appears, with no visible arithmetic | Computed from published weights you can read and disagree with |
| What sits under a claim | "Sources considered", or a citation that opens a search page | A link to somebody saying the thing — and an explicit mark when there is none |
Our argument rather than a measurement — no tool was scored against this. It is a checklist to apply yourself, and the third row is the one most tools fail quietly.
Is the verdict bound to you? A judgement about an idea in the abstract is a judgement about nobody. The same one-liner should come back differently for a solo founder and a funded team, and if it does not, the tool is rating the idea rather than the decision you have to make.
Are the weights published? A score computed from a rubric you can read is one you can argue with. A number that appears with no visible arithmetic is a vibe with a decimal point. Ours are on a page, which is an invitation to tell us they are wrong.
Can you open the evidence? Not "sources considered" — a link that resolves to somebody saying the thing. When a claim has no source behind it, the honest output marks it unverified rather than letting it read like the sourced ones.
What to do with an encouraging answer
Not discard it. Re-ask it in a frame that removes the flattery, and see whether the answer survives.
Paste the idea as a competitor's product and ask what is wrong with it. Ask what would have to be true for it to fail, then go and check the cheapest of those. Ask which single check it comes closest to failing — the answer is usually the same one you already suspected and were hoping not to hear.
If the encouraging verdict survives all three reframings, it might be right. If it evaporates the moment the idea stops being yours, you learned what the first answer was measuring.
The short version
- The warmth is the objective, not a defect. Helpful and judgemental are different goals, and one of them was trained for.
- A false yes costs the tool nothing today and costs you months. That asymmetry is the whole mechanism.
- On 2026-06-10, one idea through 3 tools: 72 / 100 and "excellent potential" from one, 15 / 100 and 0 / 100 from ours. One run, one day, no rate claimed.
- Judge a tool by whether the verdict is bound to you, whether the weights are visible, and whether the evidence opens — never by how confident it sounds.

