Skip to content
WhittleOSWhittleOS

Where startup ideas actually come from

A list of startup ideas is worth less than one idea you can defend. This is how candidates get sourced from real complaints, and what survived a 30-market sweep.

By Boris Binyaminov ·

All guides

Everyone has a list of startup ideas. Almost nobody has one idea they can defend for ten minutes under questioning. Those are different problems, and only the second one is worth solving — which is why this is not another list.

Why a list of startup ideas is worth so little

A list is cheap because generating candidates is cheap. Ask any language model for fifty startup ideas and you get fifty, immediately, all plausible, none of them attached to a person who has said out loud that they have this problem.

The expensive part is elimination. It is expensive because it requires going and looking — reading what people actually complain about, in public, with dates and links — and because most of what you find kills the idea you were hoping to keep. A generator has no incentive to do that work and no way to charge you for the disappointment.

So the useful question is not "what are some startup ideas". It is: what would have to be true for this one to work, and is it?

Start from complaints, not from a blank page

The alternative to brainstorming is sourcing. Instead of inventing candidates and then hunting for evidence that flatters them, you start from documented problems and let the candidates fall out of what people are already complaining about.

That inverts the usual order, and the inversion is the whole point. Evidence gathered after you have chosen an idea is evidence you selected; evidence gathered before is evidence you found. The first reliably confirms whatever you started with.

In practice this means reading discussion threads, review sites, job listings and complaint forums for one market, pulling out the specific recurring frictions, and only then asking which of them could support a business one person can build and sell. The corpus behind this site is organised that way — 13 clusters of recurring problems, each grouping complaints that keep reappearing across different sources.

The honest caveat, stated here rather than buried: not every problem in that corpus carries an openable link. On the published Discovery run, 32 of the 56 documented problems have a source you can click; the rest are labelled as our own estimate. A problem without a link is still a signal — it is just not evidence, and the two should never be rendered the same way.

What actually survives

Sourcing produces a lot of candidates. That is the easy half. Here is the hard half, from a sweep run on 2026-06-04 across 30 markets:

Roughly one candidate in 37 was worth writing down. That number is the argument of this entire article. If your process is keeping most of what it generates, it is not a process — it is a brainstorm with extra steps.

It also reframes what a good day looks like. Sweeping a market and coming away with two or three candidates worth investigating is a successful run. Coming away with thirty is a sign the filter is broken.

The seven questions

What does the eliminating? These, applied to every candidate. Any single failure ends it:

Notice what they have in common: not one of them is about whether the idea is clever. They are about whether a real problem exists, whether money repeats, and whether one person can reach the buyer and then survive supporting them. Cleverness is not a constraint solo founders actually hit. Distribution and support load are.

Notice also that several of them are about you. "Can it charge enough to be worth one person's time" has no answer in the abstract — it depends on how much time you have and what you need the business to earn. An idea that is workable for a funded team of four and unworkable for one person after hours is not a good idea with a caveat. It is two different questions wearing one sentence.

Reading a score without over-trusting it

Most tools in this space hand you a number. Here is what a number can and cannot tell you.

The scoring here works in two stages. A model rates seven areas — economics, distribution, sales and trust, and so on — and then the total is computed in code from a fixed weighting rather than written by the model. That second stage matters: asking a language model directly for "a score out of 100" produces a number chosen to sound reasonable, and it tends to cluster in the 70s regardless of input.

What that split buys you is that the arithmetic is stable. The same seven ratings always produce the same total and the same letter grade.

What it does not buy you is repeatability of the ratings themselves. Those come from a model sampling at temperature 1, so run the same idea twice and the individual area ratings can move, and the total moves with them. A score near a grade boundary should be read as "near a boundary", not as a measurement.

And there is no accuracy figure, here or anywhere else in this category, that means anything. Accuracy requires ground truth, and nobody knows whether an idea that was never built would have sold. Any tool quoting an accuracy percentage is quoting a number it cannot have measured. What you can check is the reasoning: open a run, read the problems it cites, and decide whether the verdict follows from them.

What to do with the one that survives

Say a candidate clears the seven questions. You still do not have a business — you have a hypothesis that has not yet been cheaply falsified.

The next step is the smallest test that could plausibly fail. Not a customer interview, where people are polite and enthusiastic and mean none of it, but a real ask: a landing page with a specific promise and a button that costs the visitor something — an email, a waitlist slot, a pre-order. What you are buying is a no that arrives in days instead of months.

Set the stop-loss before you start, in writing, with a number and a date. "If fewer than N people do the thing by this date, I stop." The single most expensive failure mode for a solo founder is not picking a weak idea; it is picking a weak idea and then discovering, eight months later, that no criterion for abandoning it was ever written down.

The short version

Sources

  • The 30-market sweep this article's numbers come fromlib/content/catalog-sweep-facts.ts
  • The seven checks a candidate must survivelib/content/kill-checks.ts
  • The scoring rubric and its fixed weightslib/pipeline/score.ts
  • The published Discovery run, rendered in full/sample-discovery
  • The problem corpus the clusters are built fromlib/content/pain-clusters.ts