Across 23 real runs of our own screening gate, 706 candidate ideas were rejected, and this page is mostly about how those two numbers were arrived at rather than about the numbers themselves. Most of the work in a census like this is exclusion, the exclusions are where a published figure quietly becomes wrong, and ours are written down here so you can re-derive the totals instead of trusting them.
Whose runs these are
Before any proportion below means anything, the reach of the sample. All 23 runs come from 2 accounts, both of them ours. 0 belong to a customer, so nothing here measures how founders' ideas fare in general — it measures how our own gate treated candidates our own pipeline went looking for. The candidates were sourced by the pipeline's own search sweep, not a submission queue, which bounds what the reasons can describe: these are the ways a machine-sourced candidate dies, not the ways a founder's own idea does.
Every verdict was produced under pipeline_triage@v3. The version is printed rather than implied because the proportions belong to it: change the prompt and this page becomes a historical record of a gate that no longer exists.
What had to be thrown away first
| Step | Removes | Runs left | Kills left |
|---|---|---|---|
| After the synthetic predicate The synthetics killed 0 candidates and produced 340 finalists — every candidate survives a stub run, so leaving them in drags every rate toward zero | Stub-mode reports, which kill nothing | 25 | 735 |
| Drop the QA-audit day An audit drives the pipeline deliberately rather than using it; those two killed 18 and 11 candidates, 29 between them | 2 runs from 2026-08-07 | 23 | 706 |
The exclusions, in the order they are applied, with what each one costs. These two steps land exactly on the published totals, and a test fails if they ever stop doing so. The step BEFORE them does not reconcile, and the next section is about that rather than about hiding it.
The synthetic exclusion is the one that matters and it is the one most likely to be skipped. A stub run is the product driven with the model switched off; it kills nothing and every candidate survives it. Leave those in and the rejection rate does not go slightly wrong, it goes toward zero — and the resulting page reads perfectly plausibly, which is the problem.
The second exclusion is a judgement rather than a mechanical one: two runs from a day spent auditing all the modes with real money. An audit drives the pipeline deliberately instead of using it. That is defensible, and it was undocumented for a while, which on a figure whose whole claim is "re-derive it yourself" is the same defect as an unpublished weight.
The step that does not add up, published rather than dropped
Writing this page put two of our own recorded numbers next to each other for the first time, and they do not reconcile. Our record says there were 91 pipeline reports in production and that 68 of them were synthetic — which leaves 23. It also says, in the same file, that removing the synthetics leaves 25 runs. Both cannot be true. For the two to agree the synthetic count would have to be 66.
The history explains it without excusing it: the synthetic figure was written on one day, derived from a non-synthetic count that was corrected the next, and nobody re-derived it. There is also a vocabulary problem underneath — the first number counts reports and the second counts runs, and nothing guarantees those are the same population.
Settling it needs a query against the production database, which is not something this page can run. So the chain above starts where the record is sound, this section states what is not, and a test pins the discrepancy so that a future edit has to say whether it fixed it or moved it. The published totals are unaffected: both steps that remain reconcile, and the headline figures were verified by running the census both ways.
That is also the argument of this page, arriving earlier than intended. Exclusions are where a published number quietly stops being reproducible, and we found ours by writing the recipe down.
The counts, and what is missing from them
The three checks the census counted, of 10 the gate runs. The line under each bar is the honest part: only 2 of the three had their sole-cause share measured, and the other 7 checks were never counted separately at all.
Two things follow from the shape rather than from any single row.
The counted reasons sum to 765 against 706 dead candidates — an overshoot of 59. A census like this does not divide the rejections between causes, because a candidate can fail several checks at once, and every percentage here is therefore a share of rejections that mention a reason rather than a slice of a pie.
7 of the 10 checks have no published count. They fire; nobody tallied them. That is a gap in the measurement and not evidence that they never matter, and a page that listed three reasons without saying so would be implying a completeness it does not have.
What this dataset can and cannot answer
It can answer: what tends to be wrong with candidate ideas as our gate sees them, and in what proportion among the checks we counted.
It cannot answer what a market is like. All 23 runs come from our own accounts rather than from customers, so this is a measurement of an instrument, not of the world. Anyone quoting it as "the top reasons startups fail" would be making a claim about reality out of a claim about a filter.
It also cannot tell you that a rejected candidate was a bad idea. The largest row means we found no documented problem, which is a statement about what a search reached. What rules out the weaker reading — that we simply had nothing to compare against — is that those runs held 943 documented problems between them, about 41 per run, and the number of runs that found none was 0.
If you want to build one of these yourself
The method transfers to any funnel you run, and it is four steps.
- Write down the population before you filter it. Not the final number — the starting one. Every later exclusion has to be subtracted from something stated.
- Name every exclusion and what it costs. A count that quietly drops a category is not reproducible, however carefully the remainder was measured.
- Check that your recipe lands on your published total. Ours is asserted by a test, because a method that stops reproducing is worse than no method: it looks checkable and is not.
- Say which cells you did not measure. The missing sole-cause splits above are the least flattering thing on this page and the most useful, because they tell you exactly how far the conclusions reach.
The short version
- 706 rejections across 23 runs, and the interesting part is the exclusions: synthetic runs kill nothing and would have dragged every rate toward zero.
- The three counted reasons sum to 765 against 706 candidates. Shares overlap; nothing here is a slice of a pie.
- Only 2 checks had a sole-cause split measured and 7 were never counted separately. Both gaps are stated rather than smoothed.
- It measures our gate, not a market. Every run is ours.

