Skip to content
WhittleOSWhittleOS
← All guidesValidation tools5 min read

How to score a business idea, with the weights published

Every guide on scoring ideas tells you to pick criteria and assign importance. None of them says what they picked. Here is the rubric this product scores with — all 7 weights, the arithmetic, and the part a score cannot settle.

By Boris Binyaminov · ·

Weighted categories
7, plus one at zero
Heaviest single area
20% of the total
Heaviest vs lightest
4× the leverage

A scoring model is only as honest as its weights. Here are ours: 7 categories that carry weight, one that is scored and carries none, and the arithmetic that turns them into a number out of a hundred. Read them and you can work out where a point of effort actually goes — which is the only thing a score is good for.

Why the weights are the whole model

Search this question and you get a page of results explaining how to build a weighted scorecard: choose criteria, assign importance, multiply, sum. All correct, and all of it stops exactly where the decision starts. Which criteria. What importance.

That gap is not an accident of writing. Publishing weights commits you. It lets a reader say you have priced trust twice as high as margin and I think you are wrong, and it lets them notice when your scores do not follow your own stated rules. A model with unpublished weights cannot be argued with, which is comfortable and useless.

Sales & Trust Fit20%
Can the buyer be reached with async marketing? How much trust is required to convert?
Distribution Fit20%
Is there a repeatable, low-cost channel the founder can operate (SEO, IH, X, niche newsletters, etc.)?
Economics / Pricing / Margin20%
Plausible price point, margin after AI costs, gross margin sustainability for a solo operator.
Pain + Demand Evidence15%
Strength of the evidence from the Pain Box + demand-proof signals. Public sources or pasted evidence beats assumption.
Operations / Support Load10%
Support burden, onboarding friction, manual work per user. Solo founders cannot scale white-glove ops.
Risk / Platform / Regulation10%
Platform-dependency risk, AI dependency, legal/regulatory exposure, ToS risk.
Cash-Flow Sustainability + Transferability5%
Months to break-even on hosting + AI costs. Transferable asset value (can the asset be sold?).
Founder Understanding Fitnot weighted
Does the founder understand the buyer's world well enough to write copy, build the product, and answer support? (Not weighted into the final score; see KF-08.)

The last row is the one to notice. Bars are relative to the heaviest, and every weight and description is imported from the module the product scores with — so this chart cannot describe a rubric other than the one being applied.

Three things fall out of that picture.

The top is flat on purpose. Three areas share the heaviest weight rather than one dominating. Whether the buyer can be reached, whether a channel exists, whether the money works — no ordering between them survived contact with real ideas, so none was invented.

Evidence is weighted below all three. That reads wrong until you notice what it prevents: an idea with excellent evidence of a problem nobody can be reached about or paid for is still a bad idea, and weighting evidence at the top would let a well-documented dead end outscore a reachable one.

One category is scored and weighted at zero. Founder Understanding Fit is rated, shown, and contributes nothing to the total. It belongs on the page because it changes what YOU should do about a verdict; it stays out of the arithmetic because an idea does not become a better business by being handed to someone who understands it.

What a point is worth, which is the useful part

Take an idea rated 7 in every area except one. Distribution Fit comes back at 3 — a decent product nobody can be reached about, which is the most common shape there is. The total is 62, grade C.

Now spend the same three rating points twice, and watch them buy different things:

Three points into Distribution Fit

6270

+8 points, and grade C becomes B. The hard fix, and the one that moves the verdict.

Three points into Cash-Flow Sustainability + Transferability

6264

+2 points and the same grade. Real work, and it changes nothing about the decision.

Both totals come from computeFinalScore run on the same scorecard with one rating changed, so the comparison is the product's arithmetic rather than ours restated.

The heaviest area carries 4 times the leverage of the lightest. That ratio is the entire practical value of publishing a rubric: without it you cannot tell the expensive fix that moves a verdict from the satisfying one that does not, and effort goes where it feels productive rather than where it counts.

The letter, and what it is for

A
80–100
B
65–79
C
50–64
D
under 50

Read from the same array the product's grade ring renders from. Until 2026-08-07 these bands existed only in code and a design prototype, so a reader looking at four Ds had no way to know whether D meant "bad" or "normal here".

A letter is a compression of a number that is already a compression of seven judgements, and it earns its place only because it is the fastest way to see that two ideas are not in the same class. If you are choosing between ideas, the letters are enough. If you are deciding what to fix, they are the wrong instrument and the per-area ratings are the right one.

What the number cannot settle

Here is the part a published rubric obliges you to say.

The arithmetic is fixed. The ratings are not. The weighting, the sum and the band are pure code and will do the same thing every time. The zero-to-ten ratings underneath are model judgements sampled at temperature one, so the same idea run twice can score differently, and a total is a reading rather than a measurement. Anyone claiming otherwise is describing their arithmetic and hoping you hear it as their whole system.

Uncertainty is not a middling score. A category the model cannot judge does not average out to five — it is marked low-confidence, and enough of them together push the verdict to HOLD rather than letting a confident-looking total emerge from seven shrugs.

A total hides the thing that kills you. Our own worked example is a C whose real content is one specific problem and six areas that are fine. That is why the run's recommended first move names the weakest area rather than the score: an average is exactly the wrong summary of a set where one entry is fatal.

The short version

  • Weights are the model. A scorecard that will not publish them cannot be argued with, which is the point of not publishing them.
  • Ours: 7 weighted categories, top three tied at 20%, evidence deliberately below all three, and Founder Understanding Fit scored at zero weight.
  • The heaviest area is worth 4× the lightest, which is how you tell an expensive fix that changes the verdict from a satisfying one that does not.
  • The arithmetic repeats; the ratings sample. Treat the total as a reading, and read the areas.