Skip to content
WhittleOSWhittleOS
← All guidesProblems5 min read

Where to find real customer problems, and what makes one quotable

The stock answer names a website. The useful answer names a property: can you open the thing the claim rests on? Here is what one published run actually read, the five tiers we grade a source by, and why 43% of its problems carry no link at all.

By Boris Binyaminov · ·

Hosts one run reached
3, and none of them is Reddit
Problems it recorded
56
With an openable link
32 — 57%
Provenance tiers
Five, and one means not verified

The stock answer is "go read Reddit", and it names a destination when the question is about a property. What makes a problem usable is not where you found it but whether you can hand someone the thing it rests on. On the run we publish in full, 32 of 56 recorded problems carry a link you can open. The other 24 are marked, not deleted, and that distinction is the whole method.

The property, not the place

A problem you can act on has three parts: somebody described it, in their own words, somewhere you can point at. Lose the third and you have a claim; lose the second and you have a summary of a claim; lose the first and you have your own hypothesis wearing a citation.

Most lists of places to find customer problems skip straight to destinations. Destinations change — a platform closes its API, a forum goes private, a review site adds a login wall — and a method built on a list of sites expires with the list. A method built on the property does not.

The five tiers, in the product's own words

This vocabulary is shared by every surface here, and it exists because it once was not: five hand-maintained copies described the same tier three different ways, and one page rendered two of them at once.

1
from a real page
Our own fetcher read the whole page.
2
from search
A search result's summary. Real, but we never opened the page.
3
you provided this
You handed us the URL.
4
you pasted this
You pasted the text.
5
not verified
Inferred with nothing behind it. Never reads as confirmed.

Strongest first, rendered from the one map the whole site imports. Notice the last one: it is written so it can never be misread as confirmed, because that is the tier a confident tool would quietly promote.

The rule that makes the ladder mean anything is that a claim cannot be promoted up it. A search snippet is real evidence that a page exists and says something — it is not evidence of what the whole page says, so anything grounded in one stays capped at that tier no matter how convincing the summary reads. Tools that collapse these tiers are not being sloppy; the collapse is the product.

What a source has to be to be reachable at all

Three mechanisms, and each has a limit that decides what you can honestly quote from it:

MechanismWhat it doesIts limit
searchA search API returns results for a query we generated.A snippet is not a page. Anything grounded in one is capped at the weaker tier.
scraperOur own reader fetches a listing page and extracts the entries.Gated by that site's robots rules per fetch, and skipped when they say no.
apiA documented public API, called directly — no scraping, no key.Only exists where a site offers one. Most do not.

The limits are the useful column. A source that cannot be reached by any of the three is not a bad source — it is one you will have to read yourself, and know that you did.

Two consequences worth stating plainly, because they cut against us.

Robots rules decide reachability, and we obey them. A site that says no is skipped, and whatever it knew is absent from the run. That is a real hole in coverage and the honest response is to say so rather than to route around it.

The best evidence is often the least reachable. Long complaint threads behind logins, support forums that render client-side, private communities — the places people are most candid are the places a fetcher cannot go. Anyone claiming complete coverage of where customers complain is describing a product that does not exist.

What one run actually read

41
sources consulted
24
pages actually opened and read
57%

of recorded problems carry an openable link

From the run published in full, not a typical run. The gap between sources consulted and pages read is the snippet tier doing its job.

The hosts that run reached were saashub.com, trustmrr.com, remotive.com. That is not a recommendation list, and treating it as one would be the exact move this page argues against — it is where one run's generated queries happened to land, for one niche, on one day. A different niche reaches different places.

Reddit is not among them, and we do not source it. If you came here expecting a Reddit-mining technique, that is the honest answer, and the Reddit guide explains what the method can and cannot tell you when you do it by hand.

Doing it yourself

  • Keep the URL, not the summary. The moment you paraphrase into a doc without the link, you have converted evidence into memory.
  • Record the date. A complaint from 2019 about a product that shipped the fix in 2021 is not a problem, and nothing in the text will tell you.
  • Mark what you inferred. Not everything will have a source. The failure is not having unsourced hunches; it is letting them look like the sourced ones three weeks later.
  • Count what you could not reach. A search that returned nothing is a finding. A site that refused you is a hole in your coverage, and knowing its shape is worth more than pretending it closed.

The short version

  • The question is not where. It is whether you can open the thing your claim rests on.
  • Five tiers, and a claim never gets promoted up them — a snippet stays a snippet however convincing it reads.
  • Three ways to reach a source, each with a limit; the most candid places are usually the least reachable, and no tool covers them all.
  • On the published run, 32 of 56 problems carry a link. The other 24 are labelled unverified rather than quietly promoted.