Evidence methodology
Two kinds of evidence, weighed differently, because they fail differently.
The store's own signals
A count becomes an observation only when it is outside this store's own range. Observations carry a weight: strong for a signal that on its own describes a broken revenue path (a webhook failing repeatedly, a gateway enabled but unavailable, a scheduler tens of jobs behind), supporting for one that corroborates a story without telling it (note reason codes, a slower transition, a status mix).
What confidence counts is how many DIFFERENT families of signal point the same way. Three observations of one kind is one direction repeated; orders piling up in pending AND a webhook failing is two directions agreeing. One family is a coincidence with a narrative attached, and produces an abstention.
Independence, not row count
Public evidence names its origin group — the publisher, not the post. A GitHub issue and the WordPress.org thread quoting it are one witness; the second is recorded as derived and adds nothing to the independent count. Two threads on the same forum count once. Every answer shows how many distinct origins agreed, not how many rows matched.
Freshness
An observation's weight decays with age — full weight inside 45 days, decaying to a floor by 240 days, never to zero. A verified fixed version does not decay: a release note does not become less true with age. On the private side the rule is harsher, because a store changes by the hour: a snapshot over an hour old is "aging", over six hours it is "stale" and no diagnosis is attempted at all.
Attribution, and refusing to attribute
When exactly one component changed before the symptom, it is named — and only then can public evidence agree or disagree, because there is something to look up. When two changed, neither is named: picking the more plausible one is precisely the fluent guess this service exists not to make, so the ambiguity is reported instead.
Confidence is a floor, never an average
Two signal families and at least one strong signal reach medium. Only independent public agreement — two or more origins on a named component — reaches high, because one store's data cannot tell a bug apart from a local misconfiguration. Aging data lowers it. Nothing strong and nothing public is not "low confidence", it is abstained: naming a culprit there would put a wrong answer in front of someone during an outage.
What is stored
On the public side, a derived fact and a short quotation with the source URL — never full page text or a forum author's identity. On the private side, aggregate counts and health states only; see Privacy.