Loading the page.
Method
Every claim in this graph came out of a paper somebody read, carries the sentence it came from, and holds a grade nobody typed. This is how that happens, and what it refuses to do.
A finding becomes a row in five steps. Each one can reject it, and most rejections happen at the third.
A search term is measured before it is trusted.
One query per node, scored against PubMed for title precision before anything is fetched. The difference is not marginal: "creatine"[tiab] returns 43% on-subject titles and creatine[majr] returns 97%, and a term built on a common name rather than a binomial once hid 29% of a species’ literature from the graph.
A claim that cannot point at its own sentence is not a claim.
Every proposed finding carries the exact quoted span it was read from. The quote is stored on the edge and rendered beside it, so a reader can check the sentence rather than the summary. A claim whose quote does not support it on its own is rejected as unsupported, even when it is probably true.
Run independently, in different orders, and compared afterwards.
The same paper is read twice, by two workers who do not see each other’s answers — one traversing the abstract first, one the vocabulary first. Where they disagree about a direction, a population or a number, the claim goes to a human rather than to a tie-break. Agreement is not evidence that a claim is right; it is the only layer that catches a finding read backwards, and a second pass written with the first in view is agreement with oneself.
Mechanical checks first, then an argument against.
Identifiers must resolve, numbers must parse, a confidence interval must contain its own point estimate, and the population must be one the vocabulary knows. What survives is then argued against deliberately — the refuter’s job is to find the reason the claim should not stand, and a claim that survives a real attempt at refutation is worth more than one nobody tried.
No confidence score is an input to anything.
The database computes the grade from the row and rejects a write that tries to set one. Nothing a model reports about its own certainty is used at any layer — a model’s confidence is a fact about the model, not about the world, and treating it as evidence is how a system starts believing itself.
A grade is derived, never asserted. It starts at a ceiling set by the study design, drops one level for every weakness the row admits to, and can rise one level once. The arithmetic is the same in the application and in the database, so the two cannot disagree.
This is the structural refusal, and it is the reason the platform exists. No volume of cell-culture work promotes a claim about people: in vitro evidence never grades above D, whatever else is true of it.
| Design | Best possible grade | In the graph |
|---|---|---|
| Systematic review with synthesismeta_analysis | AWell established | 106 |
| Randomised controlled trialrct | AWell established | 45 |
| Randomised crossovercrossover | BLikely | 27 |
| Prospective observationalcohort | BLikely | 0 |
| Retrospective observationalcase_control | CSuggestive | 0 |
| Uncontrolled human observationcase_series | DPreliminary | 0 |
| Whole-organism non-humananimal_invivo | DPreliminary | 3 |
| Tissue outside the organismex_vivo | DPreliminary | 0 |
| Cell culture or isolated enzymein_vitro | DPreliminary | 0 |
| Asserted from known biochemistrymechanistic | FSpeculative | 0 |
| Searched, nothing foundabsent | FSpeculative | 0 |
Each of these costs one level, and they stack. Four of them on a randomised trial lands it at F, which is the correct answer rather than a bug.
Exactly one thing, at most once: three or more replications, no contradictions, and at least two research groups that are not each other. Replication by the same group is not independent and does not count. The bonus cannot beat the ceiling — a crossover trial with five replications is still a B.
181 findings, graded by that function and by nothing else.
| A | Well established | Replicated human trial evidence, consistent direction | 2313% |
|---|---|---|---|
| B | Likely | Human evidence, limited replication or design limits | 12670% |
| C | Suggestive | Weak human evidence, or strong evidence in a mismatched population | 2313% |
| D | Preliminary | Non-human or in vitro only | 84% |
| F | Speculative | Mechanistic inference, or no evidence found | 11% |
Most of what makes this graph useful is what it will not say. These are rules in the schema rather than intentions, which means a surface cannot break one by being redesigned.
Application code has no grant to set one, and the database rejects a write that tries. If a grade looks wrong, some field on the row is wrong, and the fix is to the field.
A connection assembled from two findings through shared biology is a hypothesis and is labelled one, in words and in the shape of the line. It never renders as something a source reported, because no source reported it.
A search that returned nothing is recorded as having returned nothing, and grades as speculative. That is different from a question nobody asked, and the difference is visible.
Eighteen relations exist and no nineteenth can be added without a written decision. A schema that grows to fit whatever arrived is a schema that cannot be queried consistently.
Third-party graph data is loaded into its own table and cannot become a finding. The only route into the evidence layer is somebody reading the paper.
What can be sold is decided by what the evidence supports. If a product has no evidence behind it, the answer is that it cannot be listed — not that the bar moves.
What line weight, dash and shape mean on every figure, and why none of them is a colour.
Every reference database behind the graph, and the terms each one is used under.
18 compounds carry findings. Each one shows every quote the grade was computed from.