How we grade the evidence

The two-axis rubric every breakdown on this site is assigned under: how strong the research is, and what that research actually found.

The Compound Codex · Editorial Standards

Every breakdown carries two marks: how strong the research is, and what that research actually found. This page defines exactly what earns each one, so you can check our work — and disagree with it.

Version 2.3 Applies to all breakdowns Last revised August 2026

The system

Two marks, not one

A single score can't describe research honestly. "Nobody has studied this" and "four large trials found nothing" are opposite situations, and a one-dimensional scale collapses them into the same low number. So we grade on two independent axes.

Evidence strength — four dots — measures how good the research base is. It rises with the quantity and quality of the studies, regardless of what those studies found. A well-refuted claim scores high on strength.

Verdict — one word — states what that research concluded: supported, mixed, or not supported.

Compound name

For the specific claim being examined

Strong evidence Not supported

Header block — appears at the top of every breakdown

A third element joins them when it applies. Some compounds carry a documented safety concern that has nothing to do with whether they work, and the two axes above cannot express it — so a safety flag sits alongside them:

Compound name

For the specific claim being examined

Preliminary evidence Untested in humans ⚑ Restricted by a regulator

Header block with a safety flag — shown only when one applies

Read together, the two marks say something a single score cannot: we know a great deal about this, and what we know is that it doesn't work. That combination is the most useful sentence this publication can write, and it needs both axes to exist.

Axis one

Evidence strength

Assessed against the specific claim, not the compound in general. The same compound can hold different strength grades for different outcomes.

Independent group
A research team sharing no senior authors, no institutional affiliation, and no common funding source with the other trials being counted. Two papers from the same lab count once.
Qualifying trial
A randomized controlled trial in humans, with a control arm, reporting the outcome under examination. Registered, with results posted or published.
Participant count
Total enrolled across all qualifying trials, summed. Per-trial minimums are stated separately where they apply.

Preliminary

grade-preliminary

Interesting in a dish or a mouse. Nobody has properly tested it in people.

Requires

  • Zero completed qualifying trials

Typical evidence base

  • Cell culture, tissue, or animal models
  • Case reports, uncontrolled observations, or anecdote
  • Human trials registered or underway but not yet reporting

Most research peptides sit here. A rodent result is a reason to run a human trial, not a reason to believe the finding will hold in humans.

Emerging

grade-emerging

Tested in people once or twice, in a small way. Real data, not yet confirmed.

Requires

  • At least 1 qualifying trial

Capped here when any of these apply

  • Fewer than 2 independent groups have reported on the claim
  • Total participants across qualifying trials is under 200
  • All qualifying trials are open-label, lack a placebo arm, or are manufacturer-funded without independent replication

The most commonly misread grade. A single promising trial gets written up as though the question is settled. It is not.

Moderate

grade-moderate

Several decent human trials, run by different people, tested properly.

Requires all of

  • At least 2 qualifying trials, randomized and placebo-controlled
  • At least 2 independent groups reporting
  • 200 or more total participants

Capped here when any of these apply

  • Total participants remain under 1,000
  • No systematic review or meta-analysis has been published
  • Findings hold only in one narrow population
  • Trial durations are too short to speak to sustained use
  • The verdict is Mixed — unresolved disagreement caps strength here

A realistic ceiling for most supplements that genuinely do something.

Strong

grade-strong

Repeatedly tested at scale by people with no stake in the answer.

Requires all of

  • At least 3 qualifying trials, randomized and placebo-controlled
  • Scale, by either route: at least 2 trials with 100 or more participants each — or at least 10 qualifying trials pooled by meta-analysis
  • At least 3 independent groups — or 2 groups plus a published systematic review
  • 1,000 or more total participants
  • A published systematic review or meta-analysis reaching a clear conclusion
  • Replication across 2 or more distinct populations
  • Verdict is Supported or Not supported — never Mixed

Two routes to scale exist because two research traditions do. Drug literature earns confidence through a handful of large trials; supplement literature earns it through many small independent replications, pooled. Both are legitimate, and a scale that recognised only the first would mis-grade some of the best-evidenced compounds we cover.

Very few compounds will reach this grade in either direction. That is the point of having it. If everything on a site scores well, the scale is decoration.

Axis two

Verdict

What the qualifying trials concluded. Assigned only where strength is Emerging or above — at Preliminary there is no human evidence to draw a verdict from, and the verdict reads Untested in humans.

Supported

The effect showed up

The weight of qualifying evidence found the effect, and the pooled direction is positive.

Meta-analysis finds a statistically significant effect — or, absent one, trials representing the majority of pooled participants report the effect.
Mixed

The trials disagree

Qualifying trials point in conflicting directions and the conflict cannot be resolved by dose, population, duration, or formulation.

Caps evidence strength at Moderate, however large the literature.
Not supported

It didn't show up

The weight of qualifying evidence found no effect. This is a finding, not an absence of one — and at higher strength grades it is a firm conclusion.

Meta-analysis finds no statistically significant effect — or, absent one, trials representing the majority of pooled participants report no effect.

Scope

What the scale does not grade

Both marks describe evidence about outcomes in humans. The strength axis is built from counts of human trials, participants, and independent replications. Point it at anything else and it returns a number that looks meaningful and is not.

So two categories of claim are reported in our breakdowns but never carry marks:

  • Mechanism claims — statements about how a compound works: receptor binding, enzyme activation, a biochemical pathway. These are settled in laboratories, not trials, and have no participants to count.
  • Preclinical claims — findings in cell culture or animal models, however extensive. A large animal literature raises a question worth answering; it does not answer it.

Where such a claim matters to a compound's story we describe it plainly and say what the laboratory evidence shows — refuted in vitro, demonstrated in rodents, disputed — without dots and without a verdict pill.

The distinction that makes this worth a rule: a refuted mechanism is not a refuted compound. A substance can work by a route nobody has identified yet, and plenty do. What collapses when a mechanism fails is the reason to expect a clinical benefit — which means the human trials have to carry the whole argument alone, and are usually asked to do so with far less enthusiasm behind them.

The reverse error is more common and more expensive: treating an elegant mechanism as though it were a result.

When trials disagree

Conflicting results are the normal state of this literature, not an exception. We resolve them in a fixed order, and we never settle a disagreement by counting how many papers landed on each side — a large, well-run trial and a tiny flawed one are not one vote each.

  1. A meta-analysis decides. If a published systematic review or meta-analysis addresses the claim, its conclusion sets the verdict. If several exist and disagree, we prefer the most recent that includes the largest trial set.
  2. Absent one, weight by size and quality. Larger trials outweigh smaller. Placebo-controlled outweighs open-label. Independently funded outweighs manufacturer-funded. Pre-registered outweighs not.
  3. Look for the variable that explains the split. If the disagreement tracks dose, formulation, population, or duration, we say so explicitly and grade the narrower claim that the evidence actually supports — not the broad one it doesn't.
  4. If it still won't resolve, the verdict is Mixed and strength is capped at Moderate. We say plainly that the question is open rather than picking the answer we find more interesting.

A failed replication never lowers the strength grade. It raises it — more research exists now than before — while pushing the verdict toward Mixed or Not supported. Strength describes how much we know. Verdict describes what we know.

Confidence is not magnitude

This distinction causes more confusion than anything else in health reporting, so we separate the two explicitly in every breakdown.

The marks tell you

How good the research is, and which way it points. Driven by trial count, trial quality, independent replication, and pooled direction.

The marks do not tell you

How large the effect is, or whether it's large enough to matter to you. That belongs in the body of the breakdown, in real units.

A worked example: a compound may carry strong, replicated evidence that it improves a sleep-onset measure — and the actual improvement may be four minutes. Solid evidence, negligible benefit. Both facts belong in the same sentence, and we write them that way.

Every breakdown states the effect in the units the trials used, alongside the smallest change a person would plausibly notice. Percentages without baselines, and relative risks without absolute ones, are not reporting — they're decoration.

Third element

The safety flag

Both marks describe evidence about whether something works. Neither says anything about whether it is safe, and for some compounds that is the more important question by a wide margin.

This produced a real failure. A compound whose development was terminated over animal carcinogenicity, and which one national regulator prohibits from sale outright, still displayed as Preliminary · Untested in humans — marks a reader could reasonably take to mean nobody knows, proceed carefully. The marks were accurate. The impression was wrong.

So a safety flag may appear beside the two marks. It is not a third grade. It records what has been documented and who documented it, nothing more.

⚑ Not characterised

No human safety data exists. Applied whenever no completed human trial has assessed safety, regardless of how much animal work exists.

Typical trigger: a regulator states it lacks information to know whether the substance would cause harm in humans

⚑ Identified in trials

A specific harm reached statistical significance in a controlled human trial.

Typical trigger: an adverse event significantly more common in the treatment arm of a published RCT

⚑ Documented in case reports

Specific harms are described in the published case-report literature, outside of any trial.

Typical trigger: peer-reviewed case reports of a defined injury following use

⚑ Catalogued by a health authority

A national or international body has formally assessed the harm and entered it in a reference database.

Typical trigger: an NIH LiverTox likelihood score, or an equivalent formal assessment

⚑ Restricted by a regulator

A regulator or sponsor has acted on safety grounds — scheduling, a compounding prohibition, or a trial halted early for safety.

Typical trigger: a scheduling decision, a bulk-substance category placement, or an early trial termination

Where more than one applies, the breakdown shows the one carrying the greatest institutional weight, and the safety section describes the rest. Where none applies, no flag is shown — and the absence of a flag is not a claim of safety, only the absence of a documented concern meeting these criteria.

Three things the flag deliberately is not.

It is not a severity ranking. A regulator restriction is not automatically worse than a case report; the two describe different kinds of body acting, not different degrees of danger. Read the list as a record of provenance, not a thermometer.

It is not a recommendation. We do not tell readers what to take, and a flag is not a warning against use any more than its absence is permission.

It is not a grade. It carries no dots, does not combine with the verdict, and cannot raise or lower either mark.

Safety is reported, never graded

The flag points at something; this section is where it gets explained. Grading safety would imply a judgment about whether a person should take something, which is not a judgment this publication makes. Every breakdown carries a required Safety and Unknowns section covering five points.

  • Adverse events reported in the trials we cite, with the rates they occurred at
  • Longest human exposure studied — the duration of the longest qualifying trial, stated plainly
  • Known interactions with medications or conditions where the literature documents them
  • Regulatory status — approved, unapproved, sold as a research chemical, or banned in sport
  • What has not been characterized — stated explicitly rather than left as white space

The standing rule on unstudied compounds: where no human trials exist, the honest sentence is "the safety profile in humans has not been characterized." It is never "no side effects have been reported." Absence of reported harm in rodents is not evidence of human safety, and writing it that way would mislead the reader on the single point where the stakes are highest.

What counts as a source

Marks are assigned from primary literature. We read the studies themselves, not coverage of them.

  • We use: peer-reviewed papers indexed on PubMed and PMC, registered trial records with posted results, systematic reviews and meta-analyses, and regulatory assessments.
  • We use with caveats, always labelled: preprints that have not cleared peer review, conference abstracts, and trials with manufacturer or undisclosed funding. These can support a verdict but never lift strength above Emerging on their own.
  • We do not use: vendor product pages, press releases, podcast claims, influencer protocols, or forum reports as evidence for any mark.

Every factual claim links to the paper it came from. If a claim carries no citation, treat it as our opinion and weigh it accordingly.

Marks change

Both marks are a snapshot of the literature on a date, not a permanent verdict. Every breakdown carries the date its marks were last reviewed.

TriggerWhat happens
A significant new trial reportsBreakdown is revisited and both marks re-assessed within 30 days
A replication failsStrength rises, verdict moves toward Mixed or Not supported, change noted at the top of the page
A cited paper is retractedCitation is struck, marks re-derived without it, correction posted
Routine reviewEvery breakdown re-checked at least annually

Changes are logged publicly. We would rather be visibly wrong and corrected than quietly wrong.

What we don't publish

These limits are permanent and are not commercial decisions. They keep this a publication about research rather than a guide to self-experimentation.

  • No dosing protocols for human use. We report the doses used in published trials, in the course of describing those trials. We do not tell readers what to take.
  • No vendor affiliate links. We take no commission on the sale of any compound we cover, ever. Affiliate revenue and honest grading cannot coexist.
  • No medical advice. Nothing here substitutes for a clinician who knows your history.
  • No sponsored breakdowns. No company can pay for coverage, for a mark, or for a revision to one.

Corrections

When we get something wrong, the correction appears at the top of the affected breakdown with the date and a plain description of what changed. We don't silently edit. If you spot an error — a misread result, a bad citation, a mark you think is indefensible — write to corrections@compoundcodex.org and we will look at it. If you want the argument on the public record instead, file it as a formal challenge.

Revisions to this rubric

The rubric is versioned and its changes are logged here. A publication that grades other people's evidence should be willing to show its own working.

VersionDateChange
2.3 Aug 2026 Added the safety flag. Both marks describe whether something works, so a compound with a serious documented safety concern could display marks that read as merely unproven. Surfaced by Cardarine, where a terminated development programme and an outright national prohibition sat behind a one-dot Preliminary. The flag records what was documented and by whom; it is not a grade, not a severity ranking, and not a recommendation.
2.2 Aug 2026 Placed mechanism and preclinical claims outside the scale. The strength axis counts human trials and participants; applied to a biochemistry finding it returned a grade with nothing behind it. Surfaced while grading resveratrol, where the refuted SIRT1 mechanism was briefly given a clinical grade it could not support.
2.1 Aug 2026 Added a second route to the scale requirement at Strong. The original wording demanded at least two trials of 100+ participants, which mis-graded compounds whose evidence rests on many small independent replications rather than a few large trials. Surfaced by applying the rubric to creatine monohydrate, where the strict reading returned Moderate for one of the deepest evidence bases in the field.
2.0 Aug 2026 Split grading into two axes — evidence strength and verdict — so that well-refuted claims are no longer scored identically to unstudied ones. Added numeric thresholds, the conflict-resolution rule, and the required Safety and Unknowns section.
1.0 Aug 2026 Initial four-point evidence scale.

The Compound Codex summarizes published research for informational purposes only. Nothing here is medical advice, and nothing here endorses the use of any compound discussed. Always consult a qualified clinician before starting any peptide, nootropic, or supplement regimen.

© 2026 The Compound Codex · Every claim cited, none of them prescriptive.