Methodology
Every product we review is evaluated against the same six weighted criteria, the same evidence hierarchy and the same dosing benchmarks. This page documents that framework in full — including its limitations.
1. The Problem We're Correcting For
Two supplements can list nearly identical ingredients and behave very differently in practice. The reasons are usually structural, not mysterious: one uses an ingredient form with established bioavailability while the other uses a cheaper salt; one discloses a clinically studied dose while the other hides its quantities inside a proprietary blend; one is built around ingredients with human trial evidence while the other leans on compounds studied mostly in animals or not at all.
Marketing compresses all of that nuance into the same front-of-pack vocabulary — "clinically studied," "doctor formulated," "science-backed." Our methodology decompresses it. We read the Supplement Facts panel the way a skeptical formulator would, and we score what is actually disclosed, not what is implied.
2. The Six Criteria
Each product receives six subscores from 0–100. The composite NutraBenchmark Score is the weighted average below. Weights reflect our view that what is in the product and at what dose matters more than how pleasant it is to take.
Does the ingredient selection make mechanistic and clinical sense for the product's stated purpose?
Is each active present at a quantity comparable to the doses used in published human trials, in a form with credible absorption data?
How strong is the human research behind the specific ingredients — not the category, and not the brand's own framing of it?
Can a customer determine exactly what they are buying from public information alone?
Cost per effective daily serving — not cost per container.
Adherence determines results, so practicality is scored — lightly.
3. The Evidence Hierarchy
When a brand says an ingredient is "backed by research," we ask: what kind, in whom, at what dose? We grade ingredient evidence on a four-tier scale adapted from standard evidence-based-medicine practice:
| Grade | What it requires | How we treat it |
|---|---|---|
| Strong | Multiple randomized, placebo-controlled human trials — ideally pooled in meta-analysis — showing a consistent effect at a defined dose. | Full evidence credit when the product matches the studied dose and form. |
| Moderate | At least one well-run RCT, or several smaller controlled trials with mostly consistent results. | Substantial credit; flagged where replication is thin. |
| Preliminary | Open-label human studies, small pilots, or strong animal data with plausible human mechanism. | Limited credit. An ingredient at this tier cannot carry a formula's evidence score. |
| Insufficient | In-vitro data, tradition-of-use only, or marketing claims without published human trials. | No evidence credit. Heavily marketed ingredients at this tier are called out in reviews. |
Three rules keep this honest. Dose specificity: evidence is attached to the dose studied — a formula using a quarter of the trial dose does not inherit the trial's grade. Population relevance: results in the studied population (e.g., stressed adults) are weighted above extrapolations from unrelated groups. Outcome relevance: a trial measuring subjective relaxation is not treated as evidence for hormone regulation, and vice versa.
4. Dosing Benchmarks
For the categories we currently cover, these are the daily dose ranges that published human trials have most commonly used. A product scores full dosing credit inside the range, partial credit just below it, and minimal credit for token inclusions. Ranges are periodically revised as new trials publish.
| Ingredient | Commonly studied daily range | What we look for on the label |
|---|---|---|
| Ashwagandha (root extract) | 250–600 mg | Standardized extract (e.g., KSM-66®, % withanolides stated); most stress-outcome trials cluster at 300–600 mg. |
| Rhodiola rosea | 200–600 mg | Standardized extract with % rosavins / % salidroside disclosed; fatigue and stress trials typically 200–400 mg. |
| L-Theanine | 100–400 mg | Acute relaxation studies commonly 200 mg; higher intakes used for daily supplementation. |
| Magnesium (elemental) | 200–350 mg | Elemental amount, not compound weight; bisglycinate/glycinate and citrate preferred over oxide for absorption and tolerability. |
| Vitamin D3 | 1,000–2,000 IU | Maintenance-range supplementation; higher doses are clinical-context decisions, not defaults. |
Two labels can both say "Magnesium 300 mg" and deliver very different things. Magnesium oxide is dense but poorly absorbed; chelated forms such as bisglycinate trade density for absorption and gentler digestion. The same logic applies to botanical extracts: an unstandardized root powder and a standardized extract are not interchangeable, because the active-compound content of raw botanicals varies widely. When a label names a standardized, branded extract, the dose can be compared against research directly; when it doesn't, we discount accordingly.
A proprietary blend discloses a total weight for a group of ingredients without individual amounts. Regulation permits this, but it makes dose verification impossible: a 2,000 mg blend listing six actives could be dominated by the cheapest one. Because our entire framework depends on comparing doses to research, undisclosed doses receive the scoring treatment they earn — the benefit of the doubt goes to the customer, not the formula.
5. Scoring Mechanics
Subscores are set on a 0–100 scale using the rubrics above, then combined using the criterion weights (25 / 25 / 20 / 15 / 10 / 5). The composite is rounded to a whole number and mapped to a verdict band:
Clinically dosed, transparent, evidence-led. Rare by design.
Solid formula with specific, identifiable compromises.
Works as marketed, but under-dosed or under-disclosed where it counts.
Material transparency, dosing or value failures.
Within a comparison article, the same evaluator applies the same rubric to every product in the lineup, and products are always compared against their own category — a cortisol drink is benchmarked against cortisol-support research, not against sleep aids that share two ingredients.
Letter grades that appear in some of our comparison articles (A+ through D) are a condensed presentation of the same underlying assessment, used where a full score breakdown would clutter the format.
6. Limitations & Disclosures
Our scores assess disclosed formulas, doses, pricing and positioning. We do not run identity, potency or contaminant assays on finished products, so we cannot verify that contents match labels — that is what third-party certification programs exist for, and we credit products that carry them.
Even our "Strong" tier describes bodies of research that include null results, small samples and industry funding. A high evidence grade means the balance of published human trials supports an effect at a given dose — not that the effect is guaranteed, large or universal.
Baseline status, medications, genetics and expectations all shape outcomes. Nothing we publish is medical advice, and no score substitutes for a conversation with a qualified clinician — particularly if you are pregnant, nursing or managing a condition.
NutraBenchmark may receive compensation when you purchase through links on our site, and some content is published as clearly labeled advertorial. Our containment measures: the rubric and weights on this page are fixed before any product is scored; criteria are applied identically to every product in a comparison; and the reasoning behind every score — doses, forms, disclosures, price math — is shown in the article so you can audit it yourself and disagree.
Manufacturers reformulate and reprice without notice. Every review carries a last-updated date, and we revise when we become aware of material changes. If you spot one first, we want to know.
The best way to judge a framework is to watch it work — doses checked, blends flagged, price math shown.
Read Our Featured Comparison →