How vendor scores are computed
Every number on the vendor board comes from a published formula with published inputs — no manual overrides, no paid placement, and no term you can't recompute yourself. This page is the formula.
Every published score is machine-verified to reproduce exactly from its published inputs before any deploy goes live. An automated check recomputes all 400+ vendor scores from the same data files this site ships and refuses the release on any mismatch. Each vendor profile shows its own component breakdown under Rank Composition — the terms there add up to the displayed score.
Evidence first; reviews bounded; nothing invented
Since August 28, 2026 every vendor is scored on one evidence scale:
- Verified evidence is the backbone. The 0–1 testing-evidence signal below × 8, plus a review baseline (+1: a flat calibration constant granted to every vendor the evidence model scores at all — it carries no signal of its own and is listed so each receipt sums; pricing plays no role in scoring), a verified-quality term (≤ +2.5: provider-verifiable share × verified purity above 98% × a volume ramp), an evidence kicker (+0.5 above signal 0.30), and a verified small-catalog bridge (≤ +2 for catalogs under 30 products whose documents mostly re-check at the laboratory). Capped at 9.2, with band-consistency floors (a Strong-band vendor never ranks below 7.0; Moderate below 6.2).
- Vendors we have not yet reviewed sit at a neutral 5.0 prior, lifted by an independent lab-data bridge (≤ +3 from third-party lab-test volume and quality, shrinking as our own verification grows), capped at 9.0.
- All outside review sources together — lab-test aggregators, review directories, Trustpilot, Google — form ONE bounded adjustment: at most +0.5 upward, up to −1.5 downward, scaled by review volume. Reviews can warn; they can never buy points.
- Insufficient verifiable evidence → unranked. No number is invented; review signals stay visible.
The testing-evidence signal (0 to 1)
Our review of a vendor's published testing evidence produces the 0–1 signal at the heart of the backbone, from these weighted components:
| Component | Weight | What it measures |
|---|---|---|
| Independent verification surface | 22% | Certificates that can be re-checked at the issuing laboratory — exact report lookups weigh more than generic links. |
| Document coverage | 18% | Distinct testing documents relative to the size of the catalog. |
| Catalog breadth | 14% | Share of the catalog with reviewable testing evidence. |
| Named laboratories | 14% | How many distinct, identified labs appear across the corpus (saturates at 4). |
| Certificate imagery | 12% | Published certificate images and embedded reports (saturates at 120 documents). |
| Certification depth | 8% | Distinct certification types documented (saturates at 4). |
| Reviewed catalog scale | 8% | Overall catalog size (saturates at 60 products). |
| Published report files | 8% | Directly readable report documents (saturates at 80). |
| Account-verified archive | 8% | Evidence confirmed through customer-accessible archives (bounded bonus). |
| Safety panel coverage | up to +0.08 | Average share of standard panels (purity, identity, net content, endotoxin, sterility, heavy metals) reported per product. |
The weighted terms can nominally exceed 1; the signal is clamped to a 0–1 range. A corpus consisting mostly of vendor-hosted documents with no independent verification path takes a proportional reduction (down to ×0.65) — self-published paperwork alone cannot buy a strong signal. A small number of vendors reviewed through customer-account archives are scored on a variant of this signal (coverage 45% · certificate density 35% · breadth 20%, with a ×0.4 score lift ceiling). Trust bands shown on profiles come from this signal plus the verification and purity bonuses: Strong ≥ 0.70 · Moderate ≥ 0.40 · Light > 0.
Outside sources: one adjustment, asymmetric by design
Each source's rating becomes a deviation from a 7.5/10 anchor, scaled by its review volume (log-scaled; a 1-review five-star cannot move the cap), then mixed across whichever sources exist. The class total is clamped to +0.5 / −1.5. Missing sources contribute nothing — never a synthetic neutral. Independent lab-test aggregator data additionally feeds the lab-data bridge above at reduced review-class weight, because lab measurements are evidence, not opinion. Fixed adjustments outside the class: curated program +0.15; overseas bulk suppliers without a reviewed US or curated presence−2.00 (curated vendors exempt).
Score bands
| Band | Score |
|---|---|
| A | 9.0 and above |
| B | 8.0 – 8.99 |
| C | 7.0 – 7.99 |
| D | 6.0 – 6.99 |
| E | below 6.0 |
What never affects a score
- Sponsorship and affiliate relationships. Sponsored placements are labeled, and no paid relationship touches any scoring input — how sponsorship works.
- Business-identity facts. Domain age and published contact details shown on curated profiles are context, not score inputs.
- Review-score shopping. Outside review sources (lab-test aggregators, review directories, Trustpilot, Google) are capped as one bounded class — combined they can add at most +0.5, so accumulating listings cannot buy a ranking.
The evidence-first cutover — August 28, 2026
The previous formula leaned on outside benchmarks as its backbone and blended thin data toward a neutral anchor — which could soften a genuinely poor source rating. After a public dual-display beta, drift gates (median absolute change 0.34 points on the non-redistributed cohort, rank correlation 0.96, zero unexplained band migrations), and a per-vendor migration review, the evidence-first formula above became the live score. Every profile shows its own term-by-term receipt under Rank Composition; vendors whose score moved also show their previous-formula score for a grace period. Fourteen vendors with no verifiable evidence are shown unranked instead of carrying a number — five of them scored under the previous formula, and their profiles still disclose that history.
Also on August 28: the tracked-pricing signal was retired — pricing now plays no role in scoring, in line with peer methodologies. It was replaced by the flat review baseline above, chosen so that every existing score was preserved exactly; the only moves were two vendors gaining the point our own pricing coverage had denied them, and one vendor whose score rested on pricing alone becoming unranked.
The announcement post covers the why and the biggest moves: Evidence-first scoring is live.
- Testing-transparency classifications — how vendors' publishing models are grouped.
- How to verify lab testing yourself — the checks behind our verification surface.
- Any vendor profile → Rank Composition — that vendor's own terms, adding up to its displayed score.