How the Vendor Benchmark Works
The current PeptideBenchmark scoring methodology: source-backed and PB-evidence paths, confidence weighting, testing trust, review signals, pricing coverage, and scoring limits.
What the PB score is—and is not
The PB score is a directional public-evidence score on a 10-point scale. It helps readers compare how much source-backed information PeptideBenchmark can currently verify about a vendor.
It is not a laboratory result, pharmacy credential, regulatory approval, or promise that the next vial or order will perform as expected.
The score currently considers different combinations of:
- Finnrick testing records;
- Peptide Critic ratings and review depth;
- Trustpilot rating and review depth;
- Pep Review Pro rating and review depth;
- PeptideBenchmark’s structured testing-transparency review;
- whether public pricing coverage exists;
- whether PB has completed a curated review; and
- whether an otherwise anonymous overseas supplier lacks a reviewed U.S. or curated presence.
Not every vendor has every input. The model therefore uses two scoring paths instead of pretending all vendors arrive with the same evidence.
One board, two scoring paths
Source-backed
Used when Finnrick or Peptide Critic covers the vendor.
A normalized outside-source rating, weighted by the depth of supporting tests and reviews.
PB evidence
Used when PB tracks the vendor but neither outside source does.
Testing transparency, provider-verifiable coverage, independent reviews, and visible pricing.
Path 1: source-backed vendors
Step 1: normalize the outside ratings
Finnrick now headlines percentage-based results. PB calculates a test-count-weighted average of Finnrick’s published per-peptide percentages, then divides that percentage by 10. When no rated-product breakdown is available, Finnrick’s public vendor composite is used as the fallback before the same conversion.
Peptide Critic’s 5-point rating is doubled. Both sources therefore enter the benchmark on the same 10-point scale.
- If both sources cover the vendor, the benchmark score is their average.
- If only one covers the vendor, that normalized source score remains the benchmark score.
This raw benchmark is shown separately on vendor profiles. It is not necessarily the number that orders the board.
Step 2: measure confidence
A perfect score supported by one thin source should not automatically outrank a strong score supported by many tests and reviews.
The confidence model therefore considers:
- Finnrick test count, with diminishing returns;
- Peptide Critic review count, also with diminishing returns;
- whether both sources cover the vendor;
- Peptide Critic verification and batch-tracking indicators; and
- a minimum confidence floor based on the type of source coverage.
Consensus vendors receive more starting confidence than one-source entries. The counts use logarithmic scaling, so the first meaningful body of evidence matters more than the difference between, for example, the 200th and 201st record.
Step 3: pull thin evidence toward a neutral anchor
The model starts from a neutral reference point:
- 7.2 for a Finnrick + Peptide Critic consensus vendor;
- 7.0 for a Peptide Critic-only vendor; and
- 6.9 for a Finnrick-only vendor.
The raw benchmark score moves away from that anchor only as confidence increases:
base rank = neutral anchor
+ (benchmark score − neutral anchor) × confidence
This is why a lightly supported 10.0 can rank below a broadly supported 9.0.
Step 4: apply bounded, labeled adjustments
The base rank can then receive:
PB testing-transparency lift
EvidenceRewards provider verification, archive coverage, named labs, certificate depth, and related testing evidence.
Trustpilot
+0.4 / +0.6Maximum varies by source count and is scaled by rating and review depth.
Pep Review Pro
Up to +0.2Scaled by the reported rating and the depth of reviews behind it.
Curated review
+0.15Applied after PB completes its deeper curated-review workflow.
Overseas-supplier adjustment
−2.0Applies to non-curated suppliers without a reviewed U.S. presence.
The final result is capped between 0 and 10.
The curated adjustment is intentionally small. It recognizes that PB has completed deeper identity, testing, pricing, and policy review; it is not an affiliate or sponsorship bonus. Curated vendors are exempt from the anonymous-overseas-supplier adjustment because they have already entered the reviewed-vendor workflow.
Path 2: PB-evidence vendors
Some vendors have meaningful public testing and pricing evidence but no Finnrick or Peptide Critic score. Giving them a fake outside-source rating would be misleading, so PB uses a separate evidence formula.
The PB-evidence score currently combines:
Testing-transparency evidence
PB measures the visible testing surface, including:
- public laboratory or provider lookup links;
- Janoshik public and verification links;
- certificate coverage across the catalog;
- named-lab and certification breadth;
- vendor-hosted documents versus provider-verifiable records;
- authenticated certificate archives when access has been reviewed;
- available safety-panel coverage; and
- the share of testing documents that can be corroborated at the provider.
Curated vendors use an 8× testing-signal multiplier; other PB-discovered vendors use 6× because curated coverage has gone through a deeper review workflow.
Trustpilot contribution
Trustpilot can contribute up to:
- 4.0 points for curated PB-evidence vendors; or
- 3.0 points for other PB-evidence vendors.
The contribution is scaled by both the rating and review-count confidence. Twenty reviews currently provide full review-count credit for this path; fewer reviews receive proportionally less.
Public pricing presence
A PB-evidence vendor receives one point when at least one public pricing row exists.
This is deliberately binary. A larger catalog does not earn more score, and a cheaper price does not earn more score. Price competition belongs on the Pricing board and in the Stack Builder, not inside a trust score.
Provider-verifiable coverage quality
A quality term can contribute up to 2.5 points when:
- a meaningful share of testing documents can be verified through the laboratory or provider;
- average verified purity is at least 98%; and
- multiple documents are available.
Full count credit begins at five provider-verifiable documents. This prevents a vendor with one perfect certificate from receiving the same quality credit as a consistently verified catalog.
Remaining guardrails
- Testing signals above 0.3 receive a small 0.5 depth bonus.
- Curated review can add 0.15.
- Pep Review Pro can add up to 0.2.
- Trustpilot is already included in the PB-evidence base and is not counted a second time.
- The completed score remains capped at 9.2.
Benchmark score, PB score, trust band, and curated assessment are different
These labels answer different questions:
Benchmark score
What Finnrick and/or Peptide Critic report after scale normalization.
PB score / rank score
The directional score after confidence weighting and disclosed adjustments.
Testing trust band
How strong the testing-transparency evidence reviewed by PB appears.
PB Curated Assessment
PB’s separate conclusion about the vendor’s broader trust and transparency posture.
Testing trust bands currently resolve as:
A high PB score and a Strong testing band can coexist, but they are not synonyms. The Trust System exposes testing evidence separately so readers can inspect the underlying records instead of relying on a badge.
The PB Curated Assessment is also kept separate from the PB score. It is an editorial rubric, not another hidden scoring multiplier.
What affiliate relationships and sponsorship do not change
PeptideBenchmark may earn commissions when readers use disclosed affiliate links or coupon codes. Those relationships:
- do not increase the PB score;
- do not change the testing trust band;
- do not change the PB Curated Assessment;
- do not buy a curated listing; and
- do not change the calculated ranking score.
Sponsored visibility is labeled and walled off under our sponsorship policy. An active sponsored placement may be pinned in a clearly labeled display slot, but it does not alter the vendor’s score or the underlying score-sorted order. A vendor can be curated without being an affiliate, and an affiliate relationship alone does not qualify a vendor for curation.
Coupon and quantity-discount data can change the effective price shown in Pricing or Stack Builder. It does not make the vendor more trustworthy.
The separation is intentional: commercial relationships can affect which outbound link or discount a reader sees. They cannot alter the score, testing band, or editorial assessment.
Why the same vendor can show several different numbers
Consider a hypothetical consensus vendor:
- Finnrick’s published percentages normalize to 8.8/10.
- Peptide Critic reports 4.6/5, normalized to 9.2/10.
- Its benchmark score is therefore 9.0.
- Evidence depth determines how far the base rank moves from the 7.2 neutral anchor toward 9.0.
- Reviewed testing transparency, Trustpilot, Pep Review Pro, and the small curated adjustment may then move the rank within the 10-point cap.
Now consider a vendor absent from both source directories:
- It has no legitimate Finnrick/Peptide Critic benchmark score.
- PB instead evaluates public testing evidence, provider-verifiable coverage, review signals, and whether pricing is observable.
- The result is labeled PB evidence and capped at 9.2.
The two vendors can appear on the same board, but their profiles disclose which path produced the number.
Freshness, fallbacks, and corrections
The vendor benchmark workflow is scheduled daily. Inputs do not all refresh on the same cadence:
- Finnrick is retrieved through its public vendor source;
- Peptide Critic may require a last-known-good snapshot when its live directory blocks hosted infrastructure;
- Pep Review Pro is synchronized into the benchmark workflow;
- Trustpilot is refreshed through its own scheduled process;
- testing evidence and gated archives have separate review workflows; and
- vendor pricing pipelines run independently of the score sync.
The generated benchmark records source timestamps and status. A stale upstream source is not silently presented as a fresh observation, and the pipeline can fail rather than accept an over-age fallback.
Vendor aliases and domain changes can create mismatches. We maintain explicit identity mappings and audit generated profiles for duplicate records, score-component drift, placeholder text, and public leakage of internal pipeline terminology.
What the board still does not prove
Even a 10.0 source-backed score or a 9.2 PB-evidence score does not prove:
- the identity or concentration of the vial currently for sale;
- sterility, endotoxin, or heavy-metals status unless a relevant current report supports it;
- that a certificate belongs to the batch a buyer receives;
- fulfillment quality on the next order;
- long-term business durability or complete ownership transparency;
- legality, approval, or medical suitability for human use; or
- that reviews are representative, independent, or immune from manipulation.
The board is a comparison and due-diligence tool—not a warranty. Read What a High Vendor Score Does Not Prove before treating any score as a safety conclusion.
Why this methodology will continue to change
The original version of this page described Finnrick and Peptide Critic as the entire benchmark. That stopped being true as PB added first-party testing extraction, pricing coverage, Trustpilot, Pep Review Pro, curated profiles, and explicit scoring guardrails.
Future changes should follow the same rules:
- publish the inputs and caps;
- distinguish source-derived facts from PB-authored judgments;
- prevent thin evidence from reading like certainty;
- keep price, affiliate revenue, sponsorship, and trust in separate lanes; and
- update this page when production logic changes.
If the code and this explanation ever disagree, the discrepancy is a methodology bug. Please contact us so we can correct it.