GuidesJuly 20, 2026

How the Vendor Benchmark Works

The current PeptideBenchmark scoring methodology: source-backed and PB-evidence paths, confidence weighting, testing trust, review signals, pricing coverage, and scoring limits.


Current production methodology Code-checked July 20, 2026 · Replaces the original two-source model
10-pointcommon scale
2 pathsclearly labeled
Dailyscheduled refresh

What the PB score is—and is not

The PB score is a directional public-evidence score on a 10-point scale. It helps readers compare how much source-backed information PeptideBenchmark can currently verify about a vendor.

It is not a laboratory result, pharmacy credential, regulatory approval, or promise that the next vial or order will perform as expected.

The score currently considers different combinations of:

  • Finnrick testing records;
  • Peptide Critic ratings and review depth;
  • Trustpilot rating and review depth;
  • Pep Review Pro rating and review depth;
  • PeptideBenchmark’s structured testing-transparency review;
  • whether public pricing coverage exists;
  • whether PB has completed a curated review; and
  • whether an otherwise anonymous overseas supplier lacks a reviewed U.S. or curated presence.

Not every vendor has every input. The model therefore uses two scoring paths instead of pretending all vendors arrive with the same evidence.

One board, two scoring paths

01Outside sources 10.0max

Source-backed

Used when Finnrick or Peptide Critic covers the vendor.

What anchors the score

A normalized outside-source rating, weighted by the depth of supporting tests and reviews.

02PB research 9.2max

PB evidence

Used when PB tracks the vendor but neither outside source does.

What anchors the score

Testing transparency, provider-verifiable coverage, independent reviews, and visible pricing.

Path 1: source-backed vendors

Step 1: normalize the outside ratings

Finnrick now headlines percentage-based results. PB calculates a test-count-weighted average of Finnrick’s published per-peptide percentages, then divides that percentage by 10. When no rated-product breakdown is available, Finnrick’s public vendor composite is used as the fallback before the same conversion.

Peptide Critic’s 5-point rating is doubled. Both sources therefore enter the benchmark on the same 10-point scale.

Finnrick88%
÷ 10
PB conversion8.8
Common scale8.8 / 10
Peptide Critic4.6 / 5
× 2
PB conversion9.2
Common scale9.2 / 10
  • If both sources cover the vendor, the benchmark score is their average.
  • If only one covers the vendor, that normalized source score remains the benchmark score.

This raw benchmark is shown separately on vendor profiles. It is not necessarily the number that orders the board.

Step 2: measure confidence

A perfect score supported by one thin source should not automatically outrank a strong score supported by many tests and reviews.

The confidence model therefore considers:

  • Finnrick test count, with diminishing returns;
  • Peptide Critic review count, also with diminishing returns;
  • whether both sources cover the vendor;
  • Peptide Critic verification and batch-tracking indicators; and
  • a minimum confidence floor based on the type of source coverage.

Consensus vendors receive more starting confidence than one-source entries. The counts use logarithmic scaling, so the first meaningful body of evidence matters more than the difference between, for example, the 200th and 201st record.

Step 3: pull thin evidence toward a neutral anchor

The model starts from a neutral reference point:

  • 7.2 for a Finnrick + Peptide Critic consensus vendor;
  • 7.0 for a Peptide Critic-only vendor; and
  • 6.9 for a Finnrick-only vendor.

The raw benchmark score moves away from that anchor only as confidence increases:

base rank = neutral anchor
          + (benchmark score − neutral anchor) × confidence

This is why a lightly supported 10.0 can rank below a broadly supported 9.0.

Step 4: apply bounded, labeled adjustments

The base rank can then receive:

PB testing-transparency lift

Evidence

Rewards provider verification, archive coverage, named labs, certificate depth, and related testing evidence.

Trustpilot

+0.4 / +0.6

Maximum varies by source count and is scaled by rating and review depth.

Pep Review Pro

Up to +0.2

Scaled by the reported rating and the depth of reviews behind it.

Curated review

+0.15

Applied after PB completes its deeper curated-review workflow.

Overseas-supplier adjustment

−2.0

Applies to non-curated suppliers without a reviewed U.S. presence.

The final result is capped between 0 and 10.

The curated adjustment is intentionally small. It recognizes that PB has completed deeper identity, testing, pricing, and policy review; it is not an affiliate or sponsorship bonus. Curated vendors are exempt from the anonymous-overseas-supplier adjustment because they have already entered the reviewed-vendor workflow.

Path 2: PB-evidence vendors

Some vendors have meaningful public testing and pricing evidence but no Finnrick or Peptide Critic score. Giving them a fake outside-source rating would be misleading, so PB uses a separate evidence formula.

The PB-evidence score currently combines:

Testing-transparency evidence

PB measures the visible testing surface, including:

  • public laboratory or provider lookup links;
  • Janoshik public and verification links;
  • certificate coverage across the catalog;
  • named-lab and certification breadth;
  • vendor-hosted documents versus provider-verifiable records;
  • authenticated certificate archives when access has been reviewed;
  • available safety-panel coverage; and
  • the share of testing documents that can be corroborated at the provider.

Curated vendors use an 8× testing-signal multiplier; other PB-discovered vendors use 6× because curated coverage has gone through a deeper review workflow.

Trustpilot contribution

Trustpilot can contribute up to:

  • 4.0 points for curated PB-evidence vendors; or
  • 3.0 points for other PB-evidence vendors.

The contribution is scaled by both the rating and review-count confidence. Twenty reviews currently provide full review-count credit for this path; fewer reviews receive proportionally less.

Public pricing presence

A PB-evidence vendor receives one point when at least one public pricing row exists.

This is deliberately binary. A larger catalog does not earn more score, and a cheaper price does not earn more score. Price competition belongs on the Pricing board and in the Stack Builder, not inside a trust score.

Provider-verifiable coverage quality

A quality term can contribute up to 2.5 points when:

  • a meaningful share of testing documents can be verified through the laboratory or provider;
  • average verified purity is at least 98%; and
  • multiple documents are available.

Full count credit begins at five provider-verifiable documents. This prevents a vendor with one perfect certificate from receiving the same quality credit as a consistently verified catalog.

Remaining guardrails

  • Testing signals above 0.3 receive a small 0.5 depth bonus.
  • Curated review can add 0.15.
  • Pep Review Pro can add up to 0.2.
  • Trustpilot is already included in the PB-evidence base and is not counted a second time.
  • The completed score remains capped at 9.2.

Benchmark score, PB score, trust band, and curated assessment are different

These labels answer different questions:

Outside sources

Benchmark score

What Finnrick and/or Peptide Critic report after scale normalization.

Board order

PB score / rank score

The directional score after confidence weighting and disclosed adjustments.

Testing surface

Testing trust band

How strong the testing-transparency evidence reviewed by PB appears.

Editorial rubric

PB Curated Assessment

PB’s separate conclusion about the vendor’s broader trust and transparency posture.

Testing trust bands currently resolve as:

Strong≥ 0.70
Moderate0.40–0.69
Light0.01–0.39
NoneNo signal

A high PB score and a Strong testing band can coexist, but they are not synonyms. The Trust System exposes testing evidence separately so readers can inspect the underlying records instead of relying on a badge.

The PB Curated Assessment is also kept separate from the PB score. It is an editorial rubric, not another hidden scoring multiplier.

What affiliate relationships and sponsorship do not change

PeptideBenchmark may earn commissions when readers use disclosed affiliate links or coupon codes. Those relationships:

  • do not increase the PB score;
  • do not change the testing trust band;
  • do not change the PB Curated Assessment;
  • do not buy a curated listing; and
  • do not change the calculated ranking score.

Sponsored visibility is labeled and walled off under our sponsorship policy. An active sponsored placement may be pinned in a clearly labeled display slot, but it does not alter the vendor’s score or the underlying score-sorted order. A vendor can be curated without being an affiliate, and an affiliate relationship alone does not qualify a vendor for curation.

Coupon and quantity-discount data can change the effective price shown in Pricing or Stack Builder. It does not make the vendor more trustworthy.

The separation is intentional: commercial relationships can affect which outbound link or discount a reader sees. They cannot alter the score, testing band, or editorial assessment.

Why the same vendor can show several different numbers

Consider a hypothetical consensus vendor:

  1. Finnrick’s published percentages normalize to 8.8/10.
  2. Peptide Critic reports 4.6/5, normalized to 9.2/10.
  3. Its benchmark score is therefore 9.0.
  4. Evidence depth determines how far the base rank moves from the 7.2 neutral anchor toward 9.0.
  5. Reviewed testing transparency, Trustpilot, Pep Review Pro, and the small curated adjustment may then move the rank within the 10-point cap.

Now consider a vendor absent from both source directories:

  1. It has no legitimate Finnrick/Peptide Critic benchmark score.
  2. PB instead evaluates public testing evidence, provider-verifiable coverage, review signals, and whether pricing is observable.
  3. The result is labeled PB evidence and capped at 9.2.

The two vendors can appear on the same board, but their profiles disclose which path produced the number.

Freshness, fallbacks, and corrections

The vendor benchmark workflow is scheduled daily. Inputs do not all refresh on the same cadence:

  • Finnrick is retrieved through its public vendor source;
  • Peptide Critic may require a last-known-good snapshot when its live directory blocks hosted infrastructure;
  • Pep Review Pro is synchronized into the benchmark workflow;
  • Trustpilot is refreshed through its own scheduled process;
  • testing evidence and gated archives have separate review workflows; and
  • vendor pricing pipelines run independently of the score sync.

The generated benchmark records source timestamps and status. A stale upstream source is not silently presented as a fresh observation, and the pipeline can fail rather than accept an over-age fallback.

Vendor aliases and domain changes can create mismatches. We maintain explicit identity mappings and audit generated profiles for duplicate records, score-component drift, placeholder text, and public leakage of internal pipeline terminology.

What the board still does not prove

Even a 10.0 source-backed score or a 9.2 PB-evidence score does not prove:

  • the identity or concentration of the vial currently for sale;
  • sterility, endotoxin, or heavy-metals status unless a relevant current report supports it;
  • that a certificate belongs to the batch a buyer receives;
  • fulfillment quality on the next order;
  • long-term business durability or complete ownership transparency;
  • legality, approval, or medical suitability for human use; or
  • that reviews are representative, independent, or immune from manipulation.

The board is a comparison and due-diligence tool—not a warranty. Read What a High Vendor Score Does Not Prove before treating any score as a safety conclusion.

Why this methodology will continue to change

The original version of this page described Finnrick and Peptide Critic as the entire benchmark. That stopped being true as PB added first-party testing extraction, pricing coverage, Trustpilot, Pep Review Pro, curated profiles, and explicit scoring guardrails.

Future changes should follow the same rules:

  • publish the inputs and caps;
  • distinguish source-derived facts from PB-authored judgments;
  • prevent thin evidence from reading like certainty;
  • keep price, affiliate revenue, sponsorship, and trust in separate lanes; and
  • update this page when production logic changes.

If the code and this explanation ever disagree, the discrepancy is a methodology bug. Please contact us so we can correct it.