Endpoint Trust Score

Reproducible reliability history for x402 endpoints: the verdict surface of the Forum Labs network.

Thousands of paid endpoints are appearing faster than any durable record of how they behave. The Trust Score is that record, made consultable: an append-only probe history reduced to a reproducible verdict an agent can check before it pays. GET /v1/trust/score?url=<endpoint> costs $0.005; /v1/trust/report adds history detail for $0.02. Payfetch consults this signal automatically (advisory by default) before releasing a payment.

We measure endpoints; we do not rank, endorse, or accept payment from the endpoints we measure. Metrics are mechanical and every score ships with its underlying counts.

Read the full methodology →


Methodology

Scores are computed by a pure function over an append-only history of periodic probes. Weights, bands, and windows are frozen and versioned (thresholdsVersion: p2t-score-1.2.0) and stamped on every score, so any third party holding a history slice can recompute any score. Ratings that can't be audited are marketing. The full methodology gives the component formulas, the evidence-gated rating rule, and the score caps in detail.

What we measure

Most signals come from the endpoint's unpaid protocol handshake (the HTTP 402 challenge), which costs nothing and reveals the advertised terms:

ComponentWeightWhat it captures
Availability0.50Does it respond over time (confirmed observations only), weighted by wall-clock time, so a long outage counts as the days it lasted, not the few times we polled it. An endpoint that is currently down cannot score “reliable” on a stale average.
Latency0.15p50/p95/p99 from raw timings, the tail an agent pipeline actually feels.
Terms stability0.20Price drift; a change of payment recipient (custody) is flagged, never averaged away.
Challenge integrity0.15Does it present a well-formed, parseable payment challenge.

Today’s score is computed from the unpaid protocol handshake above; a capped, fully-paid settlement spot-check (paying the advertised price like any customer) is on our roadmap.

The negative record

Most of what the record has to say today is negative, and that is the point. The prober’s append-only history captures the events that matter to a buyer and that nobody can reconstruct after the fact: endpoints that die and endpoints that resurrect, payment recipients that quietly change (a custody change is flagged, never averaged away), advertised prices that drift, and paywalls that never once present a valid challenge. A directory listing tells you an endpoint exists; the record tells you whether it has ever behaved.

Verdicts

reliable (≥80) · mixed (≥50) · unreliable · and unrated. An endpoint with too little history is unrated with no score, never a fake-neutral number. An endpoint we have never once observed presenting a payment challenge cannot be called reliable, however fast and stable its front door: we have no evidence it is a working paywall, so it is capped at mixed and flagged. Rather than suppress a confirmed-negative endpoint for weeks waiting on an observation quota, we rate it unreliable as soon as the evidence is decisive (marked provisional until the quota is met). A failure history recorded under only one HTTP method, when the endpoint answers but the complementary method was never tried, is not decisive: it returns unrated (method_unverified) until both methods have been probed.

Principles we hold ourselves to


For endpoint operators

Our monitor identifies itself honestly (User-Agent forum-labs-trust-prober) and is a good citizen: conservative cadence, exponential backoff, and it honors 429 / Retry-After. We probe only endpoints publicly listed on public directories, in the manner they advertise for consumption.

Re-probe, correction, or opt-out: email ops@forum-labs.com. Opt-outs are honored within 24 hours; corrections are appended to the record (append-only; we never silently edit it).

Forum Labs · Payfetch · Trust Score · Token Safety · Methodology · GitHub · ops@forum-labs.com · @shopforumlabs
© 2026 Forum Labs. Not financial advice. Metrics are informational and mechanical.