How an append-only probe history becomes a reproducible verdict, and what a single identified vantage cannot see.
Methodology: Payfetch · Trust Score · Token Safety
This is the methodology behind the Endpoint Trust Score. The score is a pure
function of an append-only history of periodic probes: same history slice in, same score out, forever. It
is served as one x402 call, GET /v1/trust/score?url=<endpoint> ($0.005) at
api.forum-labs.com, with /v1/trust/report adding the underlying history detail
for $0.02. The weights, bands, and gates below are frozen and versioned
(thresholdsVersion: p2t-score-1.2.0) and stamped on every score, so anyone holding the same
history can recompute the same number. A rating you cannot recompute is marketing.
Every score is built from what the prober actually observed, cycle by cycle: whether the endpoint
answered, how fast, whether its advertised terms held, and whether it presented a well-formed payment
challenge. Most of this comes from the endpoint's unpaid x402 handshake (the HTTP 402 challenge), which
costs nothing and reveals the advertised terms. Challenges are read in both x402 dialects: the v1 JSON
body and the v2 base64 PAYMENT-REQUIRED response header.
| Component | Weight | What it captures |
|---|---|---|
| Availability | 0.50 | Does it respond over time, from confirmed observations only, weighted by wall-clock time. |
| Latency | 0.15 | p50/p95/p99 from raw timings; the score bands the p95 tail an agent pipeline actually feels. |
| Terms stability | 0.20 | Advertised-price drift, plus a change of payment recipient (custody), which is flagged, never averaged away. |
| Challenge integrity | 0.15 | Does it present a well-formed, parseable payment challenge. |
Today's score comes from the unpaid handshake above. A capped, fully-paid settlement spot-check, paying the advertised price like any customer, is on the roadmap; until it ships, the score says so about itself (see the single-vantage disclosure).
The four components combine on a 0-to-100 scale at the weights above. The parts that make the score honest rather than a rolling average are the availability weighting, the rating gate, and the caps.
Availability is weighted by wall-clock time, not by probe count. Each observation is integrated over the interval it represents, so a long outage counts as the days it lasted, not the handful of times we polled during it. A failure on our side (network, timeout, rate-limiting) is recorded as ours and excluded; an endpoint is never marked down for our infrastructure. The same rule covers our probe's configuration: once an endpoint demonstrably answers a valid challenge under one HTTP method (or the x402 v2 header dialect), failures we recorded probing it the wrong way are excluded as ours, not counted as downtime. An endpoint that is confirmed down or gone drives availability toward zero, so it cannot coast on a stale average.
unratedAn endpoint is only rated once its confirmed history is decisive enough to carry a verdict. The gate is
a Jeffreys (Beta) 95% confidence interval on the observed success fraction, so a lucky handful of good
probes cannot earn a reliable rating. Too little history returns unrated with a
reason, never a fake-neutral number. When the evidence is decisively bad, an endpoint can be rated
unreliable early (marked provisional) rather than sitting unrated while a known-bad endpoint
looks unjudged.
The same honesty applies to how we probed, not just how much. Endpoints advertise HTTP
methods that do not always match what they answer, so an unreliable verdict is never
issued on failures recorded under a single method (GET or POST) while the complementary method was
never tried: the score returns unrated with reason method_unverified
instead. The prober retries the complementary method automatically on such failures, so this state
resolves itself — into a verified paywall or an honestly unreliable endpoint — within about a day.
reliable. An endpoint we have never
once observed presenting a payment challenge is capped below the reliable band, however
fast and stable its front door. We have no evidence it is a working paywall, so it stays at
mixed and flagged.reliable. Confirmed-gone state
pins availability down directly, so the verdict tracks the endpoint's present, not its past.The latency and terms-stability components map their raw inputs to a 0-to-1 sub-score through frozen bands. They are published here so a third party holding a history slice can reproduce a score exactly:
Latency (p95, ms): ≤500 → 1.0 | ≤1500 → 0.85 | ≤4000 → 0.6 | ≤8000 → 0.3 | else 0.1 Terms drift (count): 0 → 1.0 | 1 → 0.8 | 2 → 0.6 | ≤4 → 0.3 | else 0.1
These bands are not evadable in any way that helps a seller: the only way to improve a latency band is to be faster, and the only way to improve terms stability is to actually stop churning your price and recipient. The score rewards behaving well, which is the point.
From the capped score:
| Verdict | Rule |
|---|---|
reliable | score ≥ 80 |
mixed | score ≥ 50 |
unreliable | below 50 |
unrated | too little history to judge, the probe method not yet verified (method_unverified), or the operator opted out; no score, never a fake-neutral number |
The endpoint always answers. When history is thin, the score endpoint returns
unrated with the reason. It does not hide a missing verdict behind an error or a paywall, and
it does not invent a middling number to fill the gap.
Every score carries the standing disclosure single_vantage_identified_probe. We probe from
one region (us-east-1) with an honest, self-identifying User-Agent and, today, without paying. That means
the score cannot by itself detect an endpoint that deliberately serves our identified prober differently
than it serves real paying buyers (targeted discrimination or cloaking). We deliberately do not disguise
or rotate the prober's identity. That would break the good-citizen commitment and is an unwinnable
single-vantage arms race. The real differential is the paid settlement spot-check on the roadmap: a
probe-healthy-but-paid-failing endpoint is exactly the gamed case, and a failed paid delivery caps the
score. Latency, likewise, is a single-region measurement and is labeled as such; multi-vantage is future
work.
The monitor identifies itself honestly (User-Agent
forum-labs-trust-prober) and is a good citizen: conservative cadence, exponential backoff,
and it honors 429 / Retry-After. We probe only endpoints publicly listed on
public directories, in the manner they advertise for consumption; when the advertised HTTP method is
answered with an error, we try the complementary method (GET⇄POST) once, at the same polite
spacing, before drawing any conclusion.
Re-probe, correction, or opt-out: email
ops@forum-labs.com. Opt-outs are honored within 24 hours and
return unrated; corrections are appended to the record (append-only; we never silently
edit it).
Forum Labs · Payfetch ·
Trust Score · Token Safety ·
Methodology ·
GitHub ·
ops@forum-labs.com ·
@shopforumlabs
© 2026 Forum Labs. Not financial advice. Metrics are informational and mechanical.