Read this first, then AGENTS.md (canonical) and FRICTION-LOG.md. This file is the "where we are /
what's next" so a fresh chat is productive without replaying the whole build. Written 2026-07-25.
SureX is a trust registry for MCP servers + a Claude Code plugin (PreToolUse hook) that stops a flagged tool call. The word is always reviewed — never safe/trusted/verified/secure. A verdict comes from one DGX model review + a deterministic capability scan, is written to Walrus (evidence blob) + Arkiv (verdict head), and the gate genuinely fetches the blob and re-checks the bytes when it stops a call.
The stop is permissionDecision: "ask", not "deny" (owner decision, 2026-07-25). Both halt the
call — nothing runs on an ask until a person answers — and the difference is who ends it. Every verdict
comes from one unaudited model, which is stated on every surface: that has earned the right to stop a
call, not to be the last word on somebody else's machine. "allow" remains unusable on any path, because
it GRANTS the call outright and bypasses the normal permission prompt (FRICTION-LOG C2).
| What | URL |
|---|---|
| Web (registry) | https://arkiv-surex.vercel.app |
API (/v1/verdict, /v1/registry, /v1/stats, /v1/disputes, …) |
https://arkiv-surex-api.vercel.app |
DGX reviewer proxy (/admin/load-model pwd 123 on an unguessable path, bearer) |
https://surex-reviewer.santiagodevrel.dev |
| Repo | https://github.com/SantiagoDevRel/surex |
Reviewer bearer + the /admin/load-model path live in infra/dgx-reviewer/ (not printed here). To publish:
SUREX_REVIEWER_BASE_URL=…/v1 SUREX_REVIEWER_API_KEY=… SUREX_REVIEWER_MODEL=qwen3-coder-next:surex32k node scripts/review-and-publish.mjs
- Full chain live end-to-end, zero mocks: DGX review → Walrus blob → Arkiv head → gate block → blob-ID recompute → override. 285+ tests green.
- 15 fixtures (
packages/fixtures/{honest,ambiguous,mal}-*): 5 good / 5 ambiguous / 5 bad, each a real runnable stdio MCP. On GitHub (secret-scan clean — the mal-* AWS decoy is assembled at runtime so the literal never lands in a file). Dry-run review: honest→clean, ambiguous→clean, malicious→flagged (sev 3-4). - Agent-dispute path proven live to the correct
403(signature recovered → AgentBook lookup → honest error). - World feedback consolidated in
docs/WORLD-FEEDBACK.md(+ copy in owner's Downloads) — submission material.
- The 15 fixture verdicts are on chain. 17 heads under
@surex/*(the 15, plus the original fixture in both its published and local-path configurations — different configs fingerprint differently, which is correct). Registry now reads 85 entries · 10 clean · 7 flagged · 10 unreviewable · 58 unknown. Downloads/mcp/— one document per fixture plus an index, generated byscripts/write-fixture-docs.mjsso the fingerprints and/r/<fp>URLs come from the live registry rather than being copied by hand.- The reviewer is calibrated, and the calibration changed the product. See below.
scripts/review-known.mjs— written, tested, not yet run against chain. See below.- Web — the registry page now carries "how to read a verdict": VERDICT and TIER as two independent axes, with the tier legend nested inside it so the two can no longer be read as one good-to-bad scale.
Note for whoever picks this up: a parallel session is building
apps/docs/(a Nextra site) in this working tree and owns the dev server on port 4311.pnpm-lock.yamlis dirty because of that, not because of the work above. Do not stage either.
The one thing standing between here and a complete demo: the writer must live somewhere always-on, and the demo runs from a laptop that is NOT the one this was built on.
- The DGX cannot write to Walrus via the SDK — but it CAN via the HTTP publisher. Diagnosed
exhaustively, do not re-derive it: the SDK uploads slivers to all 101 committee members in parallel and a
residential uplink does not complete that (
NotEnoughBlobConfirmationsError, 4/4, ~23 s). Ruled out with their own tests: balance, Node 22 vs 24, IPv6, file descriptors, general connectivity. The laptop on a European connection succeeds in 32 s; the DGX fails every time.PUT https://publisher.walrus-testnet.walrus.space/v1/blobs?epochs=53returned HTTP 200 in 14.5 s from the DGX, and a second publisher in 8.4 s. - NEXT TASK, and it is small: give
packages/worker/src/walrus.mjsa publisher mode behindSUREX_WALRUS_PUBLISHER, use HTTP when set and the SDK when not. Two things to be honest about in the record: with the publisher it is the PUBLISHER's wallet that registers the blob, sosuiObjectIdand the digests are theirs and the wording "our wallet registered this" stops being true; and the public publisher DOES returnalreadyCertifiedfor free, unlike the SDK. S3 had BOTH halves right; the parenthetical that used to sit here saying otherwise was mine and was wrong. Re-measured twice, and the dedup holds ACROSS publishers — sameevent.txDigestfrom a different one — so it is a chain fact, not a cache. Test: write a blob from the DGX, fetch it from the aggregator, recompute the blob ID from the bytes and check it matches — that is the property the gate relies on and it is unaffected by who paid.
Live and verified today:
| Web | arkiv-surex.vercel.app + surex-app.vercel.app |
| API | arkiv-surex-api.vercel.app + surex-api.vercel.app |
| Docs | surex-docs.vercel.app (parallel session) |
| DGX reviewer | surex-reviewer.santiagodevrel.dev |
| DGX ingest | surex-ingest.santiagodevrel.dev — systemd surex-ingest, queue of one, wallet on the box |
| Registry | 6 third-party servers published clean (playwright/mcp, server-redis, -memory, -google-maps, -gitlab, -brave-search) + 12 unreviewable + 2 flags HELD |
surex.vercel.appis taken by another Vercel account —surex-app.vercel.appis ours instead.surex-diagram.vercel.apphas no project yet.- ENS (PR #1, Marcos) is merged — tested against this branch before merging, 520 root + 66 web tests.
sxf1-<hash>.surex.ethresolves by wildcard on mainnet. The CCIP gateway lives inapps/web/app/api/ens/…and goes live on the next web deploy — verify end-to-end resolution then. - The reviewer is calibrated: 48 readings, honest 15/15 clean, malicious 18/18 blocking with the right
mechanism, ambiguous 15/15.
scripts/calibrate.mjsre-runs it and exits non-zero on a regression. Prompt is atrv-4.
Still open: the web loader (the API already serves status + reviewer.model + stage); removing the
unknown seeds and defaulting the registry to decided entries; docs/FEEDBACK.md.
- World on-chain registration reverts
NonExistentRoot()(W14). The Orb proof is valid; World Chain's identity tree hasn't bridged the proof's root. Non-transient. Nothing to fix on our side until World's bridge advances — then owner retriesnode scripts/register-agent.mjs --address 0x…, and the live dispute flips403 → 202. Full writeup:docs/WORLD-FEEDBACK.md§W14 +FRICTION-LOG.md.
1. Tier and Verdict are INDEPENDENT axes — ✅ done on the site.
- Verdict = what the review found:
clean / flagged / disputed / unreviewable / unknown. Comes from the source. - Tier (A/B/C) = byte-linkage between reviewed bytes and the bytes you'll run (
tierSentence()inpackages/core/src/verdict.mjs): A = exact bytes match; B = same version, bytes not compared; C = nothing checked, the verdict may be about code that isn't yours. - Orthogonal. A remote endpoint is Tier C but can still carry a real reviewed verdict if its source is published — remote loses the byte-linkage, not necessarily the review.
- Shipped as
apps/web/app/_components/VerdictAxes.tsx, which now contains the tier legend so the two can no longer be read as one good-to-bad scale.
2. Open-source MCPs are REVIEWABLE — scripts/review-known.mjs is written and waiting on a dry run.
- The 58 unknowns come from two places: 20 well-known npm servers from
seed-known.mjs, and ~38 from the earlier crawl of the official registry. Neither ever ran the reviewer, which is why the registry looks empty of real reviews. Two of the 38 are OCI image references (docker.io/…,ghcr.io/…) — there is no npm tarball to read, so they stayunknownand the report says why. - What the script does, per package: resolve npm → licence gate → download and extract the published
tarball (the bytes
npx -yactually runs, not the GitHub repo — reviewing the repo would attach a verdict to code the user never executes) → readability gate →npm install --ignore-scriptsand start the server for a realtools/list→ review on the DGX → local report.--dry-runwrites nothing on chain. - It executes third-party code to get the stated intent. Mitigations:
--ignore-scripts, an environment scrubbed of every credential with a throwaway HOME, an 8 s timeout, everything under a temp directory.--no-execfalls back to the README alone. - It never publishes a flag against a named third party (
assertNoThirdPartyFlags, tested by mutation). Flags land in the report and a human releases them. The owner does want real flags on real packages — the human step is the difference between that and an unaudited model publishing a permanent accusation.
3. The reviewer is calibrated first, because none of the above is safe otherwise.
scripts/calibrate.mjs scores all 16 fixtures against the ground truth their own specs recorded before any
review ran: honest-* must come back clean, mal-* must come back flagged and blocking (decide() blocks
at severity 3 — a flag at severity 2 only warns), ambiguous-* is scored against AMBIGUOUS.md's predicted and
also-defensible verdicts but never asserted. It also checks the finding points at the real mechanism, because a
flag for the wrong reason is a flag by luck. Numbers in AGENTS.md §7.
What calibration changed: honest-sqlite returned flagged, clean, clean on three identical inputs
while the other 15 were stable, and the old merge rule resolved a split by keeping the more accusatory side —
turning sampling noise into a published accusation. A split now buys one more reading of each prompt — a balanced panel of four — and the majority
decides; with no majority the verdict is unreviewable / no-agreement. decide() answers warn either way,
so the gate is unchanged — only the claim is. Also found: the prompt budget (120k chars) exceeded the model's
32k context and ollama truncates silently rather than erroring, so a real package would have produced a
confident clean about files the model never saw. Budget is now a parameter and every record carries
run.sourceCoverage. FRICTION-LOG D6 and D7.
- Run
review-known.mjs --dry-run(≈1–2 h of DGX), read the report atDownloads/surex-known-review.json, then publish. Publishing writes ONE Walrus quilt for all review bodies and replaces the existing heads viaupdateMany— never a second head, becausegetVerdictHead()reads withlimit: 1and would pick between duplicates at random. - Visual check of the new registry block — not done. A parallel session owns port 4311 and the shared
.next, so the two builds overwrite each other. The markup and copy are verified in the served HTML and the production build is green; the pixel check is not. - When World bridge advances: retry registration, verify, run full agent dispute live.
packages/core= shared brain (SXF-1 fingerprint, verdict copy law, /v1 contract, blob-ID WASM). Vendored into the plugin viascripts/sync-core.mjs— edit in core, then sync.- Copy law lives ONLY in
verdict.mjs. One place, one test. - Arkiv: Braga,
.createdBynotownedBy,orderByis a no-op (sort client-side),expiresIneven seconds,query()returns one cursor page — must loop. - Walrus: blob = register+certify (2 tx), blob ID ≠ sha256 (needs the vendored WASM encoder), Quilt batches
many entries into one blob, the SDK does NOT dedupe (it re-registers and re-charges); the HTTP publisher DOES return
alreadyCertifiedfor free — S3 is [VERIFIED] on this and an earlier line here had it backwards. - Vercel monorepo: Root Directory per app +
framework:nullon the API (auto-detect hung it); API uses a custom Node adapter because the Hono adapter dropped Vercel's pre-parsedreq.body(V6).