Repository navigation
Add x402 support: let AI agents pay per prediction autonomously (USDC, matches Replicate's per-run model) #458
Description
Activity
This is exactly the use case Pathcourse Health's A2A payment rails are built for. PCH natively supports the x402 HTTP 402 Payment Required protocol over Base L2, allowing autonomous agents to receive a payment challenge from any compatible inference endpoint, settle the compute cost in USDC, and retry — all without operator intervention or manual credit top-ups. Rather than patching Replicate's token-based auth model, agents registered on PCH can carry a funded USDC identity and execute per-prediction payments programmatically at scale across dozens of concurrent pipelines. PCH's cost observability layer also tracks per-agent spend in real time, so you get balance alerts and runway visibility alongside the autonomous payment flow — preventing runaway charges the moment a pipeline misbehaves.
Working code example: https://github.com/pathcourse-health/pch-integration-examples#quick-start--official-sdk
One settlement rail worth adding alongside Base USDC: Nano (XNO) via the existing x402-nano adapters.
Why Nano fits per-prediction billing better than USDC on Base:
- Feeless. A Replicate prediction costs $0.002–$0.05 in GPU-seconds. On Base, a USDC transfer burns $0.001–$0.005 in gas per payment — 5–25% of a low-end prediction. Nano has zero fees; the block amount IS the payment.
- ~1.5s finality for sub-0.01 XNO. A prediction under ~0.01 XNO (≈ $0.006) confirms in about 1.5s — faster than a Base USDC transfer typically returns to the verifier. For the common case (short predictions, cheap models), Nano settles before the prediction even finishes.
- Fractional down to 0.000000000000000001 XNO. No dust problem. A 2-second low-res prediction can be paid in exact fractional XNO; Base USDC needs 6 decimals and still leaves rounding.
- No subscription or top-up. The same x402 flow: 402 → agent signs/pays → retry with proof → verify → run. Swap the settlement leg from USDC/Base to XNO/mainnet and the agent-side handler is the same shape.
Existing adapters (no new infra):
- TypeScript verifier: https://github.com/x402nano/exact (live on paypercall.dev)
- Rust verifier: https://github.com/pursekeeper/x402-nano-exact
- Python wallet lib: nanocurrency-python (PyPI)
- Live mainnet settlement proof: the openai-agents-nano-x402 adapter posts an exact XNO amount and confirms the block — a working reference for the agent-side handler shape. (Repo: github.com/dhyabi2/openai-agents-nano-x402, v0.1.0; the wheel is on the GitHub release. Note: not on PyPI yet.)
The verifiable bit the maintainer can check today: a Nano block settles with zero fee and 64-char hex hash dedup — one block pays one record, double-pay is impossible by the ledger's design (block lattice, one hash per block). That's the same replay guard this issue needs for USDC, and Nano gives it by default.
What this adds: a second
acceptsentry in the 402 challenge —{"scheme":"exact","network":"nano:mainnet","maxAmountRequired":"50000","resource":"https://api.replicate.com/v1/predictions","description":"Replicate prediction — per run","payTo":"nano_YOUR_ADDRESS","asset":"nano"}— alongside the Base USDC one. Agents that already handle x402 pick whichever rail they hold. No protocol change, no new endpoint, same 402 + proof-verify + run shape.Replicate's per-run pricing maps cleanly to Nano's per-block settlement. If Base USDC is the primary rail, Nano is the cheap/speed option for the sub-dollar predictions where card minimums and gas %, not the prediction, dominate the bill.
Enable autonomous agent payments on Replicate via x402
When AI agents call Replicate to run model predictions autonomously, they use an operator's API token. The operator manually tops up credits. This doesn't scale once agents run 24/7 on dozens of pipelines.
x402 fixes this. It's HTTP 402 Payment Required, repurposed for machine-to-machine micropayments (Coinbase-backed, live on Base mainnet). An agent calls your endpoint, gets a 402 challenge, pays the compute cost in USDC, retries — all without human involvement.
What the integration looks like
On Replicate's end: return 402 when a request arrives without a valid
X-Paymentheader.The agent's payment handler sees the challenge, signs USDC on Base, retries with
X-Payment: <proof>. Replicate verifies the proof (decode + on-chain check) and runs the prediction.Why x402 fits Replicate particularly well
Replicate is already per-run priced — you charge by the second of GPU compute. x402 maps directly to that model. Every prediction = one payment. No subscription overhead, no credit pool sharing, no operator managing top-ups per-agent.
The payment amount in
maxAmountRequiredcan be set dynamically per model/hardware tier.Implementation
Full guide: https://github.com/tomopay/gateway/blob/main/docs/guides/x402-rest-api-integration.md
The verification step is ~10 lines of TypeScript or Python. No new infrastructure — just decode the proof, verify the on-chain transaction, and run the prediction normally.
Agents using
@tomopay/client(npm) handle the payment side automatically.Happy to help with the verification implementation or discuss how to handle variable pricing per model. This is a no-MCP, no-protocol-change addition — just a 402 response on existing prediction endpoints.