# Toll Compute Router (beta)

Discover hosted inference providers without an account or payment:

```sh
curl https://tollpay.shop/api/compute \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen/qwen3-32b","context":32768,"output":512,"budget":0.05,"latency":0,"sort":"price","region":"any"}'
```

The response includes eligible provider tags, catalog prices per token, cost
ceilings based on the full requested input allowance (context minus output),
maximum output and any per-request fee, plus available historical latency.
`latency: 0` permits unknown latency. Positive latency thresholds exclude unknown
or slower observations, but do not guarantee future speed. `sort` accepts `price`
or `latency`. Only `region: any` is supported; provider location is not verified.
An empty routes array means no match. 400 means invalid requirements, 429 means
rate limited, and 502 means the upstream catalog is unavailable.

## Execution

Use /inference#compute-router to connect your OpenRouter account, compare routes,
review one provider and authorize a request. A review expires in 60 seconds.
Toll refreshes provider metadata before execution and blocks changed prices or
unavailable routes. Execution uses `provider.only: [tag]`,
`allow_fallbacks: false`, `require_parameters: true`, and `max_price` matching
the reviewed prompt/completion rates (converted to USD per million tokens) and
per-request price. No automatic paid retries are performed.

Agents can execute directly against OpenRouter's chat/completions API using their
own key and these same provider controls. Re-query before execution. Never send
keys to the public Toll discovery endpoint.

The cost calculation is a planning envelope, not an exact tokenizer quote or a
provider account spending cap. Additional account fees may apply. Configure a
key spending limit at OpenRouter. The browser applies a conservative UTF-8 byte
check with a 256-token formatting allowance; provider tokenization is authoritative.

## Billing and receipts

Execution spends OpenRouter credits. It does not settle in Arc USDC or directly
charge X Money. Downloadable receipts contain the provider response ID, route,
model, review amounts and returned usage/cost. They are usage records, not proof
of blockchain settlement. Missing cost remains unknown. Receipts and responses
are held in component memory; download before leaving. Prompts and keys are not
included in receipts. On a timeout, inspect OpenRouter activity before retrying:
upstream work may have been charged even if no response reached the browser.

## Arc payments and GPU rental

Open `/compute` for two additional paths:

- Direct Arc compute: enter an actual MPP/x402 provider endpoint and request
  body. The live payment challenge must use Arc USDC and fit your per-request
  ceiling (maximum 1 USDC). Review the recipient and sign in your wallet. Toll
  reuses the saved authorization on retry. The provider handles settlement and
  compute delivery; a response record is not independent settlement proof.
  Toll does not turn an ordinary OpenRouter or Runpod endpoint into an Arc
  merchant, provide operator-funded credits, or promise third-party supply.
- Runpod GPU rental: bring a scoped Runpod API key, load Secure Cloud offers,
  review one GPU and your Runpod template, then explicitly create a pod.
  Template settings, including persistent storage, can incur extra charges.
  Catalog hourly rates are not total spending caps. Poll status with Refresh my
  pods; creation is asynchronous. Use the Runpod console to connect to the pod.
  Termination requires typing the exact pod ID and deletes local pod data.
  Independent volumes remain billable until removed in Runpod.

Runpod keys are held in browser memory and forwarded only to api.runpod.io by
the fixed-host `/api/gpu` adapter. No raw keys, template environments or full
provider responses are saved. Creation requires durable Vercel Blob storage:
a hashed-key/request-ID marker prevents resubmitting the same creation after a
timeout. It is not released on an ambiguous response. Refresh pods or inspect
the Runpod console before preparing a separate rental. There are no automatic
provisioning retries, spending loops or browser-dependent shutdown guarantees.

Verification uses simulated provisioning/termination and wallet settlement;
live paid GPU creation requires the user's funded account. Verified regional
routing and cross-provider Arc-to-Runpod billing are not offered.
