What is the CPI?
The Compute Price Index ("CPI") is a public benchmark that tracks the cost of AI inference across major providers. It produces a single equal-weighted reference price — the Standard Compute Unit ("SCU") — calculated from a basket of 22 models across 9 providers.
The Problem
LLM inference pricing is fragmented. Every provider quotes pricing differently — per token, per character, per request, per model. There is no standardized way to benchmark or compare the cost of AI computation across models and providers. You need a verifiable pricing reference you can audit independently.
The Solution
The CPI is an equal-weighted basket of 22 AI models across 9 providers that produces a single, verifiable unit of account: the Standard Compute Unit (SCU). The SCU represents the USD cost of a reference workload, 1,000 input tokens + 500 output tokens, taken as the geometric mean across the full basket.
Key Properties
- Diversified — 22 models across 9 providers prevent any single provider from dominating the index.
- Outlier-resistant — The geometric mean dampens any single model's impact, and the one-family-one-slot rule prevents version pile-up. No caps needed; the math does the work.
- Versioned — The equal-weighted basket launches as v1.0 on June 18, 2026. Reconstitution-driven, published on-chain when basket composition or provider rate-card prices change. There is no fixed cadence.
Basket Composition
The CPI tracks 22 models from 9 providers. Every model family gets one slot, the latest version always, and each model is weighted equally. A family is a provider's distinct product line (e.g. openai.gpt, anthropic.claude); the latest released model in each family is its representative. On a new release the representative auto-rolls. The basket updates with a new version number when families are added, removed, or replaced.
alibaba.qwen-flash
Qwen3.8 Flash
alibaba.qwen-max
Qwen3.8-Max
alibaba.qwen-plus
Qwen3.7 Plus
anthropic.claude-fable
Claude Fable 5.1
anthropic.claude-haiku
Claude Haiku 4.5
anthropic.claude-opus
Claude Opus 5.5
anthropic.claude-sonnet
Claude Sonnet 5
deepseek.v-flash
V4 Flash
deepseek.v-pro
V4 Pro
google.gemini
Gemini 2.5 Pro
google.gemini-flash
Gemini 3.8 Flash
google.gemini-flash-lite
Gemini 3.5 Flash-Lite
minimax.m
MiniMax M3
moonshot.kimi
Kimi K3
openai.gpt
GPT-6 Astra
openai.gpt-luna
GPT-6 Luna
openai.gpt-mini
GPT-5.4 Mini
openai.gpt-nano
GPT-5.4 Nano
openai.gpt-sol
GPT-6.1 Sol
openai.gpt-terra
GPT-5.6 Terra
xai.grok
Grok 4.6
xiaomi.mimo
MiMo V2.5 Pro
SCU Formula
The Standard Compute Unit is calculated in two steps. The reference workload is fixed at 1000 input + 500 output tokens, evaluated across all basket models.
The SCU is reconstitution-driven. It is published on-chain whenever the basket changes, either because a model is added, removed, or replaced, or because a provider updates a rate-card price. Both triggers produce a new on-chain version. There is no fixed cadence and no continuous off-chain feed.
Step 1 — Per-model cost
For each model in the basket, compute the cost of the reference workload using the provider's published per-token pricing. Prices are USD per 1M tokens.
Step 2 — Equal-Weight Geometric Mean
Take the geometric mean of all N model workload costs. Every model is weighted equally; there are no tiers, no caps, and no provider weights. Each model family holds one slot at its latest version, so a provider cannot inflate its representation by shipping extra SKUs.
Live SCU: $0.003214 — the geometric mean of 22 equal-weighted model costs.
Every published revision permanently carries the methodology version that produced it (currently v1). Definitions and the changelog are served by the methodology endpoint; every oracle response carries the X-Methodology-Version header.
Anyone with access to provider pricing pages can independently reproduce this number.
Read the full methodology specification →Governance
The model-family rule reduces governance to a minimum. One eligibility rule, one family rule, one emergency trigger. No weight committees, no tier reviews.
| Rule | Specification | Frequency |
|---|---|---|
| Listing Criteria | Public GA + first-party USD pricing + in-scope text model + seasoning + liveness. Applied identically to all. | Continuous |
| Model Family Rule | One slot per family, latest GA version auto-wins. No governance needed. | Auto |
| Edge-Case Review | Published, dated decision for family-vs-version edge cases. Stated reasoning. | As Needed |
| Reconstitution | Additions and removals applied by the criteria. Logged. | Scheduled |
| Price Updates | Reflected automatically from first-party price cards. Daily scan. | Daily |
| Emergency Trigger | Any single model that moves price by more than the set threshold in 24 hours enters a short hold before the change is reflected. | As Needed |
Every basket change increments the revisionVersion counter. The full history is available via the GET /v1/oracle/reconstitutions endpoint. On-chain, the OracleRegistry contract records each published revision with its timestamp and metadata hash.
On-Chain Verification
All CPI pricing is recorded on-chain via the OracleRegistry contract on Base. You can independently verify prices without trusting the Compute Finance API.
Contract Address
| Network | Base (Chain ID 8453) |
| Contract | 0x1b91c0961928a14a2eD6c1985bC11aF1b302714D |
Read Functions
| Function | Description |
|---|---|
version() | Oracle implementation version — assert before deserializing tuples |
decimals() | Scale of scuUsd and baseline values (18) |
lastRevisionVersion() | Returns the highest confirmed revision number |
getLatestRevision() | Returns the latest revision header: revisionVersion, methodologyVersion, scuUsd, contentHash, metadataHash, publishedAt |
getRevision(uint256 revisionVersion) | Returns the same revision header tuple for a specific revisionVersion |
getRevisionScuUsd(uint256 revisionVersion) | Returns the SCU value in USD 18-dec for a specific revision |
getBaseline() | Returns the SCU value of the first revision (denominator of the inverse purchasing-power index) |
getComputeIndex() | Returns the latest (baseline / SCU) × 100 |
getRevisionAt(uint256 timestamp) | Returns the revision version active at the given Unix-seconds timestamp |
getComputeIndexAt(uint256 timestamp) | Returns the atomic tuple (scuUsd, indexValue, publishedAt, methodologyVersion, revisionVersion) at the timestamp |
Verification with ethers.js
import { ethers } from "ethers";
const ORACLE_REGISTRY = "0x1b91c0961928a14a2eD6c1985bC11aF1b302714D";
const ABI = [
"function lastRevisionVersion() view returns (uint256)",
"function getLatestRevision() view returns (tuple(uint256 revisionVersion, uint16 methodologyVersion, uint256 scuUsd, bytes32 contentHash, bytes32 metadataHash, uint64 publishedAt))",
"function getBaseline() view returns (uint256)",
"function getComputeIndex() view returns (uint256)"
];
const provider = new ethers.JsonRpcProvider("https://mainnet.base.org");
const oracle = new ethers.Contract(ORACLE_REGISTRY, ABI, provider);
// Read latest revision header
const latest = await oracle.getLatestRevision();
console.log("revisionVersion:", latest.revisionVersion.toString());
console.log("SCU (USD, 18-dec):", ethers.formatUnits(latest.scuUsd, 18));
console.log("metadataHash:", latest.metadataHash);
// Inverse purchasing-power index — (baseline / SCU) × 100
const computeIndex = await oracle.getComputeIndex();
console.log("Compute Index:", ethers.formatUnits(computeIndex, 18));
// Per-model prices live in the off-chain manifest fetched by metadataHash —
// resolve at /v1/oracle/manifest/{metadataHash} and verify with JCS+keccak256.Verification via Basescan
You can also read the contract directly on Basescan. Since OracleRegistry is an upgradeable proxy, use the Read as Proxy tab so the implementation's functions are available:
- Go to Basescan → Contract → Read as Proxy
- Call
getLatestRevision()— its first return value lists every registered model pricing key - Read the on-chain revision header (scuUsd, methodologyVersion, contentHash, metadataHash, publishedAt) via
getRevision(version). Per-model prices live in the off-chain manifest fetched by metadataHash. - Prices are returned in $COMPUTE wei (18 decimals) per 1M tokens. Divide by 10^18 to get the $COMPUTE amount.
Verification via Sourcify
The same source is independently verified on Sourcify, a decentralized verification repository: View on Sourcify ↗
Public API
All oracle endpoints are public and require no authentication. Read access is free and unrestricted. Base URL: https://api.compute.finance
Endpoints
Every price the catalogue publishes carries a provenance mark — verified, inferred or promotional — on the number itself rather than on the model, so one model can be promotional at its flat rate and inferred on a cache component at the same time. A mark records whether an operator captured a vendor source for that number; it is set by hand and holds as of their last pass rather than as a live check, and it never changes what a request is billed.
/v1/oracle/scuCurrent SCU value, methodology version, and family-representative breakdown/v1/oracle/modelsCatalog of basket models plus retired ex-members — `inBasket` flag and `retiredAtRevision`; retired models are priced from the live catalog. `usdPricePerMillion` and the cache / reasoning blocks are the vendors' list prices; `markedUpUsdPricePerMillion` is the same price with the routing markup, as billed/v1/oracle/models/{key}Single catalog model (basket member or retired) by model id — the canonical `vendor/model` form, or the bare model name/v1/oracle/models/{key}/price-historyPer-model input/output USD price time series — one entry per on-chain revision for basket members, one per catalog price change for every other model (per-revision, daily, or weekly granularity); Accept: text/csv returns a CSV attachment/v1/oracle/catalogEvery tracked model (incl. catalog-only and pricing-only providers) with current price, integrated flag, index-member flag, the contextTiers ladder on models priced per context length, and a provenance mark on every number — verified (an operator recorded a vendor source for it), inferred (derived, no source recorded) or promotional (a discount that will end). Prices here are the vendors' own, without the routing markup./v1/oracle/models/{key}/price-at?date={iso8601}Per-model input/output USD price effective at the requested ISO-8601 timestamp; basket members resolve from the attested manifest, all other models from the live catalog — every response labels its source/v1/oracle/basketFull basket composition: all models with equal weights, revision and methodology version/v1/oracle/methodologyActive methodology version and full changelog/v1/oracle/methodology/{version}Single methodology record (formula, family rule, reference workload, spec URL)/v1/oracle/baselineFrozen SCU of the first confirmed revision — the denominator for the inverse computeIndex purchasing-power view/v1/oracle/scu-at?date={iso8601}SCU value active at a given timestamp — step-function lookup of the latest confirmed revision with publishedAt ≤ date (highest revisionVersion on ties)/v1/oracle/history?from={iso8601}&to={iso8601}&granularity={per-revision|daily|weekly}Historical SCU values, one entry per on-chain revision; granularity: per-revision, daily, or weekly; Accept: text/csv returns a CSV attachment/v1/oracle/reconstitutionsNamed basket-change events with version, models, SCU delta/v1/oracle/reconstitutions/exportExport full reconstitution history as a downloadable markdown file/v1/oracle/revisions/{revision}Full OracleRevision record for a specific on-chain revision/v1/oracle/latestLatest confirmed revision summary (version, timestamps, SCU, basket size)/v1/oracle/healthLast revision timestamp, version count, on-chain sync status/v1/oracle/contract-metadataLive OracleRegistry identity — chainId, proxy address, on-chain version(), and keccak256 of the deployed bytecode/v1/oracle/statsPublic aggregate protocol statistics (TVL, users, volume)/v1/oracle/activityLive activity feed of recent protocol events (paginated)/v1/oracle/pricingPer-model pricing as billed — every figure already carries the routing markup, cache and reasoning components included; `/v1/oracle/models` publishes the same models at the vendors' list prices without it/v1/oracle/resolve/{key}Resolve one model name to its canonical `vendor/model` id, prices and priceSource — percent-encode the separator (`openai%2Fgpt-5.5`) or pass the bare name. An unknown model is 200 with priceSource `off-basket` and null prices, never 404; an unknown vendor is 422/v1/oracle/resolve?keys={csv}Resolve up to 100 comma-separated names in one call, input order preserved; every key is isolated, so a name the single form would refuse with 422 — an unknown vendor included — comes back as an `off-basket` entry instead of failing the batch/v1/oracle/manifest/{metadataHash}Content-addressed manifest of a confirmed revision — verify `keccak256(JCS(manifest))` against the on-chain metadataHashCode Examples
All endpoints are public. No authentication required.
Get Current SCU
curl https://api.compute.finance/v1/oracle/scuList All Models
curl https://api.compute.finance/v1/oracle/modelsGet Single Model
curl https://api.compute.finance/v1/oracle/models/claude-opus-5Historical SCU (date range)
curl "https://api.compute.finance/v1/oracle/history?from=2026-05-18T00:00:00Z&to=2026-06-18T00:00:00Z&granularity=daily"Response Schemas
GET /v1/oracle/scu
{
"scuUsd": 0.003301944577548741,
"computeIndex": 73.73384151088473,
"referenceWorkload": { "inputTokens": 1000, "outputTokens": 500 },
"methodologyVersion": 1,
"breakdown": {
"methodologyVersion": 1,
"familyRepresentatives": [
{
"family": "openai.gpt",
"modelKey": "gpt-5.5",
"inputPriceUsdPerMillion": 5.0,
"outputPriceUsdPerMillion": 30.0,
"blendedCostUsd": 0.02
}
]
},
"updatedAt": "2026-08-27T15:35:16.947Z"
}GET /v1/oracle/models
{
"models": [
{
"id": "openai/gpt-5.5",
"displayName": "GPT-5.5",
"provider": { "key": "openai", "name": "OpenAI" },
"family": "openai.gpt",
"weiPricePerMillion": { "input": 0.06781146800693592, "output": 0.4068688080416155 },
"usdPricePerMillion": { "input": 5, "output": 30 },
"markedUpWeiPricePerMillion": { "input": 0.07120204140728272, "output": 0.42721224844369626 },
"markedUpUsdPricePerMillion": { "input": 5.25, "output": 31.5 },
"releasedAt": "2026-04-23T00:00:00.000Z",
"cache": {
"cachedInput": { "usdPerMillion": 0.5, "ratioOfInput": 0.1, "source": "catalog", "sourceUrl": null, "provenance": "verified", "createdAt": "2026-08-26T15:16:10.984Z" },
"cacheWrite5m": null,
"cacheWrite1h": null,
"read_multiplier": 0.1,
"write_multiplier_5m": null,
"write_multiplier_1h": null
},
"reasoning": {
"reasoningOutput": { "usdPerMillion": 30, "ratioOfInput": 6, "source": "catalog", "sourceUrl": null, "provenance": "verified", "createdAt": "2026-08-26T15:16:10.984Z" }
},
"inBasket": true,
"retiredAtRevision": null
}
]
}Error catalog
Every API response uses the envelope below on failure. Branch on error.code rather than HTTP status or error.message text — the code taxonomy is stable, message copy may change.
{
"error": {
"message": "The Bearer API key is missing, malformed, or not recognized.",
"type": "invalid_request_error",
"code": "invalid_api_key",
"param": "authorization", // optional, set when the error binds to a field
"details": { ... }, // optional, code-specific structured payload
"issues": [ ... ] // present when one or more fields fail validation
}
}type is derived from code, with one exception that separates the two 422s: a body that fails schema validation answers validation_error and lists every offending field in issues[], while a body that parsed and was then refused further in — a parameter the target provider refuses, a provider-capability mismatch, an input above the model's window — answers invalid_request_error and carries no issues[]. Both carry code: validation_failed, so type is what tells them apart.
| Code | Status | Type | Meaning | When it occurs |
|---|---|---|---|---|
unauthorized | 401 | invalid_request_error | Authentication is required to access this endpoint. | The session cookie is missing or invalid, the Bearer token is absent, or the timestamp on a signed request is outside the 5-minute freshness window. |
forbidden | 403 | forbidden | The caller is authenticated but is not allowed to perform this action. | The caller lacks the required role, or a request was signed by an address different from the authenticated wallet. |
org_role_required | 403 | forbidden | The caller's role in the active organization is below the one this action requires. | The caller attempted an action their role in the active organization does not reach — managing keys, limits or pools and connecting apps that act on the account need ADMIN or OWNER; inviting a MEMBER, resending or revoking a MEMBER's invitation and removing a MEMBER need ADMIN or OWNER, while doing the same for an ADMIN needs OWNER, so an ADMIN hands out only the MEMBER role; deleting the organization needs OWNER; `details.minimumRole` names the role required and `details.actualRole` the one held. |
invalid_signature | 401 | invalid_request_error | The provided signature could not be verified. | The EIP-191 signature is malformed, truncated, or does not recover to a valid signer. |
signature_reused | 409 | invalid_request_error | This signed request has already been submitted. | The nonce for this signed request was consumed by an earlier submission. |
invalid_api_key | 401 | invalid_request_error | The Bearer API key is missing, malformed, or not recognized. | The token is empty, does not use the `ct_live_` format, or matches no key the exchange issued. A key that exists but is revoked or frozen answers with its own code instead. |
api_key_frozen | 403 | forbidden | The API key is currently frozen and cannot be used for requests. | The key was frozen from the API keys page or by support. Reactivate it from the API keys page, or contact support if it was frozen by an administrator. |
api_key_revoked | 401 | invalid_request_error | The API key has been revoked and can no longer be used. | The key was revoked from the API keys page. Revocation is permanent — issue a new key to continue. |
invalid_app_grant | 401 | invalid_request_error | The Bearer app grant token is missing, malformed, or not recognized. | The token is empty, does not use the `cfa_live_` format, or matches no grant the exchange issued. A grant that exists but is revoked or frozen answers with its own code instead. |
app_grant_frozen | 403 | forbidden | The app grant is currently frozen and cannot be used for requests. | The grant was frozen from the connected apps section of Settings. Lift the freeze there; meanwhile the account's own session and every other grant keep working. |
app_grant_revoked | 401 | invalid_request_error | The app grant has been revoked and can no longer be used. | The grant was revoked from the connected apps section of Settings. Revocation is permanent — the app has to be granted access again, which issues a new token. |
model_restricted | 403 | forbidden | The requested model is not in this key's allowed-models list. | The key was restricted to a specific set of models; the request referenced a model outside that set. |
bad_request | 400 | invalid_request_error | The request is invalid in a way that does not match a more specific error code. | A generic 400 for request-shape issues that no other code describes more precisely. |
validation_failed | 422 | invalid_request_error | The request failed validation. A schema failure lists every issue in `error.issues[]`; a single-field failure names that field in `error.param`. | One or more fields — of the body, the query string, or the path — are missing, of the wrong type, or outside their allowed range. |
method_not_allowed | 405 | invalid_request_error | The endpoint exists but does not accept this HTTP method. | A request used a method the route does not support; the `Allow` response header lists the accepted methods. |
payload_too_large | 413 | invalid_request_error | The request body is larger than the exchange accepts. | The body exceeded the fixed body ceiling, which covers the largest context the exchange serves; it is refused before it is parsed, so no field-level detail is available. |
not_found | 404 | not_found | The requested resource does not exist. | The identifier did not match any resource, or the URL does not match any endpoint. |
already_exists | 409 | conflict | A resource with the same unique identity already exists. | A duplicate was detected during pre-check, or a concurrent write violated a unique constraint. |
conflict | 409 | conflict | The resource was modified concurrently; retry with the latest version. | A serializable transaction aborted due to concurrent writes, or an operation found the resource in an inconsistent state. |
invitation_expired | 409 | conflict | The organization invitation is past its expiry and can no longer be accepted. | The invitation link was opened after its expiry; a member who may hand out the invited role — ADMIN or OWNER for a MEMBER invitation, OWNER for an ADMIN one — has to send it again, which issues a new link and invalidates the old one. |
invitation_revoked | 409 | conflict | The organization invitation was withdrawn before it was accepted. | A member who may hand out the invited role — ADMIN or OWNER for a MEMBER invitation, OWNER for an ADMIN one — revoked the invitation; the link stays readable but can never be accepted. |
invitation_already_accepted | 409 | conflict | The organization invitation has already been accepted. | The invitation was accepted earlier, or two accepts raced and this one lost — an invitation admits exactly one member. |
insufficient_balance | 402 | insufficient_quota | The account's $COMPUTE balance is below the requested amount. | Triggered when withdrawing more than the balance, or when an inference request would exceed the balance after billing. |
spending_limit_reached | 429 | rate_limit_error | A spending cap has been reached. | This request would exceed a daily or monthly cap on the key, or a daily, weekly or monthly cap on the account; `details.scope` names which carried it. |
all_keys_exhausted | 502 | server_error | No upstream capacity is currently available for the routed model. | All contributor keys serving this model were unavailable, rate-limited, or in a cool-down; retry shortly. |
rate_limited | 429 | rate_limit_error | The caller exceeded a rate-limit window on this endpoint. | Either the per-endpoint request-rate cap or the provider-pool RPM/TPM cap was hit. |
stream_interrupted | 500 | server_error | The SSE stream aborted before completion. | The upstream provider connection dropped mid-stream. This error is delivered inside the SSE body, not as an HTTP status. |
contract_error | 500 | server_error | An on-chain call reverted or could not be confirmed. | The transaction failed to broadcast, timed out waiting for a receipt, or the receipt reported failure. Where the layer that refused is known, `details.chainFailure.reason` names it: `contributor_frozen` (the escrow refuses this contributor until an operator lifts the freeze), `contract_rejected` (the contract refused for another reason), `relayer_unfunded` (the exchange could not pay for the transaction, so retrying will not help), `chain_unavailable` (no verdict was reached, so retrying is meaningful) or `confirmation_pending` (the transaction was broadcast and its outcome is still unknown, so resubmitting risks a second charge). |
internal_error | 500 | server_error | An unexpected server-side failure occurred. | A fallback for exceptions that no other error code describes; the failure is logged for investigation. |
service_unavailable | 503 | server_error | A required upstream dependency is temporarily unavailable. | A dependency needed for this request is unreachable; security-critical paths intentionally reject rather than degrade. |
Inference API
An OpenAI- and Anthropic-compatible inference API. Point an official openai or @anthropic-ai/sdk client at Compute Finance by swapping the base URL and providing a ct_live_* key. Requests are billed from your $COMPUTE balance.
Base URL: https://api.compute.finance. Swagger UI: /v1/docs/inference · Get a key: https://compute.finance/dashboard/api-keys.
First request
Two things have to exist before any snippet works: an account and a balance on it. Signing in by email or by wallet creates the account — see Compute Finance ID. Its $COMPUTE balance is topped up at https://compute.finance/dashboard/ai-balance, by card or by swapping USDC from a connected wallet.
From there to an answer: issue a key at https://compute.finance/dashboard/api-keys, then run the cURL or the Python snippet with your key in place of ct_live_.... cURL needs nothing installed; the Python one is the official openai client with its base URL pointed here and everything else unchanged. The snippets name a model as vendor/model; how that name is formed, what auto does instead, and what the answer costs are under Integration contracts.
curl https://api.compute.finance/v1/chat/completions \
-H "Authorization: Bearer ct_live_..." \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-5.5","messages":[{"role":"user","content":"Hello"}]}'from openai import OpenAI
client = OpenAI(
base_url="https://api.compute.finance/v1",
api_key="ct_live_...",
)
resp = client.chat.completions.create(
model="openai/gpt-5.5",
messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)API keys
Keys are issued from the API keys page of your account. Issuance is free and never checks your balance, and every field on the form is optional — a key created with nothing filled in works. Its label and its spending caps stay editable afterwards.
The raw key is returned exactly once, at creation. Only its SHA-256 hash is stored, so nothing can show it again — copy it before you leave the page. From then on the list identifies each key by a short prefix: enough to tell two keys apart, useless as a credential. A lost key is replaced, not recovered — revoke it and issue another.
Freezing a key is reversible from the same page, as often as you like. Revoking is permanent — a revoked key can be neither restored nor edited, and it stays in the list only as a record. A state change binds from the next request onward; a request already streaming runs to its end. What each state returns is under Authentication.
A key can be held to a subset of the catalogue: name the models under allowedModels when you issue the key at https://compute.finance/dashboard/api-keys, and change the list later from the same page. The list replaces the previous one whole, and an empty list lifts the restriction. Each entry names a catalogue model by its vendor/model id or by the vendor's own SDK id and is stored as the vendor/model id; a spelling the catalogue does not hold is refused with 422 validation_failed naming it, and at most 50 entries survive once duplicate spellings collapse. The catalogue is a superset of GET /v1/models, so a model can be allowed before it is listed there. A request naming a model outside the list returns 403 model_restricted — it is never routed to a model the key may reach instead. Leave the list empty and the key reaches every model, including auto.
Balance is held on the account, not on the key: every key of an account spends the same $COMPUTE balance, and a key's daily and monthly caps only bound that key's share of it — exceeding one returns 429 spending_limit_reached. Because issuance never checks the balance, a new key on an unfunded account authenticates and then fails its first request with 402 insufficient_balance — top the account up at https://compute.finance/dashboard/ai-balance first.
Authentication
Every endpoint accepts Authorization: Bearer ct_live_*. POST /v1/messages also accepts x-api-key: ct_live_* for @anthropic-ai/sdk compatibility.
Every request checks the key's state. A key marked active proceeds normally; a frozen key returns 403 api_key_frozen; a revoked key returns 401 api_key_revoked; anything unknown or malformed returns 401 invalid_api_key. If a key is restricted to a specific set of models and the request references another, the response is 403 model_restricted.
Endpoints
/v1/chat/completionsBearer ct_live_*OpenAI Chat Completions wire format, streaming or non-streaming./v1/messagesBearer / x-api-keyAnthropic Messages API wire format, streaming or non-streaming./v1/modelsRoutable model catalog in OpenAI format./v1/models/availabilityBearer ct_live_* (optional)Per-model routable flag over the capacity the caller can reach, plus auto — the model the routing directive resolves to from the operator's preference list right now, or null. Advisory only: true as of computedAt, never a guarantee for a later request./v1/usageBearer ct_live_*The account's balance and cumulative spend, the calling key's own daily and monthly limits and usage, the remaining budget under the tighter of the key's and the account's caps, and the account's cumulative cache reads and writes./v1/inference/estimateReturns an up-front cost preview: resolvedModel, creditsCost, creditsWei, usdCost, reservedWei — the hold the charge is taken from — and tierFromInputTokens, the context tier the quote was priced at, null on a flat-priced model. model defaults to auto, quoted at the model the exchange chooses for the request you describe — the same per-request choice the send path runs. Describe it with promptTokens, reasoningRequested and toolsPresent (true or false, nothing else parses) and lean it with routing_bias; a bias beside a named model is 422, because there is no choice to lean. A named model asked for reasoning it does not price or tools its provider does not take is refused with 422, exactly as the send answers the same request. X-Model-Chooser names what decided: engine, list or degraded. Two things the quote cannot see still move the send — a conversation already pinned to a model is served that model, and a share of auto traffic is held on the operator's order as a control arm — and the hold covers the costliest candidate, so a different model answering never costs more than was held. Name a model to quote exactly what will run. 503 when none has serving capacity or the stored operator route list cannot be read. The quote is advisory: it prices the token counts you pass, and the bill is always the counts the provider reports.SDKs
The same key drives the OpenAI client in TypeScript and the official anthropic clients against POST /v1/messages. Each label names the SDK version its snippet is pinned to.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.compute.finance/v1",
apiKey: process.env.COMPUTE_FINANCE_API_KEY!, // ct_live_...
});
const resp = await client.chat.completions.create({
model: "openai/gpt-5.5",
messages: [{ role: "user", content: "Hello" }],
});
console.log(resp.choices[0].message.content);from anthropic import Anthropic
client = Anthropic(
base_url="https://api.compute.finance",
api_key="ct_live_...",
)
resp = client.messages.create(
model="anthropic/claude-opus-4.8",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello"}],
)
print(resp.content[0].text)import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
baseURL: "https://api.compute.finance",
apiKey: process.env.COMPUTE_FINANCE_API_KEY!, // ct_live_...
});
const resp = await client.messages.create({
model: "anthropic/claude-opus-4.8",
max_tokens: 1024,
messages: [{ role: "user", content: "Hello" }],
});
console.log(resp.content[0].text);Wire-format deltas
Requests are validated before dispatch. A malformed param returns 422 validation_failed with the offending field in the error body — no tokens are billed. A param the serving provider cannot honour is either refused the same way or ignored under the published policy below.
Provider capability gating. Advanced features (tools, structured outputs) only work on providers that support them natively — currently openai and anthropic. A single named model on another provider returns 422 validation_failed; inside a models[] list, such an entry is skipped and the request fails only if no listed entry supports the feature.
Every parameter has one fate. It takes effect, it returns 422 validation_failed naming it, or it is ignored under this policy — never a silent no-op. Anthropic-served models expose no equivalent for frequency_penalty, presence_penalty, seed, logit_bias, logprobs and top_logprobs, so those are dropped and the request is served; on every other provider they are forwarded and the vendor decides. tool_choice and parallel_tool_calls sent without a non-empty tools list govern nothing and are dropped along with it. Current Claude models accept temperature only at 1.0 and top_p only at 0.99 or above, so any other value is refused naming the accepted one rather than passed on to fail at the vendor.
End-user identifiers. user — and metadata.user_id on the Messages endpoint — reaches every provider as a stable, irreversible fingerprint computed before the request leaves the exchange, or is dropped entirely where the operator has configured no fingerprint key. A vendor can attribute repeat abuse to one end user instead of banning the whole contributor key, and the identifier you send reaches no third party.
OpenAI Chat Completions
Supported params: model, messages, stream, response_format, tools, tool_choice, parallel_tool_calls, reasoning_effort, reasoning, plus the standard sampling / limit / logprobs family.
Also accepted: models (ordered fallback list, 1–8 entries), conversation_id, routing_bias — cheaper, balanced or better, which leans an auto choice and is refused anywhere else.
Not supported (422): legacy functions / function_call (use tools / tool_choice), n > 1, non-text content parts.
Anthropic Messages
Supported params: model, messages, max_tokens, system, stream, tools, tool_choice, thinking, plus stop_sequences, metadata.user_id, and the sampling family (temperature, top_p).
Content blocks: text, tool_use, tool_result; a cache_control marker on a text block, a system block, a text block inside a tool_result, or a tools[] entry reaches Anthropic in its native form. Also accepted: models (ordered fallback list) and routing_bias. Extra fields on message objects are rejected with 422.
Prompt caching
What you may send. On /v1/chat/completions a message's content is a string or an array of up to 50 text parts ({"type":"text","text":"…"}); a part of any other type — image_url and the rest — is refused with 422 validation_failed naming the offending part type. On /v1/messages a message's content is a string or up to 200 blocks of type text, tool_use or tool_result, system is a string or up to 20 text blocks, and a tool_result holds up to 50 text blocks; any other block type is refused the same way, naming it. Every text part counts toward the input estimate that sizes the reservation and toward the bill.
Where the breakpoint goes. A text part may carry cache_control — {"type":"ephemeral"}, optionally with a ttl of 5m (the default) or 1h; any other value is refused before the request is priced. The marker is accepted on a message text part, on a system text block, on a text block inside a tool_result, and on a tools[] entry. It is refused on a part with no text — a breakpoint must ride text — and on the tool_result block itself rather than on a text block inside it. How many breakpoints one request may carry is capped by the provider; above that cap the provider refuses the request, and its own explanation comes back in the error message.
{
"model": "anthropic/claude-opus-4.8",
"messages": [
{
"role": "system",
"content": [
{
"type": "text",
"text": "You are a contract analyst. Answer only from the contract below.",
"cache_control": { "type": "ephemeral", "ttl": "1h" }
}
]
},
{
"role": "user",
"content": [
{ "type": "text", "text": "<the full contract text>", "cache_control": { "type": "ephemeral" } },
{ "type": "text", "text": "Which termination clauses changed?" }
]
}
]
}A provider that cannot express one. The marker reaches Anthropic-served models in Anthropic's own block form. On every other provider the text parts are joined into one string and the marker is removed before the call — the request is served normally, with no error and no warning. Those vendors may still cache implicitly: where a vendor reports the hit in the usage shape the exchange reads for that vendor it is billed at the model's cache-read price, and where a vendor reports a hit in any other shape, those tokens are billed as fresh input.
What it costs. A cache write is billed at the model's cache-write price for the ttl requested — 5m and 1h are priced separately — and a later hit at its cache-read price. Both are components of the same published price list as input and output, under the same routing markup. Cache usage comes back in the wire format you called: usage.prompt_tokens_details.cached_tokens on /v1/chat/completions, usage.cache_read_input_tokens with writes under usage.cache_creation on /v1/messages. GET /v1/inference/estimate is not cache-aware: a hit is not knowable before the request runs, so the quote prices the whole input as fresh. GET /v1/usage reports the same tokens cumulatively for the account behind the key — cachedInputTokens for reads, cacheWrite5mTokens and cacheWrite1hTokens for writes at each lifetime. Each is null until at least one settled request on the account recorded a non-zero count of that kind, so an account that has never used caching carries null for all three, and the counters are incremented from the same breakdown the receipt stores: exactly the token counts the charges were computed from. Both write lifetimes collapse into usage.prompt_tokens_details.cache_creation_tokens on /v1/chat/completions, so the two cumulative write counters reconcile per request only on /v1/messages, which splits them.
Reasoning
Two controls switch reasoning on and map onto the same provider wiring: reasoning_effort from the OpenAI wire format, and the structured reasoning object of the OpenRouter convention. An OpenRouter-native client works unchanged.
| Field | Meaning |
|---|---|
reasoning.effort | Reasoning depth: low, medium or high. Equivalent to reasoning_effort. |
reasoning.max_tokens | Explicit thinking budget, 1024–200000 tokens. Mutually exclusive with effort. |
reasoning.enabled | true without a budget uses effort medium; false switches reasoning off. |
reasoning.exclude | Think, but keep the reasoning out of the response. |
Any reasoning object without enabled: false switches reasoning on: a bare reasoning: {} and reasoning: {"exclude": true} both run at effort medium and are billed for the thinking tokens they generate.
Contradictory values across the two controls are 422 validation_failed naming both fields — never a silent priority. That covers two different effort levels, an effort together with reasoning.max_tokens, and enabled: false alongside any requested budget. Agreeing values are accepted, and exclude combines freely with either control.
The effort level becomes a thinking budget as a share of max_tokens (default 4096), floored at 1024 tokens — Anthropic's minimum, applied on every provider. The budget must stay below max_tokens; a request that leaves no room for an answer returns 422 validation_failed naming both numbers, before it reaches a provider and before any credits move.
| Effort | Thinking budget |
|---|---|
low | 20% of max_tokens |
medium | 50% of max_tokens |
high | 80% of max_tokens |
On OpenAI-shaped providers the effort level is forwarded as reasoning_effort; an explicit reasoning.max_tokens maps onto the closest effort bucket in the same table.
Reasoning is available exactly on models whose reasoning output is priced in the catalogue. A single named model without it returns 422 validation_failed; inside a models[] list, such an entry is skipped and the request fails only if no listed entry can reason. To check a model before you send it: GET /v1/oracle/models — and GET /v1/oracle/catalog for models outside the basket — returns a reasoning.reasoningOutput price block for exactly the models that can reason, and null for the rest.
The assistant message carries a reasoning field alongside content; streaming chunks carry delta.reasoning before the content deltas. The name is the same on every provider — a model whose own API returns its thinking under reasoning_content is renamed before the response leaves the exchange, streaming and non-streaming alike, and exclude withholds it just the same.
{
"choices": [{
"index": 0,
"message": {
"role": "assistant",
"content": "42",
"reasoning": "The question asks for the answer to life, the universe and everything..."
},
"finish_reason": "stop"
}],
"usage": { "prompt_tokens": 18, "completion_tokens": 512, "total_tokens": 530 }
}data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":...,"model":"anthropic/claude-opus-4.8","choices":[{"index":0,"delta":{"reasoning":"The question asks"},"finish_reason":null}]}
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":...,"model":"anthropic/claude-opus-4.8","choices":[{"index":0,"delta":{"reasoning":" for the answer..."},"finish_reason":null}]}
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":...,"model":"anthropic/claude-opus-4.8","choices":[{"index":0,"delta":{"content":"42"},"finish_reason":null}]}With reasoning.exclude: true the model still reasons and the reasoning tokens are still generated and billed — only the output is withheld. Hiding the reasoning does not make it free.
On /v1/messages the native thinking parameter sets the same budget as reasoning.max_tokens: {"type":"enabled","budget_tokens":N} obeys the same floor and the same below-max_tokens rule, while {"type":"disabled"} switches reasoning off and no thinking parameter is forwarded at all. If the request falls through to an OpenAI-shaped entry, the budget maps onto the closest effort bucket. Reasoning comes back as thinking content blocks ahead of the text block, carrying the signature Anthropic issued when an Anthropic model served the request and an empty signature when another provider did; streaming emits thinking_delta events on that block, plus signature_delta when a signature was issued.
Message content blocks accept text, tool_use and tool_result; a thinking block in an assistant turn is rejected with 422 validation_failed. Strip thinking blocks from an assistant turn before you replay it.
Reasoning tokens are output tokens: they are counted in completion_tokens and billed at the model's output rate. There is no separate reasoning charge.
Streaming
Both endpoints stream Server-Sent Events (Accept: text/event-stream) when stream: true. Wire formats differ (OpenAI vs Anthropic).
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":...,"model":"openai/gpt-5.5","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":...,"model":"openai/gpt-5.5","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}
...
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":...,"model":"openai/gpt-5.5","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":10,"completion_tokens":5,"total_tokens":15},"x_credits_used":"1234","x_credits_remaining":"98765","x_ratelimit_requests_remaining":123}
data: [DONE]event: message_start
data: {"type":"message_start","message":{"id":"msg_...","type":"message","role":"assistant","content":[],"model":"anthropic/claude-opus-4.8","stop_reason":null,"stop_sequence":null,"usage":{"input_tokens":10,"output_tokens":0}}}
event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}
event: content_block_stop
data: {"type":"content_block_stop","index":0}
event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"output_tokens":5}}
event: message_stop
data: {"type":"message_stop"}Stop reasons are translated between wire formats: OpenAI stop ↔ Anthropic end_turn; length ↔ max_tokens; tool_calls ↔ tool_use.
The streaming final chunk also mirrors x_credits_used, x_credits_remaining, and x_ratelimit_requests_remaining into the body.
Response headers
| Header | Meaning |
|---|---|
X-Model-Used | Model that actually answered this request, in `vendor/model` form. |
X-Model-Requested | Model the exchange set out to serve, in `vendor/model` form — the one the request named, or the head of the chain `auto` resolved to. |
X-Fallback-Used | `true` exactly when the model that answered is not the one `X-Model-Requested` names. |
X-Model-Chooser | On every `auto` answer, what picked the model: `engine` — a choice made for this request; `pin` — the model this conversation is already on; `list` — the operator's preference order; `degraded` — the per-request choice could not be made, so the preference order served. Absent when the request named a model. |
X-Provider | Vendor slug of the provider that served the request. |
X-Routing-Strategy | How the provider key was chosen: `sticky`, `weighted-random`, or `failover` after a retry or a fallback entry took over. |
X-Request-Id | Identifier of this routing decision — quote it in a support request. |
X-Compute-Used | Wei debited for this request after settlement. |
X-Compute-Remaining | Remaining $COMPUTE balance for the buyer account, in wei. |
X-Key-Daily-Remaining | Remaining daily spending cap for this key, in wei — `unlimited` if unset. |
X-Key-Monthly-Remaining | Remaining monthly spending cap for this key. |
X-RateLimit-Limit | Ceiling of the tightest active rate-limit bucket. |
X-RateLimit-Remaining | Remaining capacity in the tightest bucket. |
X-RateLimit-Reset | Epoch seconds when the tightest bucket resets. |
Errors
Every error response follows the envelope below. Branch on error.code rather than HTTP status — several codes can share the same status.
{
"error": {
"message": "The Bearer API key is missing, malformed, or not recognized.",
"type": "invalid_request_error",
"code": "invalid_api_key",
"param": "authorization", // optional, set when the error binds to a field
"details": { ... }, // optional, code-specific structured payload
"issues": [ ... ] // present when one or more fields fail validation
}
}type is derived from code, with one exception that separates the two 422s: a body that fails schema validation answers validation_error and lists every offending field in issues[], while a body that parsed and was then refused further in — a parameter the target provider refuses, a provider-capability mismatch, an input above the model's window — answers invalid_request_error and carries no issues[]. Both carry code: validation_failed, so type is what tells them apart.
| Code | Status | Type | Meaning | When it occurs |
|---|---|---|---|---|
invalid_api_key | 401 | invalid_request_error | The Bearer API key is missing, malformed, or not recognized. | The token is empty, does not use the `ct_live_` format, or matches no key the exchange issued. A key that exists but is revoked or frozen answers with its own code instead. |
api_key_frozen | 403 | forbidden | The API key is currently frozen and cannot be used for requests. | The key was frozen from the API keys page or by support. Reactivate it from the API keys page, or contact support if it was frozen by an administrator. |
api_key_revoked | 401 | invalid_request_error | The API key has been revoked and can no longer be used. | The key was revoked from the API keys page. Revocation is permanent — issue a new key to continue. |
model_restricted | 403 | forbidden | The requested model is not in this key's allowed-models list. | The key was restricted to a specific set of models; the request referenced a model outside that set. |
validation_failed | 422 | invalid_request_error | The request failed validation. A schema failure lists every issue in `error.issues[]`; a single-field failure names that field in `error.param`. | One or more fields — of the body, the query string, or the path — are missing, of the wrong type, or outside their allowed range. |
method_not_allowed | 405 | invalid_request_error | The endpoint exists but does not accept this HTTP method. | A request used a method the route does not support; the `Allow` response header lists the accepted methods. |
payload_too_large | 413 | invalid_request_error | The request body is larger than the exchange accepts. | The body exceeded the fixed body ceiling, which covers the largest context the exchange serves; it is refused before it is parsed, so no field-level detail is available. |
insufficient_balance | 402 | insufficient_quota | The account's $COMPUTE balance is below the requested amount. | Triggered when withdrawing more than the balance, or when an inference request would exceed the balance after billing. |
spending_limit_reached | 429 | rate_limit_error | A spending cap has been reached. | This request would exceed a daily or monthly cap on the key, or a daily, weekly or monthly cap on the account; `details.scope` names which carried it. |
all_keys_exhausted | 502 | server_error | No upstream capacity is currently available for the routed model. | All contributor keys serving this model were unavailable, rate-limited, or in a cool-down; retry shortly. |
rate_limited | 429 | rate_limit_error | The caller exceeded a rate-limit window on this endpoint. | Either the per-endpoint request-rate cap or the provider-pool RPM/TPM cap was hit. |
stream_interrupted | 500 | server_error | The SSE stream aborted before completion. | The upstream provider connection dropped mid-stream. This error is delivered inside the SSE body, not as an HTTP status. |
service_unavailable | 503 | server_error | A required upstream dependency is temporarily unavailable. | A dependency needed for this request is unreachable; security-critical paths intentionally reject rather than degrade. |
Rate limits
Per-key. Each API key can carry optional daily and monthly $COMPUTE caps, and the account behind it carries daily, weekly and monthly caps of its own. Exceeding any of them returns 429 spending_limit_reached; what is left of the key's caps is exposed in the X-Key-*-Remaining headers.
Per-pool. Each provider pool has RPM and TPM ceilings. Exceeding either returns 429 rate_limited.
GET /v1/inference/estimate is throttled to 10 requests per second per source IP; it requires no authentication and does not consume $COMPUTE.
Request size
Body ceiling. A request body may be up to 4194304 bytes (4 MiB) — the room the largest context the exchange serves, 1048576 input tokens, needs on the wire. A larger body is refused with 413 before it is parsed. The ceiling is fixed for the whole exchange and applies to /v1/chat/completions and /v1/messages alike.
Per-model maximum. A model may declare the context window its vendor actually serves. A request whose estimated input exceeds it fails with 422 validation_failed naming the model and its limit — before the request reaches a provider and before any $COMPUTE is reserved, so an oversized attempt costs nothing. A model that declares no maximum is bounded only by the body ceiling. GET /v1/inference/estimate runs the same check, so you can size a request before you send it. Every declared window is published as maxInputTokens on GET /v1/models and in the full catalog at https://api.compute.finance/v1/oracle/catalog, absent on a model that declares none.
Inside a fallback chain. An entry of models[] that the exchange cannot serve right now is skipped and the next candidate takes over. An input too long for an entry is not skipped: the length is a property of your request, so the request fails naming that entry rather than quietly falling through to a model you did not ask for. Only model: "auto" narrows on length — there the exchange picked the candidates, so it routes to one whose window holds the request and returns 422 only when none does.
Billing
Cost is estimated up-front, reserved from the balance, then settled against actual token usage. The X-Compute-Used and X-Compute-Remaining headers report the final debit and remaining balance. An insufficient balance returns 402; hitting a per-key or per-account cap returns 429. An account with nothing left to spend is refused before its request is estimated, so a request you could never pay for is never priced.
The estimate is advisory; the bill is not. The input estimate that sizes your reservation and picks the context tier it is held at counts your message text, your tool schema and response format, and the per-message chat framing the provider wraps every turn in. Name a model whose provider publishes its own token counter and the estimate is that vendor’s count of your request — an Anthropic model is counted by Anthropic, and no chat framing is added on top, because that count already carries it. Everything else is counted locally through the o200k_base BPE vocabulary at a pinned tokenizer version and scaled into the target model’s vocabulary by that vendor’s fixed ratio, so the same request reserves the same amount — above a small minimum floor, and outside the rare fallback where the vocabulary cannot be loaded. auto is always counted that way, since the serving model is chosen after the hold is taken, and so is a vendor counter that declines or times out; each candidate model is counted in its own vocabulary and your reservation is held at the costliest of them. It never decides what you pay: settlement charges the token counts the provider reported, at the tier those counts select. GET /v1/inference/estimate tokenizes nothing — it prices the promptTokens you hand it, so counting your prompt with the same vocabulary and adding four tokens per message plus three for the reply quotes a locally counted reservation; a named model its vendor counts holds a different amount, which only that vendor’s own count reproduces.
A model may carry a ladder of context tiers — ordered input-token thresholds, each with its own input and output price. The tier follows the whole input side of the request (prompt plus cache reads plus cache writes), and an input landing exactly on a threshold is priced at the higher tier. Settlement re-selects the tier from the tokens the provider reported and is what you pay; the reservation’s tier only sets how much was held. A cache or reasoning price the catalogue attests is single-level, but a kind it does not price separately is billed at its side’s rate, that side’s context tier included. A flat-priced model — one with no ladder — costs the same at every length. Quote a specific size with GET /v1/inference/estimate — its tierFromInputTokens names the tier your bill will use for that input.
Machine-readable specifications
OpenAPI YAML · OpenAPI JSON · llms-full.txt
Integration contracts
How a model is named and chosen, which API answers which question, and what a request costs in $COMPUTE — the rules the running system enforces, in one place. The wire-format detail behind each stays in the section it belongs to.
Model identifiers and selection
The identifier. GET /v1/models publishes the routable catalogue in OpenAI format, and every id has the form vendor/model — the identifier the Inference API routes and bills on. Pass it as model, or as an entry of the models[] fallback list. The bare model name is accepted in both places, so x-ai/grok-4.5 and grok-4.5 reach the same model; the match is case-insensitive, and an unknown vendor prefix, or a real model named under the wrong vendor, is rejected rather than resolved by guess. GET /v1/oracle/resolve/{key} answers what a spelling resolves to before you send it.
auto. Omitting model, or setting it to auto, hands the choice to the exchange: it serves a model from the operator's ordered preference list, among the entries with serving capacity — send conversation_id to group the turns of one conversation, so the exchange keeps serving them from the same capacity. When that list is empty, or nothing on it can serve, the choice falls through to the cheapest model that can rather than refusing every buyer over an operator setting; a list that can serve nothing also raises an operator alert. That value is a routing directive and never a catalogue id — it takes no vendor prefix and may not appear inside models[]. When no model has capacity the request is refused with 503 rather than served by something you did not choose, and GET /v1/models/availability names the model auto resolves to from that list right now.
routing_bias. An optional cheaper / balanced / better, defaulting to balanced, that leans an auto choice toward the cheaper or the stronger end of the operator's list. It is a bias and not a service level — it influences which model is picked and promises neither a price nor a quality. Sending it beside a named model, or beside a models[] list, is refused with 422 validation_failed: both spend the choice the dial exists to lean. It reaches that choice only where the exchange made one for the request, which X-Model-Chooser names — engine and pin carry the bias, while list and degraded mean the operator's order served and the dial did not act. Changing the value releases the model a conversation_id was being served from, so the next turn is chosen afresh. GET /v1/inference/estimate takes the same dial and runs the same choice, so a quote can be leaned exactly as the request it stands for.
Your own fallback list. Send models — an ordered list of 1 to 8 catalogue ids; model, when present, is tried first, then each entry in turn. An entry lacking capacity or a requested capability is skipped, the request is billed at the model that answered, and the hold covers the costliest listed entry with the excess released at settlement. Combining a list with model: auto is rejected. Fall-through covers capacity and capability only: a prompt above an entry's context window, or a parameter that entry rejects as invalid, aborts the request instead of moving to the next candidate.
No silent substitution. A named model is served or the request fails — the exchange never swaps it for another. Substitution happens only where you asked for it: an entry of your models[] list, or auto. Every answer says which model produced it — X-Model-Used names the model that ran, X-Model-Requested the one the exchange set out to serve, and X-Fallback-Used reads true exactly when the two differ. A named model outside a key's allowedModels is refused with 403 model_restricted rather than swapped for one the key may reach.
Parameters. What each wire format does with every parameter it is sent — take it, refuse it, or drop it under a published policy — is listed under Wire-format deltas.
Which API answers what
Two products share one host. The Inference API is paid, takes a ct_live_* key, and serves the models that can run a request. The Oracle API is free, unauthenticated and read-only, and publishes prices and the index over a catalogue that is a superset of the routable set — a model can be priced there and not be routable here.
| Question | Ask |
|---|---|
| Which models may I name? | GET /v1/models |
| Which of them can serve right now? | GET /v1/models/availability |
| What does this spelling resolve to? | GET /v1/oracle/resolve/{key} · GET /v1/oracle/resolve?keys={csv} |
| What will this request cost? | GET /v1/inference/estimate |
| What rates am I billed at? | GET /v1/oracle/pricing |
| What do the vendors charge? | GET /v1/oracle/models · GET /v1/oracle/catalog |
| What is compute worth, now and over time? | GET /v1/oracle/scu · GET /v1/oracle/history |
| What have I spent, and what is left? | GET /v1/usage |
Markup is where the two price lists differ. Every figure on GET /v1/oracle/pricing already carries the routing markup, cache and reasoning components included, so it matches what a request is charged. GET /v1/oracle/models and GET /v1/oracle/catalog publish the vendors' own list prices without it, and /v1/oracle/models and /v1/oracle/basket carry the marked-up pair alongside as markedUpUsdPricePerMillion and markedUpWeiPricePerMillion. Read a number against the endpoint it came from; the two lists are not interchangeable.
What a request costs
Tokens are split by kind, and the kind picks the rate. The catalogue prices each kind separately on the models that support it; a kind a model does not price separately is billed at its side's rate — cache components at input, reasoning output at output.
| Kind | Counts |
|---|---|
input | Prompt tokens the provider read fresh. |
output | Tokens the model generated, thinking tokens included. |
cachedInput | Prompt tokens served from a cache hit. |
cacheWrite5m | Prompt tokens written into a 5-minute cache entry. |
cacheWrite1h | Prompt tokens written into a 1-hour cache entry. |
reasoningOutput | Thinking tokens, on a model that prices them apart from ordinary output. |
The arithmetic. Each kind's token count is charged at its own rate per million tokens. On a model with a context-tier ladder, the whole input side of the request — prompt plus cache reads plus cache writes — picks the tier the input and output rates are read from. The legs are summed, the 5% routing markup is applied to that sum exactly once, and the result is divided by the $COMPUTE peg to reach the debit. Settlement re-runs the same arithmetic on the token counts the provider reported, and that is the bill; the reservation only decides how much was held.
The provenance mark a catalogue price carries plays no part in this arithmetic — Endpoints defines what it records.
# Catalogue rates for openai/gpt-5.5 — GET /v1/oracle/catalog
input $5.00 per 1M tokens
output $30.00 per 1M tokens
# Peg — GET /v1/oracle/pricing
pegUsd 73.733841 USDC per 1 $COMPUTE
input leg 1000 / 1e6 x 5.00 = $0.005
output leg 500 / 1e6 x 30.00 = $0.015
subtotal = $0.020
routing markup x 1.05 = $0.021
÷ peg / 73.733841 = 0.00028480816562913085 $COMPUTE
# GET /v1/inference/estimate?model=openai/gpt-5.5&promptTokens=1000&completionTokens=500
{
"model": "openai/gpt-5.5",
"resolvedModel": "openai/gpt-5.5",
"promptTokens": 1000,
"completionTokens": 500,
"creditsCost": 0.00028480816562913085,
"creditsWei": "284808165629131",
"usdCost": 0.020999999999999998,
"reservedWei": "284808165629131",
"tierFromInputTokens": null
}Catalogue prices and the peg both move, so the figures above are illustrative — GET /v1/inference/estimate is the quote that is always current, and its reservedWei is the hold the charge comes out of.
Model Context Protocol (MCP)
AI agents can query the oracle, estimate costs, and analyze sessions through the official Compute Finance MCP server. Stdio transport, no API key required.
Install in any MCP client: npx @compute-finance/mcp
Client config snippet: {"command":"npx","args":["@compute-finance/mcp"]}
Claude Code one-liner (registers MCP + skills + cost hook): npx @compute-finance/mcp setup
Read-only tools across five layers:
- data — Live oracle data — basket, price, SCU, CPI, reconstitutions
- compute — Cost estimation and cross-model comparison
- render — Pre-formatted session reports used by the Claude Code skills
- analyze — Raw JSON session and per-inference breakdown
- history — Aggregate stats across logged sessions
Bundled Claude Code slash skills:
/cf-session-management— Measured post-session cost analysis/cf-session-consumption— Per-inference token spend breakdown/cf-active-sessions— Multi-session overview across projects
Reference: github.com/compute-finance/mcp · /.well-known/mcp/server-card.json
Compute Finance ID
Your Compute Finance ID (CF ID) is your identity across all Compute Finance surfaces. Created on first sign-in via email or wallet, it ties together your profile, points balance, and referral code into a single record.
What CF ID stores
| Field | Description |
|---|---|
cf_id | Public identifier — format: cf_usr_XXXXXXXXXXXX |
email | Used for notifications. Optional for wallet-only sign-in. |
wallet_address | On-chain address on Base. Created via account abstraction or connected externally. |
display_name | User-chosen name. Defaults to a truncated email or wallet address. |
referral_code | Permanent 8-character code — format: cf_ref_XXXXXXXX |
points_balance | Current points total, denormalized from the points ledger |
Sign-in flows
Two sign-in methods, both produce a CF ID:
- Email — Enter your email, receive a 6-digit code, verify. A smart wallet is created on Base and associated with your email. No seed phrase required.
- Wallet — Connect MetaMask, Coinbase Wallet, or any WalletConnect-compatible wallet. Sign a SIWE message. Your wallet address becomes your CF ID's primary identifier.
Both flows converge on the same CF ID record. You can add an email to a wallet-only account later from your settings.
Organizations
An organization lets several people work on one account — its owner's. The owner's balance pays for the members' requests, and the owner's API keys, spending limits, pools and history are shared with them. Every member keeps a personal account of their own alongside.
Personal and organization context
While you are signed in, your requests run in one context, remembered on your profile. Every account page shows the active context and your role in it in the account switcher at the top of the navigation — at the top of the page on a narrow screen — and switches from there in one step; Settings → Organizations switches too:
- Personal — your own account — your balance, API keys, spending limits, pools and history.
- Organization — the owner's account. The balance, API keys, spending limits, pools and history you see are the owner's; chat you send is paid from the owner's balance, and a key issued there belongs to the owner's account. What you may change there is set by your role.
Accepting an invitation makes that organization your active context. Creating an organization makes you its owner but does not switch to it. You can own at most 10 organizations. Over the API, POST /v1/orgs/context with an orgId switches and answers the context now active, orgId: null returns to the personal context, and POST /v1/invitations/accept switches into the organization it joins.
Some things keep the context they started in, whichever one is active later: a connected app acts in the context its grant was issued for, a Spark proposal is carried out in the context it was made in, and an API key always spends the account it was issued on.
Roles
Each member holds exactly one role — Owner, Admin or Member — and every role includes everything the role below it allows. The table is generated from the same declaration the API enforces. An action your role does not allow is refused with org_role_required; the one exception is member addresses, which a Member sees masked rather than being refused.
| What the role allows | Owner | Admin | Member |
|---|---|---|---|
| Send inference requests and chat, paid from the organization's balance | Yes | Yes | Yes |
| See the organization's API keys, spending limits, pools, history and balance | Yes | Yes | Yes |
| See who belongs to the organization and the role each member holds | Yes | Yes | Yes |
| Create, edit, freeze and revoke the organization's API keys | Yes | Yes | No |
| Set the organization's spending limits | Yes | Yes | No |
| Create, edit and delete the organization's pools, and see and change who is in them | Yes | Yes | No |
| Connect apps that make changes to the organization's account | Yes | Yes | No |
| See every member's full email address | Yes | Yes | No |
| See pending invitations; invite Members, resend or revoke their invitations, and remove Members | Yes | Yes | No |
| Invite Admins, resend or revoke their invitations, and remove Admins | Yes | No | No |
| Delete the organization once every other member has been removed | Yes | No | No |
A role is never edited: to change someone's role, remove them and invite them again.
Who hands out which role
A role you can hand out is one you can take back. You invite someone to a role, resend or revoke an invitation to it, and remove a member who holds it only if your role could remove that member. The owner role is never handed out.
| Role | Invites and removes |
|---|---|
| Owner | Member, Admin |
| Admin | Member |
| Member | Nobody |
Invitations
- An invitation is addressed to one email address. Only a signed-in account whose confirmed profile email is that address can accept it, so a forwarded link admits nobody else.
- The link carries a one-time token, of which only a SHA-256 hash is stored. It expires 7 days after it is issued.
- An organization holds at most 50 pending invitations. An expired invitation takes no slot, and inviting an address whose only invitation has expired replaces it.
- Resending issues a new link and a new expiry, and the previous link stops working. Revoking withdraws an invitation for good; an accepted invitation cannot be revoked.
- An invitation is pending, expired, revoked or accepted. While it is pending, its page shows the organization, the inviter, the role, what that role allows and the expiry, and the invitation email lists the same; in any other state the page shows only the state. Reopening an accepted link while signed in as the member it admitted leads back into the organization.
What always stays personal
These act on your own account whichever context is active, so none of them reaches the owner's account from inside an organization:
- Funding — card top-ups and USDC swaps into your balance.
- Withdrawal of $COMPUTE from your balance.
- Saved cards.
- Auto top-up.
Removing a member, leaving and deleting
- Removal takes effect at once: the removed member's next request runs in their personal account, and every API key they issued on the owner's account is revoked in the same transaction.
- An Admin or a Member can leave the organization, with the same effect on the keys they issued.
- The owner can neither leave nor be removed. The owner can delete the organization instead, once every other member has been removed.
- Deleting your profile deletes an organization you own alone; while an organization you own still has other members, the profile deletion is refused.
Points
Points track your engagement with Compute Finance — signup, daily logins, oracle interactions, and referrals. As your balance grows, you unlock higher tiers. Points are append-only and recorded in a public ledger per CF ID.
Tiers
| Tier | Min Points |
|---|---|
| Explorer | 0 |
| Starter | 500 |
| Builder | 2,000 |
| Architect | 5,000 |
| Titan | 15,000 |
How points are earned
V1 supports six earning channels. New channels will be added in future versions and announced via the changelog.
| Channel | Amount | Trigger | Frequency |
|---|---|---|---|
| Signup bonus | 100 pts | CF ID created | Once per account |
| Referral (referrer) | 250 pts | Referred user completes signup | Per successful referral |
| Referral (referred user) | 50 pts | User signs up via referral link | Once per account |
| Daily login | 10 pts | User logs in on a new calendar day (UTC) | Once per day |
| 7-day login streak bonus | 50 pts | User logs in 7 consecutive days | Once per streak completion |
| Oracle interaction | 5 pts | User views a unique model’s pricing on the oracle page | Up to 12 per day (one per model) |
Ledger model
The points ledger is an append-only log. No entries are ever updated or deleted. Your canonical points balance is the sum of all your ledger entries. The denormalized points_balance field on your CF ID record is a performance optimization that is reconciled periodically.
| Field | Type | Description |
|---|---|---|
id | UUID | Primary key |
cf_id | String (FK) | The user who earned the points |
type | Enum | One of: signup, referral, referral_welcome, daily_login, streak_bonus, oracle_interaction |
amount | Integer | Points earned (always positive — append-only, no negative entries) |
source | String | Human-readable source descriptor (e.g. referral:cf_usr_a3k9m2x7p1b4, oracle:gpt-5.5) |
created_at | Timestamp | When the points were earned |
The append-only design means: no points can be silently removed or altered, the complete earning history is preserved and queryable, and any future audit can reconstruct the exact points balance at any point in time.
Streaks
Use Compute Finance on consecutive calendar days to build a streak. Longer streaks earn bonus points. Your current streak and longest streak are shown in your profile. Streaks reset if you miss a calendar day (UTC).
Leaderboard
The top 50 users by total points are displayed on the public leaderboard, refreshed periodically. The leaderboard shows display name and points only — no other profile data is exposed.
Non-transferability
Points are non-transferable in V1. They have no monetary value, are not convertible to any token or currency, and cannot be sold, traded, or assigned to another account. A points-to-credits conversion ratio for V2 will be announced before V2 ships, and the append-only ledger ensures all V1 earning history is preserved and can be converted accurately at that time.
API endpoints
Points data is queryable via the following endpoints. Authentication is required for endpoints that return personal data; the leaderboard is public.
/v1/pointsYour points summary (total, tier, current streak, longest streak)/v1/points/historyPaginated points ledger entries for your CF ID/v1/points/leaderboardTop 50 users by total points (public, no auth)Referral Program
Every CF ID includes a unique referral code (8 characters). Share your link to invite new users — both sides earn points. The program uses first-touch attribution with a 30-day cookie window.
Referral rewards
| Event | Points | Who Earns |
|---|---|---|
| New user signs up with your code | 50 pts | New user (welcome bonus) |
| You referred a new user | 250 pts | Referrer |
Sharing your referral link
// Referral URL format
https://compute.finance/r?c=cf_ref_XXXXXXXXFind your referral code in your Compute Finance ID settings. Share buttons are available for X, LinkedIn, Telegram, and copy-to-clipboard. The link uses a 30-day attribution cookie — referrals count when the referred user creates a CF ID within 30 days of clicking your link.
Attribution model
Referral attribution is first-touch with a 30-day window. The first referral link a user clicks is the one credited if they sign up within the window. Subsequent referral links from other users do not overwrite the original cookie.
How it works step by step:
- A user visits
https://compute.finance/r?c=cf_ref_XXXXXXXX - The server redirects to
https://compute.financeand sets a cookie:cf_ref=cf_ref_XXXXXXXXwithMax-Age=2592000(30 days),SameSite=Lax,Secure - The click is recorded in the database with the referral code, timestamp, hashed IP address, and user agent
- If the user signs up within 30 days — even if they navigate directly to
https://compute.financewithout the referral link — the cookie is read during CF ID creation and the referral relationship is stored - If the user has already clicked a different referral link, the first-touch cookie is preserved
Anti-gaming rules
The referral program enforces several rules to prevent farming and abuse. None of these rules surface error messages — invalid referrals are silently ignored to avoid leaking information about user accounts.
| Rule | Implementation |
|---|---|
| No self-referral | If the cf_ref cookie matches the signing-up user's own referral code, the referral is silently ignored |
| Email deduplication | One CF ID per email — a user cannot create multiple accounts with the same email to farm referral points |
| IP rate limiting | Maximum 10 CF ID creations per IP address per 24 hours, preventing mass account creation from a single source |
| Click rate limiting | Maximum 100 clicks per referral code per hour. Clicks beyond the limit are not recorded. |
| Disposable email detection | Optional: reject signups from known disposable email domains (mailinator, guerrillamail, etc.) |
Referral dashboard
Your CF ID profile includes a referral dashboard showing:
- Total referral link clicks (all-time)
- Total signups from your referral link
- Conversion rate (signups ÷ clicks)
- Total points earned from referrals
- List of referred users (display name or truncated email, signup date, status: pending or confirmed)
Status transitions
A referral has two possible states. Points are awarded when the status transitions from pending to confirmed:
- pending — The user clicked the referral link but has not yet completed signup
- confirmed — The user has created a CF ID. Points are awarded to both sides within 60 seconds of confirmation.
