COMPUTE
. FINANCE

Documentation.

What is the CPI?

The Compute Price Index ("CPI") is a public benchmark that tracks the cost of AI inference across major providers. It produces a single equal-weighted reference price — the Standard Compute Unit ("SCU") — calculated from a basket of 22 models across 9 providers.

The Problem

LLM inference pricing is fragmented. Every provider quotes pricing differently — per token, per character, per request, per model. There is no standardized way to benchmark or compare the cost of AI computation across models and providers. You need a verifiable pricing reference you can audit independently.

The Solution

The CPI is an equal-weighted basket of 22 AI models across 9 providers that produces a single, verifiable unit of account: the Standard Compute Unit (SCU). The SCU represents the USD cost of a reference workload, 1,000 input tokens + 500 output tokens, taken as the geometric mean across the full basket.

Key Properties

  • Diversified — 22 models across 9 providers prevent any single provider from dominating the index.
  • Outlier-resistant — The geometric mean dampens any single model's impact, and the one-family-one-slot rule prevents version pile-up. No caps needed; the math does the work.
  • Versioned — The equal-weighted basket launches as v1.0 on June 18, 2026. Reconstitution-driven, published on-chain when basket composition or provider rate-card prices change. There is no fixed cadence.

Basket Composition

The CPI tracks 22 models from 9 providers. Every model family gets one slot, the latest version always, and each model is weighted equally. A family is a provider's distinct product line (e.g. openai.gpt, anthropic.claude); the latest released model in each family is its representative. On a new release the representative auto-rolls. The basket updates with a new version number when families are added, removed, or replaced.

alibaba.qwen-flash

Qwen3.8 Flash

alibaba.qwen-max

Qwen3.8-Max

alibaba.qwen-plus

Qwen3.7 Plus

anthropic.claude-fable

Claude Fable 5.1

anthropic.claude-haiku

Claude Haiku 4.5

anthropic.claude-opus

Claude Opus 5.5

anthropic.claude-sonnet

Claude Sonnet 5

deepseek.v-flash

V4 Flash

deepseek.v-pro

V4 Pro

google.gemini

Gemini 2.5 Pro

google.gemini-flash

Gemini 3.8 Flash

google.gemini-flash-lite

Gemini 3.5 Flash-Lite

minimax.m

MiniMax M3

moonshot.kimi

Kimi K3

openai.gpt

GPT-6 Astra

openai.gpt-luna

GPT-6 Luna

openai.gpt-mini

GPT-5.4 Mini

openai.gpt-nano

GPT-5.4 Nano

openai.gpt-sol

GPT-6.1 Sol

openai.gpt-terra

GPT-5.6 Terra

xai.grok

Grok 4.6

xiaomi.mimo

MiMo V2.5 Pro

SCU Formula

The Standard Compute Unit is calculated in two steps. The reference workload is fixed at 1000 input + 500 output tokens, evaluated across all basket models.

The SCU is reconstitution-driven. It is published on-chain whenever the basket changes, either because a model is added, removed, or replaced, or because a provider updates a rate-card price. Both triggers produce a new on-chain version. There is no fixed cadence and no continuous off-chain feed.

Step 1 — Per-model cost

For each model in the basket, compute the cost of the reference workload using the provider's published per-token pricing. Prices are USD per 1M tokens.

cost_m = (input_price_m × 0.001) + (output_price_m × 0.0005)

Step 2 — Equal-Weight Geometric Mean

Take the geometric mean of all N model workload costs. Every model is weighted equally; there are no tiers, no caps, and no provider weights. Each model family holds one slot at its latest version, so a provider cannot inflate its representation by shipping extra SKUs.

SCU = (∏ cost_m)^(1/N) for all m in the index, N = 22

Live SCU: $0.003214 — the geometric mean of 22 equal-weighted model costs.

Every published revision permanently carries the methodology version that produced it (currently v1). Definitions and the changelog are served by the methodology endpoint; every oracle response carries the X-Methodology-Version header.

Anyone with access to provider pricing pages can independently reproduce this number.

Read the full methodology specification →

Governance

The model-family rule reduces governance to a minimum. One eligibility rule, one family rule, one emergency trigger. No weight committees, no tier reviews.

RuleSpecificationFrequency
Listing CriteriaPublic GA + first-party USD pricing + in-scope text model + seasoning + liveness. Applied identically to all.Continuous
Model Family RuleOne slot per family, latest GA version auto-wins. No governance needed.Auto
Edge-Case ReviewPublished, dated decision for family-vs-version edge cases. Stated reasoning.As Needed
ReconstitutionAdditions and removals applied by the criteria. Logged.Scheduled
Price UpdatesReflected automatically from first-party price cards. Daily scan.Daily
Emergency TriggerAny single model that moves price by more than the set threshold in 24 hours enters a short hold before the change is reflected.As Needed

Every basket change increments the revisionVersion counter. The full history is available via the GET /v1/oracle/reconstitutions endpoint. On-chain, the OracleRegistry contract records each published revision with its timestamp and metadata hash.

On-Chain Verification

All CPI pricing is recorded on-chain via the OracleRegistry contract on Base. You can independently verify prices without trusting the Compute Finance API.

Contract Address

NetworkBase (Chain ID 8453)
Contract0x1b91c0961928a14a2eD6c1985bC11aF1b302714D

Read Functions

FunctionDescription
version()Oracle implementation version — assert before deserializing tuples
decimals()Scale of scuUsd and baseline values (18)
lastRevisionVersion()Returns the highest confirmed revision number
getLatestRevision()Returns the latest revision header: revisionVersion, methodologyVersion, scuUsd, contentHash, metadataHash, publishedAt
getRevision(uint256 revisionVersion)Returns the same revision header tuple for a specific revisionVersion
getRevisionScuUsd(uint256 revisionVersion)Returns the SCU value in USD 18-dec for a specific revision
getBaseline()Returns the SCU value of the first revision (denominator of the inverse purchasing-power index)
getComputeIndex()Returns the latest (baseline / SCU) × 100
getRevisionAt(uint256 timestamp)Returns the revision version active at the given Unix-seconds timestamp
getComputeIndexAt(uint256 timestamp)Returns the atomic tuple (scuUsd, indexValue, publishedAt, methodologyVersion, revisionVersion) at the timestamp

Verification with ethers.js

JavaScript
import { ethers } from "ethers";

const ORACLE_REGISTRY = "0x1b91c0961928a14a2eD6c1985bC11aF1b302714D";
const ABI = [
  "function lastRevisionVersion() view returns (uint256)",
  "function getLatestRevision() view returns (tuple(uint256 revisionVersion, uint16 methodologyVersion, uint256 scuUsd, bytes32 contentHash, bytes32 metadataHash, uint64 publishedAt))",
  "function getBaseline() view returns (uint256)",
  "function getComputeIndex() view returns (uint256)"
];

const provider = new ethers.JsonRpcProvider("https://mainnet.base.org");
const oracle = new ethers.Contract(ORACLE_REGISTRY, ABI, provider);

// Read latest revision header
const latest = await oracle.getLatestRevision();
console.log("revisionVersion:", latest.revisionVersion.toString());
console.log("SCU (USD, 18-dec):", ethers.formatUnits(latest.scuUsd, 18));
console.log("metadataHash:", latest.metadataHash);

// Inverse purchasing-power index — (baseline / SCU) × 100
const computeIndex = await oracle.getComputeIndex();
console.log("Compute Index:", ethers.formatUnits(computeIndex, 18));

// Per-model prices live in the off-chain manifest fetched by metadataHash —
// resolve at /v1/oracle/manifest/{metadataHash} and verify with JCS+keccak256.

Verification via Basescan

You can also read the contract directly on Basescan. Since OracleRegistry is an upgradeable proxy, use the Read as Proxy tab so the implementation's functions are available:

  1. Go to Basescan → Contract → Read as Proxy
  2. Call getLatestRevision() — its first return value lists every registered model pricing key
  3. Read the on-chain revision header (scuUsd, methodologyVersion, contentHash, metadataHash, publishedAt) via getRevision(version). Per-model prices live in the off-chain manifest fetched by metadataHash.
  4. Prices are returned in $COMPUTE wei (18 decimals) per 1M tokens. Divide by 10^18 to get the $COMPUTE amount.

Verification via Sourcify

The same source is independently verified on Sourcify, a decentralized verification repository: View on Sourcify ↗

Public API

All oracle endpoints are public and require no authentication. Read access is free and unrestricted. Base URL: https://api.compute.finance

Endpoints

Every price the catalogue publishes carries a provenance mark — verified, inferred or promotional — on the number itself rather than on the model, so one model can be promotional at its flat rate and inferred on a cache component at the same time. A mark records whether an operator captured a vendor source for that number; it is set by hand and holds as of their last pass rather than as a live check, and it never changes what a request is billed.

GET/v1/oracle/scuCurrent SCU value, methodology version, and family-representative breakdown
GET/v1/oracle/modelsCatalog of basket models plus retired ex-members — `inBasket` flag and `retiredAtRevision`; retired models are priced from the live catalog. `usdPricePerMillion` and the cache / reasoning blocks are the vendors' list prices; `markedUpUsdPricePerMillion` is the same price with the routing markup, as billed
GET/v1/oracle/models/{key}Single catalog model (basket member or retired) by model id — the canonical `vendor/model` form, or the bare model name
GET/v1/oracle/models/{key}/price-historyPer-model input/output USD price time series — one entry per on-chain revision for basket members, one per catalog price change for every other model (per-revision, daily, or weekly granularity); Accept: text/csv returns a CSV attachment
GET/v1/oracle/catalogEvery tracked model (incl. catalog-only and pricing-only providers) with current price, integrated flag, index-member flag, the contextTiers ladder on models priced per context length, and a provenance mark on every number — verified (an operator recorded a vendor source for it), inferred (derived, no source recorded) or promotional (a discount that will end). Prices here are the vendors' own, without the routing markup.
GET/v1/oracle/models/{key}/price-at?date={iso8601}Per-model input/output USD price effective at the requested ISO-8601 timestamp; basket members resolve from the attested manifest, all other models from the live catalog — every response labels its source
GET/v1/oracle/basketFull basket composition: all models with equal weights, revision and methodology version
GET/v1/oracle/methodologyActive methodology version and full changelog
GET/v1/oracle/methodology/{version}Single methodology record (formula, family rule, reference workload, spec URL)
GET/v1/oracle/baselineFrozen SCU of the first confirmed revision — the denominator for the inverse computeIndex purchasing-power view
GET/v1/oracle/scu-at?date={iso8601}SCU value active at a given timestamp — step-function lookup of the latest confirmed revision with publishedAt ≤ date (highest revisionVersion on ties)
GET/v1/oracle/history?from={iso8601}&to={iso8601}&granularity={per-revision|daily|weekly}Historical SCU values, one entry per on-chain revision; granularity: per-revision, daily, or weekly; Accept: text/csv returns a CSV attachment
GET/v1/oracle/reconstitutionsNamed basket-change events with version, models, SCU delta
GET/v1/oracle/reconstitutions/exportExport full reconstitution history as a downloadable markdown file
GET/v1/oracle/revisions/{revision}Full OracleRevision record for a specific on-chain revision
GET/v1/oracle/latestLatest confirmed revision summary (version, timestamps, SCU, basket size)
GET/v1/oracle/healthLast revision timestamp, version count, on-chain sync status
GET/v1/oracle/contract-metadataLive OracleRegistry identity — chainId, proxy address, on-chain version(), and keccak256 of the deployed bytecode
GET/v1/oracle/statsPublic aggregate protocol statistics (TVL, users, volume)
GET/v1/oracle/activityLive activity feed of recent protocol events (paginated)
GET/v1/oracle/pricingPer-model pricing as billed — every figure already carries the routing markup, cache and reasoning components included; `/v1/oracle/models` publishes the same models at the vendors' list prices without it
GET/v1/oracle/resolve/{key}Resolve one model name to its canonical `vendor/model` id, prices and priceSource — percent-encode the separator (`openai%2Fgpt-5.5`) or pass the bare name. An unknown model is 200 with priceSource `off-basket` and null prices, never 404; an unknown vendor is 422
GET/v1/oracle/resolve?keys={csv}Resolve up to 100 comma-separated names in one call, input order preserved; every key is isolated, so a name the single form would refuse with 422 — an unknown vendor included — comes back as an `off-basket` entry instead of failing the batch
GET/v1/oracle/manifest/{metadataHash}Content-addressed manifest of a confirmed revision — verify `keccak256(JCS(manifest))` against the on-chain metadataHash

Code Examples

All endpoints are public. No authentication required.

Get Current SCU

curl https://api.compute.finance/v1/oracle/scu

List All Models

curl https://api.compute.finance/v1/oracle/models

Get Single Model

curl https://api.compute.finance/v1/oracle/models/claude-opus-5

Historical SCU (date range)

curl "https://api.compute.finance/v1/oracle/history?from=2026-05-18T00:00:00Z&to=2026-06-18T00:00:00Z&granularity=daily"

Response Schemas

GET /v1/oracle/scu

response
{
  "scuUsd": 0.003301944577548741,
  "computeIndex": 73.73384151088473,
  "referenceWorkload": { "inputTokens": 1000, "outputTokens": 500 },
  "methodologyVersion": 1,
  "breakdown": {
    "methodologyVersion": 1,
    "familyRepresentatives": [
      {
        "family": "openai.gpt",
        "modelKey": "gpt-5.5",
        "inputPriceUsdPerMillion": 5.0,
        "outputPriceUsdPerMillion": 30.0,
        "blendedCostUsd": 0.02
      }
    ]
  },
  "updatedAt": "2026-08-27T15:35:16.947Z"
}

GET /v1/oracle/models

response
{
  "models": [
    {
      "id": "openai/gpt-5.5",
      "displayName": "GPT-5.5",
      "provider": { "key": "openai", "name": "OpenAI" },
      "family": "openai.gpt",
      "weiPricePerMillion": { "input": 0.06781146800693592, "output": 0.4068688080416155 },
      "usdPricePerMillion": { "input": 5, "output": 30 },
      "markedUpWeiPricePerMillion": { "input": 0.07120204140728272, "output": 0.42721224844369626 },
      "markedUpUsdPricePerMillion": { "input": 5.25, "output": 31.5 },
      "releasedAt": "2026-04-23T00:00:00.000Z",
      "cache": {
        "cachedInput": { "usdPerMillion": 0.5, "ratioOfInput": 0.1, "source": "catalog", "sourceUrl": null, "provenance": "verified", "createdAt": "2026-08-26T15:16:10.984Z" },
        "cacheWrite5m": null,
        "cacheWrite1h": null,
        "read_multiplier": 0.1,
        "write_multiplier_5m": null,
        "write_multiplier_1h": null
      },
      "reasoning": {
        "reasoningOutput": { "usdPerMillion": 30, "ratioOfInput": 6, "source": "catalog", "sourceUrl": null, "provenance": "verified", "createdAt": "2026-08-26T15:16:10.984Z" }
      },
      "inBasket": true,
      "retiredAtRevision": null
    }
  ]
}

Error catalog

Every API response uses the envelope below on failure. Branch on error.code rather than HTTP status or error.message text — the code taxonomy is stable, message copy may change.

envelope
{
  "error": {
    "message": "The Bearer API key is missing, malformed, or not recognized.",
    "type": "invalid_request_error",
    "code": "invalid_api_key",
    "param": "authorization",                    // optional, set when the error binds to a field
    "details": { ... },                          // optional, code-specific structured payload
    "issues": [ ... ]                            // present when one or more fields fail validation
  }
}

type is derived from code, with one exception that separates the two 422s: a body that fails schema validation answers validation_error and lists every offending field in issues[], while a body that parsed and was then refused further in — a parameter the target provider refuses, a provider-capability mismatch, an input above the model's window — answers invalid_request_error and carries no issues[]. Both carry code: validation_failed, so type is what tells them apart.

CodeStatusTypeMeaningWhen it occurs
unauthorized401invalid_request_errorAuthentication is required to access this endpoint.The session cookie is missing or invalid, the Bearer token is absent, or the timestamp on a signed request is outside the 5-minute freshness window.
forbidden403forbiddenThe caller is authenticated but is not allowed to perform this action.The caller lacks the required role, or a request was signed by an address different from the authenticated wallet.
org_role_required403forbiddenThe caller's role in the active organization is below the one this action requires.The caller attempted an action their role in the active organization does not reach — managing keys, limits or pools and connecting apps that act on the account need ADMIN or OWNER; inviting a MEMBER, resending or revoking a MEMBER's invitation and removing a MEMBER need ADMIN or OWNER, while doing the same for an ADMIN needs OWNER, so an ADMIN hands out only the MEMBER role; deleting the organization needs OWNER; `details.minimumRole` names the role required and `details.actualRole` the one held.
invalid_signature401invalid_request_errorThe provided signature could not be verified.The EIP-191 signature is malformed, truncated, or does not recover to a valid signer.
signature_reused409invalid_request_errorThis signed request has already been submitted.The nonce for this signed request was consumed by an earlier submission.
invalid_api_key401invalid_request_errorThe Bearer API key is missing, malformed, or not recognized.The token is empty, does not use the `ct_live_` format, or matches no key the exchange issued. A key that exists but is revoked or frozen answers with its own code instead.
api_key_frozen403forbiddenThe API key is currently frozen and cannot be used for requests.The key was frozen from the API keys page or by support. Reactivate it from the API keys page, or contact support if it was frozen by an administrator.
api_key_revoked401invalid_request_errorThe API key has been revoked and can no longer be used.The key was revoked from the API keys page. Revocation is permanent — issue a new key to continue.
invalid_app_grant401invalid_request_errorThe Bearer app grant token is missing, malformed, or not recognized.The token is empty, does not use the `cfa_live_` format, or matches no grant the exchange issued. A grant that exists but is revoked or frozen answers with its own code instead.
app_grant_frozen403forbiddenThe app grant is currently frozen and cannot be used for requests.The grant was frozen from the connected apps section of Settings. Lift the freeze there; meanwhile the account's own session and every other grant keep working.
app_grant_revoked401invalid_request_errorThe app grant has been revoked and can no longer be used.The grant was revoked from the connected apps section of Settings. Revocation is permanent — the app has to be granted access again, which issues a new token.
model_restricted403forbiddenThe requested model is not in this key's allowed-models list.The key was restricted to a specific set of models; the request referenced a model outside that set.
bad_request400invalid_request_errorThe request is invalid in a way that does not match a more specific error code.A generic 400 for request-shape issues that no other code describes more precisely.
validation_failed422invalid_request_errorThe request failed validation. A schema failure lists every issue in `error.issues[]`; a single-field failure names that field in `error.param`.One or more fields — of the body, the query string, or the path — are missing, of the wrong type, or outside their allowed range.
method_not_allowed405invalid_request_errorThe endpoint exists but does not accept this HTTP method.A request used a method the route does not support; the `Allow` response header lists the accepted methods.
payload_too_large413invalid_request_errorThe request body is larger than the exchange accepts.The body exceeded the fixed body ceiling, which covers the largest context the exchange serves; it is refused before it is parsed, so no field-level detail is available.
not_found404not_foundThe requested resource does not exist.The identifier did not match any resource, or the URL does not match any endpoint.
already_exists409conflictA resource with the same unique identity already exists.A duplicate was detected during pre-check, or a concurrent write violated a unique constraint.
conflict409conflictThe resource was modified concurrently; retry with the latest version.A serializable transaction aborted due to concurrent writes, or an operation found the resource in an inconsistent state.
invitation_expired409conflictThe organization invitation is past its expiry and can no longer be accepted.The invitation link was opened after its expiry; a member who may hand out the invited role — ADMIN or OWNER for a MEMBER invitation, OWNER for an ADMIN one — has to send it again, which issues a new link and invalidates the old one.
invitation_revoked409conflictThe organization invitation was withdrawn before it was accepted.A member who may hand out the invited role — ADMIN or OWNER for a MEMBER invitation, OWNER for an ADMIN one — revoked the invitation; the link stays readable but can never be accepted.
invitation_already_accepted409conflictThe organization invitation has already been accepted.The invitation was accepted earlier, or two accepts raced and this one lost — an invitation admits exactly one member.
insufficient_balance402insufficient_quotaThe account's $COMPUTE balance is below the requested amount.Triggered when withdrawing more than the balance, or when an inference request would exceed the balance after billing.
spending_limit_reached429rate_limit_errorA spending cap has been reached.This request would exceed a daily or monthly cap on the key, or a daily, weekly or monthly cap on the account; `details.scope` names which carried it.
all_keys_exhausted502server_errorNo upstream capacity is currently available for the routed model.All contributor keys serving this model were unavailable, rate-limited, or in a cool-down; retry shortly.
rate_limited429rate_limit_errorThe caller exceeded a rate-limit window on this endpoint.Either the per-endpoint request-rate cap or the provider-pool RPM/TPM cap was hit.
stream_interrupted500server_errorThe SSE stream aborted before completion.The upstream provider connection dropped mid-stream. This error is delivered inside the SSE body, not as an HTTP status.
contract_error500server_errorAn on-chain call reverted or could not be confirmed.The transaction failed to broadcast, timed out waiting for a receipt, or the receipt reported failure. Where the layer that refused is known, `details.chainFailure.reason` names it: `contributor_frozen` (the escrow refuses this contributor until an operator lifts the freeze), `contract_rejected` (the contract refused for another reason), `relayer_unfunded` (the exchange could not pay for the transaction, so retrying will not help), `chain_unavailable` (no verdict was reached, so retrying is meaningful) or `confirmation_pending` (the transaction was broadcast and its outcome is still unknown, so resubmitting risks a second charge).
internal_error500server_errorAn unexpected server-side failure occurred.A fallback for exceptions that no other error code describes; the failure is logged for investigation.
service_unavailable503server_errorA required upstream dependency is temporarily unavailable.A dependency needed for this request is unreachable; security-critical paths intentionally reject rather than degrade.

Inference API

An OpenAI- and Anthropic-compatible inference API. Point an official openai or @anthropic-ai/sdk client at Compute Finance by swapping the base URL and providing a ct_live_* key. Requests are billed from your $COMPUTE balance.

Base URL: https://api.compute.finance. Swagger UI: /v1/docs/inference · Get a key: https://compute.finance/dashboard/api-keys.

First request

Two things have to exist before any snippet works: an account and a balance on it. Signing in by email or by wallet creates the account — see Compute Finance ID. Its $COMPUTE balance is topped up at https://compute.finance/dashboard/ai-balance, by card or by swapping USDC from a connected wallet.

From there to an answer: issue a key at https://compute.finance/dashboard/api-keys, then run the cURL or the Python snippet with your key in place of ct_live_.... cURL needs nothing installed; the Python one is the official openai client with its base URL pointed here and everything else unchanged. The snippets name a model as vendor/model; how that name is formed, what auto does instead, and what the answer costs are under Integration contracts.

cURL
curl https://api.compute.finance/v1/chat/completions \
  -H "Authorization: Bearer ct_live_..." \
  -H "Content-Type: application/json" \
  -d '{"model":"openai/gpt-5.5","messages":[{"role":"user","content":"Hello"}]}'
Python — openai@2.45
from openai import OpenAI

client = OpenAI(
    base_url="https://api.compute.finance/v1",
    api_key="ct_live_...",
)

resp = client.chat.completions.create(
    model="openai/gpt-5.5",
    messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)

API keys

Keys are issued from the API keys page of your account. Issuance is free and never checks your balance, and every field on the form is optional — a key created with nothing filled in works. Its label and its spending caps stay editable afterwards.

The raw key is returned exactly once, at creation. Only its SHA-256 hash is stored, so nothing can show it again — copy it before you leave the page. From then on the list identifies each key by a short prefix: enough to tell two keys apart, useless as a credential. A lost key is replaced, not recovered — revoke it and issue another.

Freezing a key is reversible from the same page, as often as you like. Revoking is permanent — a revoked key can be neither restored nor edited, and it stays in the list only as a record. A state change binds from the next request onward; a request already streaming runs to its end. What each state returns is under Authentication.

A key can be held to a subset of the catalogue: name the models under allowedModels when you issue the key at https://compute.finance/dashboard/api-keys, and change the list later from the same page. The list replaces the previous one whole, and an empty list lifts the restriction. Each entry names a catalogue model by its vendor/model id or by the vendor's own SDK id and is stored as the vendor/model id; a spelling the catalogue does not hold is refused with 422 validation_failed naming it, and at most 50 entries survive once duplicate spellings collapse. The catalogue is a superset of GET /v1/models, so a model can be allowed before it is listed there. A request naming a model outside the list returns 403 model_restricted — it is never routed to a model the key may reach instead. Leave the list empty and the key reaches every model, including auto.

Balance is held on the account, not on the key: every key of an account spends the same $COMPUTE balance, and a key's daily and monthly caps only bound that key's share of it — exceeding one returns 429 spending_limit_reached. Because issuance never checks the balance, a new key on an unfunded account authenticates and then fails its first request with 402 insufficient_balance — top the account up at https://compute.finance/dashboard/ai-balance first.

Authentication

Every endpoint accepts Authorization: Bearer ct_live_*. POST /v1/messages also accepts x-api-key: ct_live_* for @anthropic-ai/sdk compatibility.

Every request checks the key's state. A key marked active proceeds normally; a frozen key returns 403 api_key_frozen; a revoked key returns 401 api_key_revoked; anything unknown or malformed returns 401 invalid_api_key. If a key is restricted to a specific set of models and the request references another, the response is 403 model_restricted.

Endpoints

POST/v1/chat/completionsBearer ct_live_*OpenAI Chat Completions wire format, streaming or non-streaming.
POST/v1/messagesBearer / x-api-keyAnthropic Messages API wire format, streaming or non-streaming.
GET/v1/modelsRoutable model catalog in OpenAI format.
GET/v1/models/availabilityBearer ct_live_* (optional)Per-model routable flag over the capacity the caller can reach, plus auto — the model the routing directive resolves to from the operator's preference list right now, or null. Advisory only: true as of computedAt, never a guarantee for a later request.
GET/v1/usageBearer ct_live_*The account's balance and cumulative spend, the calling key's own daily and monthly limits and usage, the remaining budget under the tighter of the key's and the account's caps, and the account's cumulative cache reads and writes.
GET/v1/inference/estimateReturns an up-front cost preview: resolvedModel, creditsCost, creditsWei, usdCost, reservedWei — the hold the charge is taken from — and tierFromInputTokens, the context tier the quote was priced at, null on a flat-priced model. model defaults to auto, quoted at the model the exchange chooses for the request you describe — the same per-request choice the send path runs. Describe it with promptTokens, reasoningRequested and toolsPresent (true or false, nothing else parses) and lean it with routing_bias; a bias beside a named model is 422, because there is no choice to lean. A named model asked for reasoning it does not price or tools its provider does not take is refused with 422, exactly as the send answers the same request. X-Model-Chooser names what decided: engine, list or degraded. Two things the quote cannot see still move the send — a conversation already pinned to a model is served that model, and a share of auto traffic is held on the operator's order as a control arm — and the hold covers the costliest candidate, so a different model answering never costs more than was held. Name a model to quote exactly what will run. 503 when none has serving capacity or the stored operator route list cannot be read. The quote is advisory: it prices the token counts you pass, and the bill is always the counts the provider reports.

SDKs

The same key drives the OpenAI client in TypeScript and the official anthropic clients against POST /v1/messages. Each label names the SDK version its snippet is pinned to.

TypeScript — openai@^4.76
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.compute.finance/v1",
  apiKey: process.env.COMPUTE_FINANCE_API_KEY!, // ct_live_...
});

const resp = await client.chat.completions.create({
  model: "openai/gpt-5.5",
  messages: [{ role: "user", content: "Hello" }],
});
console.log(resp.choices[0].message.content);
Python — anthropic@0.116
from anthropic import Anthropic

client = Anthropic(
    base_url="https://api.compute.finance",
    api_key="ct_live_...",
)

resp = client.messages.create(
    model="anthropic/claude-opus-4.8",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello"}],
)
print(resp.content[0].text)
TypeScript — @anthropic-ai/sdk@^0.120.0
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  baseURL: "https://api.compute.finance",
  apiKey: process.env.COMPUTE_FINANCE_API_KEY!, // ct_live_...
});

const resp = await client.messages.create({
  model: "anthropic/claude-opus-4.8",
  max_tokens: 1024,
  messages: [{ role: "user", content: "Hello" }],
});
console.log(resp.content[0].text);

Wire-format deltas

Requests are validated before dispatch. A malformed param returns 422 validation_failed with the offending field in the error body — no tokens are billed. A param the serving provider cannot honour is either refused the same way or ignored under the published policy below.

Provider capability gating. Advanced features (tools, structured outputs) only work on providers that support them natively — currently openai and anthropic. A single named model on another provider returns 422 validation_failed; inside a models[] list, such an entry is skipped and the request fails only if no listed entry supports the feature.

Every parameter has one fate. It takes effect, it returns 422 validation_failed naming it, or it is ignored under this policy — never a silent no-op. Anthropic-served models expose no equivalent for frequency_penalty, presence_penalty, seed, logit_bias, logprobs and top_logprobs, so those are dropped and the request is served; on every other provider they are forwarded and the vendor decides. tool_choice and parallel_tool_calls sent without a non-empty tools list govern nothing and are dropped along with it. Current Claude models accept temperature only at 1.0 and top_p only at 0.99 or above, so any other value is refused naming the accepted one rather than passed on to fail at the vendor.

End-user identifiers. user — and metadata.user_id on the Messages endpoint — reaches every provider as a stable, irreversible fingerprint computed before the request leaves the exchange, or is dropped entirely where the operator has configured no fingerprint key. A vendor can attribute repeat abuse to one end user instead of banning the whole contributor key, and the identifier you send reaches no third party.

OpenAI Chat Completions

Supported params: model, messages, stream, response_format, tools, tool_choice, parallel_tool_calls, reasoning_effort, reasoning, plus the standard sampling / limit / logprobs family.

Also accepted: models (ordered fallback list, 1–8 entries), conversation_id, routing_bias — cheaper, balanced or better, which leans an auto choice and is refused anywhere else.

Not supported (422): legacy functions / function_call (use tools / tool_choice), n > 1, non-text content parts.

Anthropic Messages

Supported params: model, messages, max_tokens, system, stream, tools, tool_choice, thinking, plus stop_sequences, metadata.user_id, and the sampling family (temperature, top_p).

Content blocks: text, tool_use, tool_result; a cache_control marker on a text block, a system block, a text block inside a tool_result, or a tools[] entry reaches Anthropic in its native form. Also accepted: models (ordered fallback list) and routing_bias. Extra fields on message objects are rejected with 422.

Prompt caching

What you may send. On /v1/chat/completions a message's content is a string or an array of up to 50 text parts ({"type":"text","text":"…"}); a part of any other type — image_url and the rest — is refused with 422 validation_failed naming the offending part type. On /v1/messages a message's content is a string or up to 200 blocks of type text, tool_use or tool_result, system is a string or up to 20 text blocks, and a tool_result holds up to 50 text blocks; any other block type is refused the same way, naming it. Every text part counts toward the input estimate that sizes the reservation and toward the bill.

Where the breakpoint goes. A text part may carry cache_control — {"type":"ephemeral"}, optionally with a ttl of 5m (the default) or 1h; any other value is refused before the request is priced. The marker is accepted on a message text part, on a system text block, on a text block inside a tool_result, and on a tools[] entry. It is refused on a part with no text — a breakpoint must ride text — and on the tool_result block itself rather than on a text block inside it. How many breakpoints one request may carry is capped by the provider; above that cap the provider refuses the request, and its own explanation comes back in the error message.

Marked request body
{
  "model": "anthropic/claude-opus-4.8",
  "messages": [
    {
      "role": "system",
      "content": [
        {
          "type": "text",
          "text": "You are a contract analyst. Answer only from the contract below.",
          "cache_control": { "type": "ephemeral", "ttl": "1h" }
        }
      ]
    },
    {
      "role": "user",
      "content": [
        { "type": "text", "text": "<the full contract text>", "cache_control": { "type": "ephemeral" } },
        { "type": "text", "text": "Which termination clauses changed?" }
      ]
    }
  ]
}

A provider that cannot express one. The marker reaches Anthropic-served models in Anthropic's own block form. On every other provider the text parts are joined into one string and the marker is removed before the call — the request is served normally, with no error and no warning. Those vendors may still cache implicitly: where a vendor reports the hit in the usage shape the exchange reads for that vendor it is billed at the model's cache-read price, and where a vendor reports a hit in any other shape, those tokens are billed as fresh input.

What it costs. A cache write is billed at the model's cache-write price for the ttl requested — 5m and 1h are priced separately — and a later hit at its cache-read price. Both are components of the same published price list as input and output, under the same routing markup. Cache usage comes back in the wire format you called: usage.prompt_tokens_details.cached_tokens on /v1/chat/completions, usage.cache_read_input_tokens with writes under usage.cache_creation on /v1/messages. GET /v1/inference/estimate is not cache-aware: a hit is not knowable before the request runs, so the quote prices the whole input as fresh. GET /v1/usage reports the same tokens cumulatively for the account behind the key — cachedInputTokens for reads, cacheWrite5mTokens and cacheWrite1hTokens for writes at each lifetime. Each is null until at least one settled request on the account recorded a non-zero count of that kind, so an account that has never used caching carries null for all three, and the counters are incremented from the same breakdown the receipt stores: exactly the token counts the charges were computed from. Both write lifetimes collapse into usage.prompt_tokens_details.cache_creation_tokens on /v1/chat/completions, so the two cumulative write counters reconcile per request only on /v1/messages, which splits them.

Reasoning

Two controls switch reasoning on and map onto the same provider wiring: reasoning_effort from the OpenAI wire format, and the structured reasoning object of the OpenRouter convention. An OpenRouter-native client works unchanged.

FieldMeaning
reasoning.effortReasoning depth: low, medium or high. Equivalent to reasoning_effort.
reasoning.max_tokensExplicit thinking budget, 1024–200000 tokens. Mutually exclusive with effort.
reasoning.enabledtrue without a budget uses effort medium; false switches reasoning off.
reasoning.excludeThink, but keep the reasoning out of the response.

Any reasoning object without enabled: false switches reasoning on: a bare reasoning: {} and reasoning: {"exclude": true} both run at effort medium and are billed for the thinking tokens they generate.

Contradictory values across the two controls are 422 validation_failed naming both fields — never a silent priority. That covers two different effort levels, an effort together with reasoning.max_tokens, and enabled: false alongside any requested budget. Agreeing values are accepted, and exclude combines freely with either control.

The effort level becomes a thinking budget as a share of max_tokens (default 4096), floored at 1024 tokens — Anthropic's minimum, applied on every provider. The budget must stay below max_tokens; a request that leaves no room for an answer returns 422 validation_failed naming both numbers, before it reaches a provider and before any credits move.

EffortThinking budget
low20% of max_tokens
medium50% of max_tokens
high80% of max_tokens

On OpenAI-shaped providers the effort level is forwarded as reasoning_effort; an explicit reasoning.max_tokens maps onto the closest effort bucket in the same table.

Reasoning is available exactly on models whose reasoning output is priced in the catalogue. A single named model without it returns 422 validation_failed; inside a models[] list, such an entry is skipped and the request fails only if no listed entry can reason. To check a model before you send it: GET /v1/oracle/models — and GET /v1/oracle/catalog for models outside the basket — returns a reasoning.reasoningOutput price block for exactly the models that can reason, and null for the rest.

The assistant message carries a reasoning field alongside content; streaming chunks carry delta.reasoning before the content deltas. The name is the same on every provider — a model whose own API returns its thinking under reasoning_content is renamed before the response leaves the exchange, streaming and non-streaming alike, and exclude withholds it just the same.

Non-streaming response
{
  "choices": [{
    "index": 0,
    "message": {
      "role": "assistant",
      "content": "42",
      "reasoning": "The question asks for the answer to life, the universe and everything..."
    },
    "finish_reason": "stop"
  }],
  "usage": { "prompt_tokens": 18, "completion_tokens": 512, "total_tokens": 530 }
}
Streaming chunks
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":...,"model":"anthropic/claude-opus-4.8","choices":[{"index":0,"delta":{"reasoning":"The question asks"},"finish_reason":null}]}
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":...,"model":"anthropic/claude-opus-4.8","choices":[{"index":0,"delta":{"reasoning":" for the answer..."},"finish_reason":null}]}
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":...,"model":"anthropic/claude-opus-4.8","choices":[{"index":0,"delta":{"content":"42"},"finish_reason":null}]}

With reasoning.exclude: true the model still reasons and the reasoning tokens are still generated and billed — only the output is withheld. Hiding the reasoning does not make it free.

On /v1/messages the native thinking parameter sets the same budget as reasoning.max_tokens: {"type":"enabled","budget_tokens":N} obeys the same floor and the same below-max_tokens rule, while {"type":"disabled"} switches reasoning off and no thinking parameter is forwarded at all. If the request falls through to an OpenAI-shaped entry, the budget maps onto the closest effort bucket. Reasoning comes back as thinking content blocks ahead of the text block, carrying the signature Anthropic issued when an Anthropic model served the request and an empty signature when another provider did; streaming emits thinking_delta events on that block, plus signature_delta when a signature was issued.

Message content blocks accept text, tool_use and tool_result; a thinking block in an assistant turn is rejected with 422 validation_failed. Strip thinking blocks from an assistant turn before you replay it.

Reasoning tokens are output tokens: they are counted in completion_tokens and billed at the model's output rate. There is no separate reasoning charge.

Streaming

Both endpoints stream Server-Sent Events (Accept: text/event-stream) when stream: true. Wire formats differ (OpenAI vs Anthropic).

OpenAI /v1/chat/completions
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":...,"model":"openai/gpt-5.5","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":...,"model":"openai/gpt-5.5","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}
...
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":...,"model":"openai/gpt-5.5","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":10,"completion_tokens":5,"total_tokens":15},"x_credits_used":"1234","x_credits_remaining":"98765","x_ratelimit_requests_remaining":123}
data: [DONE]
Anthropic /v1/messages
event: message_start
data: {"type":"message_start","message":{"id":"msg_...","type":"message","role":"assistant","content":[],"model":"anthropic/claude-opus-4.8","stop_reason":null,"stop_sequence":null,"usage":{"input_tokens":10,"output_tokens":0}}}

event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}

event: content_block_stop
data: {"type":"content_block_stop","index":0}

event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"output_tokens":5}}

event: message_stop
data: {"type":"message_stop"}

Stop reasons are translated between wire formats: OpenAI stop ↔ Anthropic end_turn; length ↔ max_tokens; tool_calls ↔ tool_use.

The streaming final chunk also mirrors x_credits_used, x_credits_remaining, and x_ratelimit_requests_remaining into the body.

Response headers

HeaderMeaning
X-Model-UsedModel that actually answered this request, in `vendor/model` form.
X-Model-RequestedModel the exchange set out to serve, in `vendor/model` form — the one the request named, or the head of the chain `auto` resolved to.
X-Fallback-Used`true` exactly when the model that answered is not the one `X-Model-Requested` names.
X-Model-ChooserOn every `auto` answer, what picked the model: `engine` — a choice made for this request; `pin` — the model this conversation is already on; `list` — the operator's preference order; `degraded` — the per-request choice could not be made, so the preference order served. Absent when the request named a model.
X-ProviderVendor slug of the provider that served the request.
X-Routing-StrategyHow the provider key was chosen: `sticky`, `weighted-random`, or `failover` after a retry or a fallback entry took over.
X-Request-IdIdentifier of this routing decision — quote it in a support request.
X-Compute-UsedWei debited for this request after settlement.
X-Compute-RemainingRemaining $COMPUTE balance for the buyer account, in wei.
X-Key-Daily-RemainingRemaining daily spending cap for this key, in wei — `unlimited` if unset.
X-Key-Monthly-RemainingRemaining monthly spending cap for this key.
X-RateLimit-LimitCeiling of the tightest active rate-limit bucket.
X-RateLimit-RemainingRemaining capacity in the tightest bucket.
X-RateLimit-ResetEpoch seconds when the tightest bucket resets.

Errors

Every error response follows the envelope below. Branch on error.code rather than HTTP status — several codes can share the same status.

envelope
{
  "error": {
    "message": "The Bearer API key is missing, malformed, or not recognized.",
    "type": "invalid_request_error",
    "code": "invalid_api_key",
    "param": "authorization",                    // optional, set when the error binds to a field
    "details": { ... },                          // optional, code-specific structured payload
    "issues": [ ... ]                            // present when one or more fields fail validation
  }
}

type is derived from code, with one exception that separates the two 422s: a body that fails schema validation answers validation_error and lists every offending field in issues[], while a body that parsed and was then refused further in — a parameter the target provider refuses, a provider-capability mismatch, an input above the model's window — answers invalid_request_error and carries no issues[]. Both carry code: validation_failed, so type is what tells them apart.

CodeStatusTypeMeaningWhen it occurs
invalid_api_key401invalid_request_errorThe Bearer API key is missing, malformed, or not recognized.The token is empty, does not use the `ct_live_` format, or matches no key the exchange issued. A key that exists but is revoked or frozen answers with its own code instead.
api_key_frozen403forbiddenThe API key is currently frozen and cannot be used for requests.The key was frozen from the API keys page or by support. Reactivate it from the API keys page, or contact support if it was frozen by an administrator.
api_key_revoked401invalid_request_errorThe API key has been revoked and can no longer be used.The key was revoked from the API keys page. Revocation is permanent — issue a new key to continue.
model_restricted403forbiddenThe requested model is not in this key's allowed-models list.The key was restricted to a specific set of models; the request referenced a model outside that set.
validation_failed422invalid_request_errorThe request failed validation. A schema failure lists every issue in `error.issues[]`; a single-field failure names that field in `error.param`.One or more fields — of the body, the query string, or the path — are missing, of the wrong type, or outside their allowed range.
method_not_allowed405invalid_request_errorThe endpoint exists but does not accept this HTTP method.A request used a method the route does not support; the `Allow` response header lists the accepted methods.
payload_too_large413invalid_request_errorThe request body is larger than the exchange accepts.The body exceeded the fixed body ceiling, which covers the largest context the exchange serves; it is refused before it is parsed, so no field-level detail is available.
insufficient_balance402insufficient_quotaThe account's $COMPUTE balance is below the requested amount.Triggered when withdrawing more than the balance, or when an inference request would exceed the balance after billing.
spending_limit_reached429rate_limit_errorA spending cap has been reached.This request would exceed a daily or monthly cap on the key, or a daily, weekly or monthly cap on the account; `details.scope` names which carried it.
all_keys_exhausted502server_errorNo upstream capacity is currently available for the routed model.All contributor keys serving this model were unavailable, rate-limited, or in a cool-down; retry shortly.
rate_limited429rate_limit_errorThe caller exceeded a rate-limit window on this endpoint.Either the per-endpoint request-rate cap or the provider-pool RPM/TPM cap was hit.
stream_interrupted500server_errorThe SSE stream aborted before completion.The upstream provider connection dropped mid-stream. This error is delivered inside the SSE body, not as an HTTP status.
service_unavailable503server_errorA required upstream dependency is temporarily unavailable.A dependency needed for this request is unreachable; security-critical paths intentionally reject rather than degrade.

Rate limits

Per-key. Each API key can carry optional daily and monthly $COMPUTE caps, and the account behind it carries daily, weekly and monthly caps of its own. Exceeding any of them returns 429 spending_limit_reached; what is left of the key's caps is exposed in the X-Key-*-Remaining headers.

Per-pool. Each provider pool has RPM and TPM ceilings. Exceeding either returns 429 rate_limited.

GET /v1/inference/estimate is throttled to 10 requests per second per source IP; it requires no authentication and does not consume $COMPUTE.

Request size

Body ceiling. A request body may be up to 4194304 bytes (4 MiB) — the room the largest context the exchange serves, 1048576 input tokens, needs on the wire. A larger body is refused with 413 before it is parsed. The ceiling is fixed for the whole exchange and applies to /v1/chat/completions and /v1/messages alike.

Per-model maximum. A model may declare the context window its vendor actually serves. A request whose estimated input exceeds it fails with 422 validation_failed naming the model and its limit — before the request reaches a provider and before any $COMPUTE is reserved, so an oversized attempt costs nothing. A model that declares no maximum is bounded only by the body ceiling. GET /v1/inference/estimate runs the same check, so you can size a request before you send it. Every declared window is published as maxInputTokens on GET /v1/models and in the full catalog at https://api.compute.finance/v1/oracle/catalog, absent on a model that declares none.

Inside a fallback chain. An entry of models[] that the exchange cannot serve right now is skipped and the next candidate takes over. An input too long for an entry is not skipped: the length is a property of your request, so the request fails naming that entry rather than quietly falling through to a model you did not ask for. Only model: "auto" narrows on length — there the exchange picked the candidates, so it routes to one whose window holds the request and returns 422 only when none does.

Billing

Cost is estimated up-front, reserved from the balance, then settled against actual token usage. The X-Compute-Used and X-Compute-Remaining headers report the final debit and remaining balance. An insufficient balance returns 402; hitting a per-key or per-account cap returns 429. An account with nothing left to spend is refused before its request is estimated, so a request you could never pay for is never priced.

The estimate is advisory; the bill is not. The input estimate that sizes your reservation and picks the context tier it is held at counts your message text, your tool schema and response format, and the per-message chat framing the provider wraps every turn in. Name a model whose provider publishes its own token counter and the estimate is that vendor’s count of your request — an Anthropic model is counted by Anthropic, and no chat framing is added on top, because that count already carries it. Everything else is counted locally through the o200k_base BPE vocabulary at a pinned tokenizer version and scaled into the target model’s vocabulary by that vendor’s fixed ratio, so the same request reserves the same amount — above a small minimum floor, and outside the rare fallback where the vocabulary cannot be loaded. auto is always counted that way, since the serving model is chosen after the hold is taken, and so is a vendor counter that declines or times out; each candidate model is counted in its own vocabulary and your reservation is held at the costliest of them. It never decides what you pay: settlement charges the token counts the provider reported, at the tier those counts select. GET /v1/inference/estimate tokenizes nothing — it prices the promptTokens you hand it, so counting your prompt with the same vocabulary and adding four tokens per message plus three for the reply quotes a locally counted reservation; a named model its vendor counts holds a different amount, which only that vendor’s own count reproduces.

A model may carry a ladder of context tiers — ordered input-token thresholds, each with its own input and output price. The tier follows the whole input side of the request (prompt plus cache reads plus cache writes), and an input landing exactly on a threshold is priced at the higher tier. Settlement re-selects the tier from the tokens the provider reported and is what you pay; the reservation’s tier only sets how much was held. A cache or reasoning price the catalogue attests is single-level, but a kind it does not price separately is billed at its side’s rate, that side’s context tier included. A flat-priced model — one with no ladder — costs the same at every length. Quote a specific size with GET /v1/inference/estimate — its tierFromInputTokens names the tier your bill will use for that input.

Machine-readable specifications

OpenAPI YAML · OpenAPI JSON · llms-full.txt

Integration contracts

How a model is named and chosen, which API answers which question, and what a request costs in $COMPUTE — the rules the running system enforces, in one place. The wire-format detail behind each stays in the section it belongs to.

Model identifiers and selection

The identifier. GET /v1/models publishes the routable catalogue in OpenAI format, and every id has the form vendor/model — the identifier the Inference API routes and bills on. Pass it as model, or as an entry of the models[] fallback list. The bare model name is accepted in both places, so x-ai/grok-4.5 and grok-4.5 reach the same model; the match is case-insensitive, and an unknown vendor prefix, or a real model named under the wrong vendor, is rejected rather than resolved by guess. GET /v1/oracle/resolve/{key} answers what a spelling resolves to before you send it.

auto. Omitting model, or setting it to auto, hands the choice to the exchange: it serves a model from the operator's ordered preference list, among the entries with serving capacity — send conversation_id to group the turns of one conversation, so the exchange keeps serving them from the same capacity. When that list is empty, or nothing on it can serve, the choice falls through to the cheapest model that can rather than refusing every buyer over an operator setting; a list that can serve nothing also raises an operator alert. That value is a routing directive and never a catalogue id — it takes no vendor prefix and may not appear inside models[]. When no model has capacity the request is refused with 503 rather than served by something you did not choose, and GET /v1/models/availability names the model auto resolves to from that list right now.

routing_bias. An optional cheaper / balanced / better, defaulting to balanced, that leans an auto choice toward the cheaper or the stronger end of the operator's list. It is a bias and not a service level — it influences which model is picked and promises neither a price nor a quality. Sending it beside a named model, or beside a models[] list, is refused with 422 validation_failed: both spend the choice the dial exists to lean. It reaches that choice only where the exchange made one for the request, which X-Model-Chooser names — engine and pin carry the bias, while list and degraded mean the operator's order served and the dial did not act. Changing the value releases the model a conversation_id was being served from, so the next turn is chosen afresh. GET /v1/inference/estimate takes the same dial and runs the same choice, so a quote can be leaned exactly as the request it stands for.

Your own fallback list. Send models — an ordered list of 1 to 8 catalogue ids; model, when present, is tried first, then each entry in turn. An entry lacking capacity or a requested capability is skipped, the request is billed at the model that answered, and the hold covers the costliest listed entry with the excess released at settlement. Combining a list with model: auto is rejected. Fall-through covers capacity and capability only: a prompt above an entry's context window, or a parameter that entry rejects as invalid, aborts the request instead of moving to the next candidate.

No silent substitution. A named model is served or the request fails — the exchange never swaps it for another. Substitution happens only where you asked for it: an entry of your models[] list, or auto. Every answer says which model produced it — X-Model-Used names the model that ran, X-Model-Requested the one the exchange set out to serve, and X-Fallback-Used reads true exactly when the two differ. A named model outside a key's allowedModels is refused with 403 model_restricted rather than swapped for one the key may reach.

Parameters. What each wire format does with every parameter it is sent — take it, refuse it, or drop it under a published policy — is listed under Wire-format deltas.

Which API answers what

Two products share one host. The Inference API is paid, takes a ct_live_* key, and serves the models that can run a request. The Oracle API is free, unauthenticated and read-only, and publishes prices and the index over a catalogue that is a superset of the routable set — a model can be priced there and not be routable here.

QuestionAsk
Which models may I name?GET /v1/models
Which of them can serve right now?GET /v1/models/availability
What does this spelling resolve to?GET /v1/oracle/resolve/{key} · GET /v1/oracle/resolve?keys={csv}
What will this request cost?GET /v1/inference/estimate
What rates am I billed at?GET /v1/oracle/pricing
What do the vendors charge?GET /v1/oracle/models · GET /v1/oracle/catalog
What is compute worth, now and over time?GET /v1/oracle/scu · GET /v1/oracle/history
What have I spent, and what is left?GET /v1/usage

Markup is where the two price lists differ. Every figure on GET /v1/oracle/pricing already carries the routing markup, cache and reasoning components included, so it matches what a request is charged. GET /v1/oracle/models and GET /v1/oracle/catalog publish the vendors' own list prices without it, and /v1/oracle/models and /v1/oracle/basket carry the marked-up pair alongside as markedUpUsdPricePerMillion and markedUpWeiPricePerMillion. Read a number against the endpoint it came from; the two lists are not interchangeable.

What a request costs

Tokens are split by kind, and the kind picks the rate. The catalogue prices each kind separately on the models that support it; a kind a model does not price separately is billed at its side's rate — cache components at input, reasoning output at output.

KindCounts
inputPrompt tokens the provider read fresh.
outputTokens the model generated, thinking tokens included.
cachedInputPrompt tokens served from a cache hit.
cacheWrite5mPrompt tokens written into a 5-minute cache entry.
cacheWrite1hPrompt tokens written into a 1-hour cache entry.
reasoningOutputThinking tokens, on a model that prices them apart from ordinary output.

The arithmetic. Each kind's token count is charged at its own rate per million tokens. On a model with a context-tier ladder, the whole input side of the request — prompt plus cache reads plus cache writes — picks the tier the input and output rates are read from. The legs are summed, the 5% routing markup is applied to that sum exactly once, and the result is divided by the $COMPUTE peg to reach the debit. Settlement re-runs the same arithmetic on the token counts the provider reported, and that is the bill; the reservation only decides how much was held.

The provenance mark a catalogue price carries plays no part in this arithmetic — Endpoints defines what it records.

1000 prompt tokens, 500 completion tokens, openai/gpt-5.5
# Catalogue rates for openai/gpt-5.5 — GET /v1/oracle/catalog
input    $5.00 per 1M tokens
output  $30.00 per 1M tokens

# Peg — GET /v1/oracle/pricing
pegUsd  73.733841 USDC per 1 $COMPUTE

input leg     1000 / 1e6 x  5.00  = $0.005
output leg     500 / 1e6 x 30.00  = $0.015
subtotal                          = $0.020
routing markup             x 1.05 = $0.021
÷ peg                / 73.733841  = 0.00028480816562913085 $COMPUTE

# GET /v1/inference/estimate?model=openai/gpt-5.5&promptTokens=1000&completionTokens=500
{
  "model": "openai/gpt-5.5",
  "resolvedModel": "openai/gpt-5.5",
  "promptTokens": 1000,
  "completionTokens": 500,
  "creditsCost": 0.00028480816562913085,
  "creditsWei": "284808165629131",
  "usdCost": 0.020999999999999998,
  "reservedWei": "284808165629131",
  "tierFromInputTokens": null
}

Catalogue prices and the peg both move, so the figures above are illustrative — GET /v1/inference/estimate is the quote that is always current, and its reservedWei is the hold the charge comes out of.

Model Context Protocol (MCP)

AI agents can query the oracle, estimate costs, and analyze sessions through the official Compute Finance MCP server. Stdio transport, no API key required.

Install in any MCP client: npx @compute-finance/mcp

Client config snippet: {"command":"npx","args":["@compute-finance/mcp"]}

Claude Code one-liner (registers MCP + skills + cost hook): npx @compute-finance/mcp setup

Read-only tools across five layers:

  • data — Live oracle data — basket, price, SCU, CPI, reconstitutions
  • compute — Cost estimation and cross-model comparison
  • render — Pre-formatted session reports used by the Claude Code skills
  • analyze — Raw JSON session and per-inference breakdown
  • history — Aggregate stats across logged sessions

Bundled Claude Code slash skills:

  • /cf-session-management — Measured post-session cost analysis
  • /cf-session-consumption — Per-inference token spend breakdown
  • /cf-active-sessions — Multi-session overview across projects

Reference: github.com/compute-finance/mcp · /.well-known/mcp/server-card.json

Compute Finance ID

Your Compute Finance ID (CF ID) is your identity across all Compute Finance surfaces. Created on first sign-in via email or wallet, it ties together your profile, points balance, and referral code into a single record.

What CF ID stores

FieldDescription
cf_idPublic identifier — format: cf_usr_XXXXXXXXXXXX
emailUsed for notifications. Optional for wallet-only sign-in.
wallet_addressOn-chain address on Base. Created via account abstraction or connected externally.
display_nameUser-chosen name. Defaults to a truncated email or wallet address.
referral_codePermanent 8-character code — format: cf_ref_XXXXXXXX
points_balanceCurrent points total, denormalized from the points ledger

Sign-in flows

Two sign-in methods, both produce a CF ID:

  • Email — Enter your email, receive a 6-digit code, verify. A smart wallet is created on Base and associated with your email. No seed phrase required.
  • Wallet — Connect MetaMask, Coinbase Wallet, or any WalletConnect-compatible wallet. Sign a SIWE message. Your wallet address becomes your CF ID's primary identifier.

Both flows converge on the same CF ID record. You can add an email to a wallet-only account later from your settings.

Organizations

An organization lets several people work on one account — its owner's. The owner's balance pays for the members' requests, and the owner's API keys, spending limits, pools and history are shared with them. Every member keeps a personal account of their own alongside.

Personal and organization context

While you are signed in, your requests run in one context, remembered on your profile. Every account page shows the active context and your role in it in the account switcher at the top of the navigation — at the top of the page on a narrow screen — and switches from there in one step; Settings → Organizations switches too:

  • Personal — your own account — your balance, API keys, spending limits, pools and history.
  • Organization — the owner's account. The balance, API keys, spending limits, pools and history you see are the owner's; chat you send is paid from the owner's balance, and a key issued there belongs to the owner's account. What you may change there is set by your role.

Accepting an invitation makes that organization your active context. Creating an organization makes you its owner but does not switch to it. You can own at most 10 organizations. Over the API, POST /v1/orgs/context with an orgId switches and answers the context now active, orgId: null returns to the personal context, and POST /v1/invitations/accept switches into the organization it joins.

Some things keep the context they started in, whichever one is active later: a connected app acts in the context its grant was issued for, a Spark proposal is carried out in the context it was made in, and an API key always spends the account it was issued on.

Roles

Each member holds exactly one role — Owner, Admin or Member — and every role includes everything the role below it allows. The table is generated from the same declaration the API enforces. An action your role does not allow is refused with org_role_required; the one exception is member addresses, which a Member sees masked rather than being refused.

What the role allowsOwnerAdminMember
Send inference requests and chat, paid from the organization's balanceYesYesYes
See the organization's API keys, spending limits, pools, history and balanceYesYesYes
See who belongs to the organization and the role each member holdsYesYesYes
Create, edit, freeze and revoke the organization's API keysYesYesNo
Set the organization's spending limitsYesYesNo
Create, edit and delete the organization's pools, and see and change who is in themYesYesNo
Connect apps that make changes to the organization's accountYesYesNo
See every member's full email addressYesYesNo
See pending invitations; invite Members, resend or revoke their invitations, and remove MembersYesYesNo
Invite Admins, resend or revoke their invitations, and remove AdminsYesNoNo
Delete the organization once every other member has been removedYesNoNo

A role is never edited: to change someone's role, remove them and invite them again.

Who hands out which role

A role you can hand out is one you can take back. You invite someone to a role, resend or revoke an invitation to it, and remove a member who holds it only if your role could remove that member. The owner role is never handed out.

RoleInvites and removes
OwnerMember, Admin
AdminMember
MemberNobody

Invitations

  • An invitation is addressed to one email address. Only a signed-in account whose confirmed profile email is that address can accept it, so a forwarded link admits nobody else.
  • The link carries a one-time token, of which only a SHA-256 hash is stored. It expires 7 days after it is issued.
  • An organization holds at most 50 pending invitations. An expired invitation takes no slot, and inviting an address whose only invitation has expired replaces it.
  • Resending issues a new link and a new expiry, and the previous link stops working. Revoking withdraws an invitation for good; an accepted invitation cannot be revoked.
  • An invitation is pending, expired, revoked or accepted. While it is pending, its page shows the organization, the inviter, the role, what that role allows and the expiry, and the invitation email lists the same; in any other state the page shows only the state. Reopening an accepted link while signed in as the member it admitted leads back into the organization.

What always stays personal

These act on your own account whichever context is active, so none of them reaches the owner's account from inside an organization:

  • Funding — card top-ups and USDC swaps into your balance.
  • Withdrawal of $COMPUTE from your balance.
  • Saved cards.
  • Auto top-up.

Removing a member, leaving and deleting

  • Removal takes effect at once: the removed member's next request runs in their personal account, and every API key they issued on the owner's account is revoked in the same transaction.
  • An Admin or a Member can leave the organization, with the same effect on the keys they issued.
  • The owner can neither leave nor be removed. The owner can delete the organization instead, once every other member has been removed.
  • Deleting your profile deletes an organization you own alone; while an organization you own still has other members, the profile deletion is refused.

Points

Points track your engagement with Compute Finance — signup, daily logins, oracle interactions, and referrals. As your balance grows, you unlock higher tiers. Points are append-only and recorded in a public ledger per CF ID.

Tiers

TierMin Points
Explorer0
Starter500
Builder2,000
Architect5,000
Titan15,000

How points are earned

V1 supports six earning channels. New channels will be added in future versions and announced via the changelog.

ChannelAmountTriggerFrequency
Signup bonus100 ptsCF ID createdOnce per account
Referral (referrer)250 ptsReferred user completes signupPer successful referral
Referral (referred user)50 ptsUser signs up via referral linkOnce per account
Daily login10 ptsUser logs in on a new calendar day (UTC)Once per day
7-day login streak bonus50 ptsUser logs in 7 consecutive daysOnce per streak completion
Oracle interaction5 ptsUser views a unique model’s pricing on the oracle pageUp to 12 per day (one per model)

Ledger model

The points ledger is an append-only log. No entries are ever updated or deleted. Your canonical points balance is the sum of all your ledger entries. The denormalized points_balance field on your CF ID record is a performance optimization that is reconciled periodically.

FieldTypeDescription
idUUIDPrimary key
cf_idString (FK)The user who earned the points
typeEnumOne of: signup, referral, referral_welcome, daily_login, streak_bonus, oracle_interaction
amountIntegerPoints earned (always positive — append-only, no negative entries)
sourceStringHuman-readable source descriptor (e.g. referral:cf_usr_a3k9m2x7p1b4, oracle:gpt-5.5)
created_atTimestampWhen the points were earned

The append-only design means: no points can be silently removed or altered, the complete earning history is preserved and queryable, and any future audit can reconstruct the exact points balance at any point in time.

Streaks

Use Compute Finance on consecutive calendar days to build a streak. Longer streaks earn bonus points. Your current streak and longest streak are shown in your profile. Streaks reset if you miss a calendar day (UTC).

Leaderboard

The top 50 users by total points are displayed on the public leaderboard, refreshed periodically. The leaderboard shows display name and points only — no other profile data is exposed.

Non-transferability

Points are non-transferable in V1. They have no monetary value, are not convertible to any token or currency, and cannot be sold, traded, or assigned to another account. A points-to-credits conversion ratio for V2 will be announced before V2 ships, and the append-only ledger ensures all V1 earning history is preserved and can be converted accurately at that time.

API endpoints

Points data is queryable via the following endpoints. Authentication is required for endpoints that return personal data; the leaderboard is public.

GET/v1/pointsYour points summary (total, tier, current streak, longest streak)
GET/v1/points/historyPaginated points ledger entries for your CF ID
GET/v1/points/leaderboardTop 50 users by total points (public, no auth)

Referral Program

Every CF ID includes a unique referral code (8 characters). Share your link to invite new users — both sides earn points. The program uses first-touch attribution with a 30-day cookie window.

Referral rewards

EventPointsWho Earns
New user signs up with your code50 ptsNew user (welcome bonus)
You referred a new user250 ptsReferrer

Sharing your referral link

// Referral URL format
https://compute.finance/r?c=cf_ref_XXXXXXXX

Find your referral code in your Compute Finance ID settings. Share buttons are available for X, LinkedIn, Telegram, and copy-to-clipboard. The link uses a 30-day attribution cookie — referrals count when the referred user creates a CF ID within 30 days of clicking your link.

Attribution model

Referral attribution is first-touch with a 30-day window. The first referral link a user clicks is the one credited if they sign up within the window. Subsequent referral links from other users do not overwrite the original cookie.

How it works step by step:

  • A user visits https://compute.finance/r?c=cf_ref_XXXXXXXX
  • The server redirects to https://compute.finance and sets a cookie: cf_ref=cf_ref_XXXXXXXX with Max-Age=2592000 (30 days), SameSite=Lax, Secure
  • The click is recorded in the database with the referral code, timestamp, hashed IP address, and user agent
  • If the user signs up within 30 days — even if they navigate directly to https://compute.finance without the referral link — the cookie is read during CF ID creation and the referral relationship is stored
  • If the user has already clicked a different referral link, the first-touch cookie is preserved

Anti-gaming rules

The referral program enforces several rules to prevent farming and abuse. None of these rules surface error messages — invalid referrals are silently ignored to avoid leaking information about user accounts.

RuleImplementation
No self-referralIf the cf_ref cookie matches the signing-up user's own referral code, the referral is silently ignored
Email deduplicationOne CF ID per email — a user cannot create multiple accounts with the same email to farm referral points
IP rate limitingMaximum 10 CF ID creations per IP address per 24 hours, preventing mass account creation from a single source
Click rate limitingMaximum 100 clicks per referral code per hour. Clicks beyond the limit are not recorded.
Disposable email detectionOptional: reject signups from known disposable email domains (mailinator, guerrillamail, etc.)

Referral dashboard

Your CF ID profile includes a referral dashboard showing:

  • Total referral link clicks (all-time)
  • Total signups from your referral link
  • Conversion rate (signups ÷ clicks)
  • Total points earned from referrals
  • List of referred users (display name or truncated email, signup date, status: pending or confirmed)

Status transitions

A referral has two possible states. Points are awarded when the status transitions from pending to confirmed:

  • pending — The user clicked the referral link but has not yet completed signup
  • confirmed — The user has created a CF ID. Points are awarded to both sides within 60 seconds of confirmation.

The CPI is a public benchmark, not financial advice. Data is sourced from public provider pricing pages, may not reflect negotiated enterprise rates, and does not constitute an offer of any token or security. Compute Finance is independent and not affiliated with any of the providers whose public pricing is tracked in the index.

Read the full disclaimer →

Have a question about the index, the basket, or the methodology?

Ask AI about compute.finance

AI compute, priced.

The public reference price for AI compute.