M82 Router

Developer documentation

Build with M82

One M82 key. Current models and NPR prices. Connect the client you already use.

01 / Quick start

Get your key and make a request

M82 uses prepaid NPR credit. Your M82 key works across the models in this catalog; you do not need an upstream provider account or key.

  1. Sign inUse your approved account to open the M82 dashboard.
  2. Check your creditCheck your wallet balance before sending requests.
  3. Create an API keyCreate a key in this deployment; local test keys cannot authenticate to production. Copy it when shown, use a separate key for each app, and revoke keys you no longer need.

Keep the key in your server environment or secret manager. Do not put it in browser code, a public repository, or a shared screenshot. The examples below read it from your environment.

Loading the published API address and live catalog…

02 / Client setup

Connect a client

Choose a live chat model. Examples and output limits update with your selection.

Request examples appear after the live catalog loads.

Normal responses include m82.request_id and the exact m82.cost_npr. Find the same request in your Usage page. SDK examples turn off automatic retries to avoid repeating a potentially billable request after an uncertain failure.

03 / Streaming

Receive tokens as they arrive

Streaming uses Server-Sent Events. Keep reading through the final usage chunk and data: [DONE]. An interrupted stream may require settlement review; check Usage before sending the request again.

Select a live model with streaming support to see its example.

04 / Live catalog

Models and current NPR prices

Rates are per million tokens. Your request keeps the price active when credit is authorized, even if prices change while it runs. Cached tokens use the reported cache rate; reasoning tokens are included in output usage.

Current prices are unavailable. No saved price table is substituted.

05 / Usage limits

Plan for both key and account limits

Every inference request must fit both scopes. Creating more keys does not bypass your account limit. Key/account overrides may differ from the deployed defaults shown here.

  • Requests and estimated tokens use fixed windows that start with the first request in each scope.
  • Chat token limits count an input estimate plus your maximum output allowance before inference. Unused token allowance is not refunded to the rate-limit window. Lower max_tokens when you need a short response.
  • Streaming holds concurrency until the request finishes or closes. Spend restrictions and available wallet credit are enforced separately.
  • 429 includes Retry-After in seconds. Keep queues bounded and avoid immediate retry loops.

06 / Troubleshooting

Keep the request ID

All API responses carry an X-Request-ID header. Include that ID when investigating an error; do not share your key or private prompt.

HTTPMeaningNext step
400Invalid requestCheck the model protocol, JSON fields, and output-token limit.
401Invalid API keyUse your M82 machine key. A dashboard login token is not an API key.
402Insufficient creditAdd credit to your wallet. Authorization reserves a conservative maximum; unused credit is released after settlement.
403Access restrictedCheck the key’s allowed models, spend limits, and account status.
404Model unavailableRefresh the catalog and use an enabled public model ID.
429Rate limitedRespect Retry-After, reduce parallel requests, and use bounded backoff with jitter.
502 / 503Service unavailableKeep the request ID. Check your Usage page before retrying an interrupted request.
500Processing failedKeep the request ID and check Usage before retrying. Settlement may require review.