Developer documentation
Build with M82
One M82 key. Current models and NPR prices. Connect the client you already use.
01 / Quick start
Get your key and make a request
M82 uses prepaid NPR credit. Your M82 key works across the models in this catalog; you do not need an upstream provider account or key.
- Sign inUse your approved account to open the M82 dashboard.
- Check your creditCheck your wallet balance before sending requests.
- Create an API keyCreate a key in this deployment; local test keys cannot authenticate to production. Copy it when shown, use a separate key for each app, and revoke keys you no longer need.
Keep the key in your server environment or secret manager. Do not put it in browser code, a public repository, or a shared screenshot. The examples below read it from your environment.
Loading the published API address and live catalog…
02 / Client setup
Connect a client
Choose a live chat model. Examples and output limits update with your selection.
Request examples appear after the live catalog loads.
Normal responses include m82.request_id and the exact m82.cost_npr. Find the same request in your Usage page. SDK examples turn off automatic retries to avoid repeating a potentially billable request after an uncertain failure.
03 / Streaming
Receive tokens as they arrive
Streaming uses Server-Sent Events. Keep reading through the final usage chunk and data: [DONE]. An interrupted stream may require settlement review; check Usage before sending the request again.
Select a live model with streaming support to see its example.
04 / Live catalog
Models and current NPR prices
Rates are per million tokens. Your request keeps the price active when credit is authorized, even if prices change while it runs. Cached tokens use the reported cache rate; reasoning tokens are included in output usage.
Current prices are unavailable. No saved price table is substituted.
05 / Usage limits
Plan for both key and account limits
Every inference request must fit both scopes. Creating more keys does not bypass your account limit. Key/account overrides may differ from the deployed defaults shown here.
- Requests and estimated tokens use fixed windows that start with the first request in each scope.
- Chat token limits count an input estimate plus your maximum output allowance before inference. Unused token allowance is not refunded to the rate-limit window. Lower
max_tokenswhen you need a short response. - Streaming holds concurrency until the request finishes or closes. Spend restrictions and available wallet credit are enforced separately.
429includesRetry-Afterin seconds. Keep queues bounded and avoid immediate retry loops.
06 / Troubleshooting
Keep the request ID
All API responses carry an X-Request-ID header. Include that ID when investigating an error; do not share your key or private prompt.
| HTTP | Meaning | Next step |
|---|---|---|
400 | Invalid request | Check the model protocol, JSON fields, and output-token limit. |
401 | Invalid API key | Use your M82 machine key. A dashboard login token is not an API key. |
402 | Insufficient credit | Add credit to your wallet. Authorization reserves a conservative maximum; unused credit is released after settlement. |
403 | Access restricted | Check the key’s allowed models, spend limits, and account status. |
404 | Model unavailable | Refresh the catalog and use an enabled public model ID. |
429 | Rate limited | Respect Retry-After, reduce parallel requests, and use bounded backoff with jitter. |
502 / 503 | Service unavailable | Keep the request ID. Check your Usage page before retrying an interrupted request. |
500 | Processing failed | Keep the request ID and check Usage before retrying. Settlement may require review. |