Fastest models
Medians over 7 days, at our gateway.
| Model | First token | Output | |
|---|---|---|---|
No model has enough requests from enough accounts this week. | |||
Most used
Tokens processed Sep 27 – Oct 3, against the 7 days before, with each model's price in rupees.
No model met the publishing threshold in this period.
Volume
Tokens processed each UTC day, stacked by model.
No published usage in the last 30 days.Days appear once a model meets the publishing threshold.
Your wallet
Pick a wallet-credit budget and a kind of request. Every count is priced exactly from today's catalog, then rounded down. These budgets are wallet credit after any top-up fee; see the Terms for the payment-to-credit calculation.
Rs. 100 buys approximately this many short replies on each model. One short reply: 1,200 tokens in, 350 out.
Catalog checked 02:21 NPT and refreshed every minute. Counts are approximate, rounded down from example request sizes, so your own requests will vary. Each request keeps the price that was active when its credit was reserved.
Patterns
Medians over 7 days, at our gateway.
| Model | First token | Output | |
|---|---|---|---|
No model has enough requests from enough accounts this week. | |||
Who made the model, not who serves it.
30 days. Cached input bills at the lower cached rate.
No model met the publishing threshold this month.
Requests by prompt length in tokens, 30 days.
Not enough accounts in any size range this month.
How we count
Tokens processed through the M82 Router API: prompt and completion tokens as each provider reports them, on requests that completed and settled. Failed, cancelled and unreconciled requests are left out. Days are UTC.
Today is the most recent complete day, compared with the day before. This week and This month cover the last 7 and 30 complete days, compared with the period just before. Rising lists models present in both weeks, by their change.
Individual account records are not published. Usage metadata contributes to combined statistics, as explained in our Privacy Policy. A figure is published only when it covers at least 5 different accounts. Smaller models are grouped as Other models, or withheld if the group is still too small. Counts are rounded to three significant digits. Rankings use token counts and request timing, not prompt or response content. Comparisons across overlapping periods may still allow approximate usage inferences.
At our gateway, on real requests. First token is the median wait for the first streamed token. Output speed is completion tokens per second after it. End to end runs from credit authorization to settlement. A model needs 10 requests in the week to appear.
Change is shown only when a model was published in both periods. A model missing from the earlier period may simply have been under the publishing threshold, so we never call it new.
From the live catalog, exactly like billing: the whole request is priced at once, rounded to 12 decimal places, then divided into spendable wallet credit and rounded down to two significant figures, so 3,303 shows as ≈ 3,300. Request sizes are examples, so your own requests will cost more or less.
No. Usage shows what developers run on M82 Router, not which model suits your task. Models also use different numbers of tokens for the same answer.