M82 Router model rankings

Most used

The models doing the work.

Tokens processed Sep 27 – Oct 3, against the 7 days before, with each model's price in rupees.

  1. No model met the publishing threshold in this period.

Volume

Thirty days of tokens.

Tokens processed each UTC day, stacked by model.

No published usage in the last 30 days.Days appear once a model meets the publishing threshold.

Your wallet

What your rupees buy.

Pick a wallet-credit budget and a kind of request. Every count is priced exactly from today's catalog, then rounded down. These budgets are wallet credit after any top-up fee; see the Terms for the payment-to-credit calculation.

Wallet credit
Request

Rs. 100 buys approximately this many short replies on each model. One short reply: 1,200 tokens in, 350 out.

  1. gpt-oss-20bGoes furthestRs. 0.0085 each≈ 11,000approx. short replies
  2. gpt-oss-120bRs. 0.0166 each≈ 6,000approx. short replies
  3. Qwen3 Coder 30B A3B InstructRs. 0.0291 each≈ 3,400approx. short replies
  4. GLM 4.7 FlashRs. 0.0341 each≈ 2,900approx. short replies
  5. DeepSeek V4.1 FlashRs. 0.0625 each≈ 1,600approx. short replies
  6. Qwen3 Coder NextRs. 0.0679 each≈ 1,400approx. short replies
  7. DeepSeek V4 Pro 0813Rs. 0.238 each≈ 420approx. short replies

Catalog checked 02:21 NPT and refreshed every minute. Counts are approximate, rounded down from example request sizes, so your own requests will vary. Each request keeps the price that was active when its credit was reserved.

Patterns

How it's used.

Fastest models

Medians over 7 days, at our gateway.

Models ranked by median time to first token, fastest first
ModelFirst tokenOutputEnd to end

No model has enough requests from enough accounts this week.

Requests by maker

Who made the model, not who serves it.

Token mix

30 days. Cached input bills at the lower cached rate.

  1. No model met the publishing threshold this month.

Request size

Requests by prompt length in tokens, 30 days.

Not enough accounts in any size range this month.

How we count

Read the numbers right.

What is counted?

Tokens processed through the M82 Router API: prompt and completion tokens as each provider reports them, on requests that completed and settled. Failed, cancelled and unreconciled requests are left out. Days are UTC.

What do Today, This week and This month cover?

Today is the most recent complete day, compared with the day before. This week and This month cover the last 7 and 30 complete days, compared with the period just before. Rising lists models present in both weeks, by their change.

Can anyone see my usage here?

Individual account records are not published. Usage metadata contributes to combined statistics, as explained in our Privacy Policy. A figure is published only when it covers at least 5 different accounts. Smaller models are grouped as Other models, or withheld if the group is still too small. Counts are rounded to three significant digits. Rankings use token counts and request timing, not prompt or response content. Comparisons across overlapping periods may still allow approximate usage inferences.

How is speed measured?

At our gateway, on real requests. First token is the median wait for the first streamed token. Output speed is completion tokens per second after it. End to end runs from credit authorization to settlement. A model needs 10 requests in the week to appear.

Why does a model have no change figure?

Change is shown only when a model was published in both periods. A model missing from the earlier period may simply have been under the publishing threshold, so we never call it new.

How are the rupee counts worked out?

From the live catalog, exactly like billing: the whole request is priced at once, rounded to 12 decimal places, then divided into spendable wallet credit and rounded down to two significant figures, so 3,303 shows as ≈ 3,300. Request sizes are examples, so your own requests will cost more or less.

Does the most-used model mean the best model?

No. Usage shows what developers run on M82 Router, not which model suits your task. Models also use different numbers of tokens for the same answer.