> ## Documentation Index
> Fetch the complete documentation index at: https://docs.cloud.vessl.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Inference pricing

> Provisioned Throughput pricing: the PTU rate, contract terms, and capacity per PTU.

Below are the current Provisioned Throughput prices, in Provisioned Throughput Units (PTUs). Prices may change over time.

### Provisioned Throughput

| Item              | Price                                           |
| ----------------- | ----------------------------------------------- |
| **PTU rate**      | **\$0.05/PTU/minute**, the same for every model |
| **PTU quantity**  | Arranged individually with sales                |
| **Contract term** | Arranged individually with sales                |
| **Billing cycle** | Monthly                                         |

**Example**: A contract of 8 PTUs costs \$576 per day, and about \$17,520 per month on the 730-hour month the **Estimate total cost** calculator uses.

### Capacity per PTU

One PTU processes up to the following tokens per minute (TPM), by model and token type. These are default values; a contract can set its own capacity per PTU, which takes precedence.

| Model          | Input (TPM) | Cached input (TPM) | Output (TPM) |
| -------------- | ----------- | ------------------ | ------------ |
| **MiniMax M3** | 166,667     | 833,333            | 41,667       |
| **GLM 5.2**    | 35,714      | 192,308            | 11,364       |

<Info>
  Use the **Estimate total cost** calculator on any model page to size a mixed workload and compare the cost with a frontier model. See [Understand PTUs and billing](/inference/ptu).
</Info>

* [Talk to sales to set up Provisioned Throughput](https://vessl.ai/en/talk-to-sales?utm_source=docs\&utm_medium=referral\&utm_campaign=inference_en\&utm_content=pricing_cta).

<Note>
  Values are indicative and subject to updates.
</Note>

<Info>
  Last updated: 2026-09-14
</Info>
