Skip to main content
Provisioned Throughput is set up with our sales team. Tell us about your traffic and we size the capacity with you. After a proof of concept (PoC) on your real traffic, your endpoint goes live.

Estimate your PTUs first

Before you reach out, get a planning number of Provisioned Throughput Units (PTUs) from the calculator. Open InferenceModels in the sidebar, pick a model, and click Estimate total cost.
Use the demo below to see how to enter your traffic and read the PTUs you need alongside the estimated cost.
The calculator sizes the raw capacity your traffic needs. The PTU quantity you contract is arranged individually with sales. See Understand PTUs and billing.

How to request

Talk to sales, or open any model you haven’t contracted yet and click Contact sales. It helps to share:

What happens next

  1. Consultation and profiling — we go through your model, volume, traffic pattern, and target SLA.
  2. Terms and contract — we fix the guaranteed TPM, SLA, and price, and agree on the terms.
  3. PoC — you get a sample API key. We optimize the model for the agreed throughput and SLA goals, then you benchmark on your own traffic and confirm performance and cost.
  4. Go live — your endpoint is provisioned and the model’s Deploy options card shows Provisioned capacity is live. See Send your first request.