Skip to main content
Effective date: October 5, 2026 · Last updated: September 28, 2026
This Service Level Agreement (“SLA”) describes the standard availability commitments VESSL makes for eligible paid Self-Service use of the Covered Services. This SLA forms part of the Master Services Agreement. It applies automatically to an eligible paid Self-Service Customer. A B2B Customer is governed by this standard SLA (including Section 4), except to the extent a Custom SLA Addendum, or Order Form applicable to that transaction expressly provides otherwise. Capitalized terms not defined here have the meanings given in the Master Services Agreement. This SLA does not apply to (i) Beta, Preview, Early Access, or other pre-general-availability Services or features (whether or not separately labeled), (ii) Services provided free of charge or under promotional credits, or (iii) periods during which Customer is in material breach of the Agreement (including the AUP) or has unpaid undisputed Fees more than thirty (30) days past due.

1. Definitions

“Billing Month” means a calendar month during the Subscription Term. “Cluster Storage Volume” means a customer-visible persistent storage volume created through the VESSL Platform, bound to a specific Cluster, and mountable by one or more Workspaces or Jobs in that Cluster. “Downtime” means a period during which a Covered Service is Unavailable under Section 3, 4 (Managed GPU Node), or 5 (Storage). “Downtime Minutes” means, for an affected unit of a Covered Service in a Billing Month, the aggregate number of minutes during which a unit is Unavailable under the applicable Section of this SLA, excluding any period that constitutes Excluded Time. Partial minutes will be counted proportionately based on VESSL’s applicable monitoring records, and overlapping periods of Unavailability will not be counted more than once. Any proration expressly permitted under this SLA will apply. “Eligible Request” means a request that Customer submits to a Model Endpoint within the scope of Section 6.1 and that (a) is properly formed and authenticated; (b) is within Customer’s quotas, rate limits, and the documented context, token, request-size, and concurrency limits for that Model Endpoint; (c) does not specify a client-side deadline shorter than the then-current server-side default; and (d) where it is a retry, conforms to the back-off requirement in Section 6.2. “Excluded Time” means the aggregate number of minutes excluded under Section 7. Partial minutes will be counted proportionately. “Failed Request” means an Eligible Request that returns an HTTP 5xx status code, that receives no response within the applicable server-side timeout, or that, being a streaming request, is terminated by VESSL before its terminal event is delivered. A response described in Section 6.4 is not a Failed Request. An Eligible Request that is not a Failed Request succeeds. “Model Endpoint” means a distinct combination of (a) a model identifier made available through an Inference Service and (b) the Region in which it is served, as invoked by Customer through the applicable API. “Monthly Fee” means: (a) for Self-Service, the monetary Fees or paid value of Purchased Credits actually charged for the affected unit in the applicable Billing Month; and (b) for B2B, the amount calculated under the applicable private transaction document. Promotional or complimentary value is excluded. “Monthly Uptime Percentage” means, for a given affected unit of a Covered Service in a Billing Month: 100% × (Total Minutes − Downtime Minutes − Excluded Time) / (Total Minutes − Excluded Time). “Service Credit” means the non-cash remedy calculated and issued under this SLA. A Service Credit is not a Purchased Credit or Promotional Credit. “Total Minutes” means the total number of minutes in a Billing Month. “Unavailable” has the meaning stated for the applicable Covered Service in Sections 3.1, 4.1 and 5.1.

2. Covered Services

This SLA applies to the following Services (“Covered Services”): For B2B Customers, managed GPU nodes provisioned under an Order Form — whether delivered as bare-metal or virtual machine (VM) nodes — are Covered Services under Section 4 of this SLA by default, unless the applicable Order Form or Custom SLA Addendum expressly provides otherwise. Higher availability commitments (e.g., for multi-Region or multi-AZ deployments) may be agreed in a Custom SLA Addendum. Service-specific commitments may be modified or supplemented in an Order Form or Custom SLA Addendum.

3. Compute Cloud SLA — Workspace and Batch Job availability

3.1 Workspace and Batch Job Unavailability

A running Workspace or Batch Job provisioned through the VESSL Platform (each, a “Workload”) is Unavailable when, due to a cause attributable to VESSL, the Workload cannot run, including any of the following:
  • the running container fails its health checks;
  • the GPU or other compute resource allocated to the Workload becomes unusable and cannot be restored by routine recovery operations (such as reboot or reset); or
  • a failure of the VESSL Platform or its control plane prevents the Workload from running.
Container Compute availability is measured at the level of the affected Workload, using VESSL’s platform health and availability monitoring. A Workload is treated as available or Unavailable as a whole; there is no partial-resource proration. Unavailability of Cluster Storage that prevents a Workload from running is addressed under Section 5 and not under this Section 3.

3.2 Provisioning Capacity Indicator (Advisory)

The VESSL Platform may display, for a given resource specification, a real-time capacity indicator (for example, “High,” “Low,” or “Unavailable”) reflecting whether new Workspaces or Batch Jobs of that specification can currently be scheduled on available capacity. This indicator is advisory only: it reflects point-in-time scheduling capacity, may change without notice, and is designed to fail open (so that only a confirmed absence of schedulable capacity blocks a new Workload’s creation). The availability of on-demand capacity to provision new Workloads is not a service-level commitment under this SLA (consistent with the Container Compute Service Terms), and the capacity indicator does not give rise to Service Credits. The Service Credits under this Section 3 apply only to the Unavailability of running Workloads as defined in Section 3.1.

3.3 Container Compute Service Credits

The Service Credit for any individual Workload in a Billing Month will not exceed 30% of that Workload’s Monthly Fee. Aggregate Service Credit caps across all Covered Services are set out in Section 8.3.

3.4 Maintenance and Replacements

Scheduled maintenance, emergency maintenance, planned hardware replacement, and emergency hardware replacement are excluded from Downtime to the extent the maintenance or replacement is reasonably necessary for the operation, security, or integrity of the Services and is conducted in accordance with this Section 3.4. (a) Notice. VESSL will provide at least seven (7) calendar days’ advance notice of scheduled maintenance that is expected to interrupt the Services (via the VESSL console, by email, or in the Documentation). For emergency maintenance or emergency hardware replacement, VESSL will provide as much advance notice as is reasonably practicable and may proceed without prior notice where necessary to protect the security, integrity, or availability of the Services. (b) Excludable maintenance cap. Downtime arising from maintenance and replacement under this Section 3.4 is excluded from Downtime calculations only up to a total of eight (8) hours in any Billing Month. Any such Downtime in excess of eight (8) hours in a Billing Month is not excluded and counts as Downtime. Notice timing, frequency limits, and detailed procedures applicable to each category are described in the Documentation. For Customers with bare-metal GPU deployments, more specific maintenance and replacement commitments may be set out in a Custom SLA Addendum. Maintenance or replacement that materially exceeds the documented procedures, or the eight (8)-hour monthly allowance in Section 3.4(b), is not excluded from Downtime.

4. Compute Cloud SLA — managed GPU node availability (B2B only)

4.1 Managed GPU Node Unavailability

A managed GPU node provisioned through VESSL’s control plane, APIs, and management software, whether provisioned as a bare-metal node or a virtual machine (VM) node, is Unavailable when, due to a cause attributable to VESSL, Customer cannot reach or use the node for its contracted workload, including any of the following:
  • The node is unreachable from VESSL’s control plane or its on-node agent is unresponsive;
  • One or more contracted GPUs cannot be recovered by routine operations (reboot, GPU reset, driver reload);
  • The pre-installed OS or VESSL-managed software stack (e.g., NVIDIA drivers, OFED, HPC-X) prevents normal node or GPU operation.
Measurement. VESSL probes each node at intervals of no greater than sixty (60) seconds. Unavailability runs from the first of three (3) or more consecutive failed probes until the next successful probe. Where some but not all GPUs on a node are Unavailable while the remainder of the node remains operational, Downtime for that node may be calculated on a prorated basis equal to the fraction of GPUs Unavailable (for example, 2 of 8 GPUs Unavailable corresponds to 25% Downtime applied to that node’s Monthly Fee). VESSL may, at its discretion or where set out in a Custom SLA Addendum, treat such partial failure as full-node Unavailability where operational considerations make per-GPU proration impractical. Unavailability of shared storage that prevents node use is addressed under Section 5 and not under this Section 4.

4.2 Managed GPU Node Service Credits

Service Credits shall only be issued upon a valid claim submitted by Customer in accordance with Section 8 and shall not be applied automatically.

4.3 Maintenance and Replacements

Scheduled maintenance, emergency maintenance, planned hardware replacement, and emergency hardware replacement are excluded from Downtime to the extent (a) VESSL provides reasonable advance notice where practicable, and (b) the maintenance or replacement is reasonably necessary for the operation, security, or integrity of the Services. The necessity and scope of any maintenance or replacement shall be determined solely by VESSL, and such determination shall be deemed reasonable. For scheduled maintenance, VESSL will provide at least seventy-two (72) hours’ prior written notice via support@vessl.ai. For emergency maintenance, VESSL will provide as much advance notice as is reasonably practicable and will notify Customer via support@vessl.ai within twenty-four (24) hours following completion. For Customers with bare-metal GPU deployments, more specific maintenance and replacement commitments may be set out in a Custom SLA Addendum. Maintenance or replacement that materially exceeds the documented procedures is not excluded from Downtime.

5. Storage SLA

5.1 Cluster Storage Unavailability

Cluster Storage availability is measured separately for each affected Cluster Storage Volume. A Cluster Storage Volume is Unavailable when, due to a cause attributable to VESSL, read, write, or file-listing operations on the mounted Volume fail continuously for five (5) minutes or more.

5.2 Cluster Storage Service Credits

The Service Credit for an individual affected Cluster Storage Volume in a Billing Month will not exceed thirty percent (30%) of that Volume’s Monthly Fee. Aggregate Service Credit caps across all Covered Services are set out in Section 8.3.

5.3 Object Storage (Not Covered)

Object Storage is not a Covered Service under this SLA. VESSL does not provide an availability commitment for Object Storage, and no Service Credits are available in respect of Object Storage. VESSL operates Object Storage using commercially reasonable measures as described in the Storage Service Terms and the Documentation; VESSL’s liability for Customer Content stored in Object Storage that is lost, corrupted, or rendered inaccessible is governed by Section 11 (Limitation of Liability) of the MSA.

6. Inference Services SLA

6.1 Scope

This Section 6 applies to Customer’s on-demand use of the Inference Services through a Model Endpoint. Inference Service availability is measured by Monthly Availability of requests (Section 6.2) rather than by Monthly Uptime Percentage. Where Customer reserves inference capacity under a Custom SLA Addendum — Provisioned Throughput (“PT Addendum”), the Service Credits and throughput commitments for that capacity are governed by the PT Addendum. The measurement framework, definitions, and exclusions in this Section 6 remain the baseline for all Inference Services, including reserved capacity. This Section 6 does not apply to (a) an Inference Service, model, or model version designated as Beta, Preview, or Early Access; (b) an endpoint serving a model, adapter, container image, or serving configuration supplied or modified by Customer; (c) Batch or asynchronous inference; or (d) an endpoint Customer has configured to scale to zero replicas, unless an Order Form or Custom SLA Addendum states otherwise.

6.2 Model Endpoint Availability

Monthly Availability, for a Model Endpoint in a Billing Month, is the percentage of the Eligible Requests submitted to that Model Endpoint in that Billing Month that succeed (that is, that are not Failed Requests). It is measured server-side at the boundary of the VESSL inference gateway, using VESSL’s request records. As stated in Section 2, VESSL commits that Monthly Availability for a Model Endpoint will be at least 99.0% in each Billing Month. Two categories of Eligible Requests are left out of that calculation. First, Eligible Requests submitted in a five (5) minute period, measured in UTC, in which Customer submitted fewer than three hundred (300) Eligible Requests to that Model Endpoint. Second, Eligible Requests submitted during Excluded Time. When a request fails, Customer must wait before retrying: at least one (1) second after the first failure, doubling for each consecutive failure up to thirty-two (32) seconds. A repeated identical request that does not follow this is not an Eligible Request. VESSL does not commit to the availability of on-demand capacity to serve any particular request volume. Guaranteed throughput is available only under a Custom SLA Addendum for provisioned throughput capacity.

6.3 Remedies for Failed Requests

(a) No charge for failed requests. Customer is not charged for a Failed Request. Where such a charge is applied, VESSL will reverse it on request or at the next billing reconciliation. (b) Sole remedy for on-demand use. The Inference Services are billed on usage, so Customer’s Fees fall as its requests fail. Section 6.3(a) is accordingly Customer’s sole remedy under this SLA for a failure to meet the SLO in Section 2. Service Credits are not issued for on-demand use of an Inference Service, and Section 8 does not apply to that use. (c) Reserved capacity. Where Customer reserves inference capacity under an Order Form, Customer pays for it whether or not it is used. The availability commitment, service levels, and Service Credits for that capacity are those in the applicable Custom SLA Addendum, and nothing in Section 6.3(a) or 6.3(b) limits them.

6.4 Responses That Are Not Failures

None of the following is a Failed Request, and none gives rise to any remedy under Section 6.3:
  • a response with an HTTP 4xx status code, including 400, 401, 403, 404, 409, 413, 422, and 424;
  • a response with an HTTP 429 status code, or any other throttling, rate limiting, queuing, admission delay, or reduced throughput arising from quotas, rate limits, or the absence of available capacity, or from Customer reaching or exceeding any capacity reserved for Customer under an Order Form;
  • a refusal, filtered response, or truncated response generated by a content-safety, moderation, or guardrail component, whether operated by VESSL or by the model provider;
  • an error resulting from Customer setting a client-side deadline or timeout shorter than the then-current server-side default for the Model Endpoint;
  • a stream terminated by Customer, by Customer’s client library, or by a Customer-side or intermediary timeout, or a request cancelled by Customer;
  • an output that is incomplete, inaccurate, non-deterministic, or otherwise unsatisfactory in content or quality; and
  • an error arising from Customer’s continued use of a model or model version after the end of its published deprecation or end-of-life period.

6.5 No Latency or Throughput Commitment

This Section 6 commits to availability only. VESSL makes no commitment as to time to first token, time per output token, token generation rate, throughput, queueing, or cold-start time. Such a commitment arises only where expressly agreed in a Custom SLA Addendum or an Order Form.

6.6 Model Availability and Deprecation

The set of models made available through an Inference Service may change. Adding, modifying, deprecating, or removing a model or model version in accordance with Section 10 (Model Lifecycle and Deprecation) of the Inference Service Terms does not count against Monthly Availability and gives rise to no remedy under this Section 6.

7. Excluded Time

The following time is excluded from Downtime calculations and from the denominator of Monthly Uptime Percentage calculations to the extent applicable:
  • Force majeure events, including natural disasters, war, pandemic, governmental action, and similar events beyond VESSL’s reasonable control.
  • Failures of public internet, carrier networks, internet backbone providers, upstream transit providers, or networks used by Customer to access the Services (including failures on the Customer side of the connection). Where connectivity is procured by Customer directly or arranged by VESSL on Customer’s behalf with a third-party carrier, the relevant carrier’s terms of service govern, and any resulting loss of external connectivity is excluded from Downtime.
  • Denial-of-service attacks, hacking, or malware not caused by VESSL’s negligence.
  • Customer acts or omissions, including misconfiguration, unauthorized changes, application errors, security incidents within Customer’s control, exceeding documented quotas, rate limits, or other technical limitations, and use outside the documented scope.
  • Scheduled maintenance, emergency maintenance, planned replacement, and emergency replacement to the extent within the conditions of Section 3.4 and Section 4.3.
  • Suspensions, throttling, or limitations imposed pursuant to MSA Section 3.4 (including for breach of the Agreement, the AUP, or past-due undisputed Fees).
For an Inference Service, the following time is also excluded, to the extent applicable:
  • Any period during which VESSL performed a deployment, model load, model version rollout, scale-out, or configuration change to the affected Model Endpoint at Customer’s request or in accordance with Customer’s configuration;
  • Any period during which the affected Model Endpoint is serving a model, model version, adapter, container image, or serving configuration supplied or modified by Customer, or is operating outside the configuration documented for that Model Endpoint; and
  • Any failure of a third-party model provider, model repository, or upstream model API on which the affected Model Endpoint depends, where the failure is not attributable to the VESSL Platform.

8. Service Credit Claims

8.1 How to Claim

To claim Service Credits, Customer must submit a request to support@vessl.ai no later than thirty (30) calendar days after the end of the Billing Month in which the Downtime occurred. A claim not submitted by that deadline is waived to the extent permitted by applicable law. Each claim must include (i) the affected Service and time periods, (ii) supporting logs or monitoring data, and (iii) the requested Service Credit.

8.2 Validation and Application

VESSL will determine Monthly Uptime Percentage and validate each claim in good faith using its monitoring, service, billing, configuration, and other available records. VESSL may also consider supporting evidence submitted by Customer. For Self-Service, an approved Service Credit will be posted to Customer’s account. For B2B, an approved Service Credit will be applied to the next invoice or, if no invoice remains, to a renewal or an equivalent extension of the affected Subscription Term as permitted by the applicable private transaction document.

8.3 Cap

A Service Credit for an individual affected unit will not exceed thirty percent (30%) of that unit’s Monthly Fee. Aggregate Service Credits across all affected Covered Services in a Billing Month will not exceed one hundred percent (100%) of the Monthly Fees for those affected Covered Services in that Billing Month.

8.4 Form of Service Credits

Service Credits are non-transferable, are not redeemable for cash, and are not Purchased Credits. A Self-Service Service Credit expires twelve (12) months after issuance or when Customer’s account terminates, whichever occurs first, except to the extent mandatory law requires otherwise. B2B validity and application are governed by the applicable private transaction document.

9. Sole Monetary Remedy; No Double Recovery

Except for the fee adjustment under Section 6.3(a), Service Credits are Customer’s sole monetary remedy for a failure to meet an SLA metric. This Section 9 does not limit: (a) Customer’s right to terminate for an uncured material breach under Section 5.2 of the MSA; (b) liability arising from VESSL’s willful misconduct or gross negligence to the extent it cannot lawfully be limited; or (c) a right that cannot lawfully be excluded. Customer may not recover both a Service Credit and another monetary remedy for the same loss or event. This SLA does not constitute a warranty or guarantee of uninterrupted or error-free service and does not extend or enhance any warranty stated in the MSA.

10. Updates

Updates to this SLA are governed by MSA Section 14.10.
VESSL AI, Inc. — support@vessl.ai