Skip to main content

One doc tagged with "inference quota"

View all tags

Manage Inference Quotas

This guide describes how to create, edit, and delete Inference Quotas. An Inference Quota is a usage budget that the PaletteAI Inference Launchpad gateway enforces on inference requests. Each quota caps the volume a caller can consume over a rolling window, measured in Requests, Tokens, or Cost.