Skip to main content

One doc tagged with "model pricing"

View all tags

Configure Model Pricing

This guide describes how to configure model pricing at the System scope. Model pricing is the install-wide catalog of per-model rates (USD per 1,000,000 tokens) that the Inference Launchpad gateway uses for cost accounting and that the Cost dimension of Inference Quotas is measured against. PaletteAI federates the catalog to every Launchpad-enabled spoke; each served model that appears in the catalog gets an explicit rate, and every other served model falls back to the Default rate. Without a Default and an explicit price for a model, that model's requests do not contribute to the Cost dimension of any quota.