Token Metering
This section describes how to manage Token Metering from PaletteAI: creating and editing the Inference Quotas that cap metered usage, issuing the API Keys that attribute each caller's traffic to a quota, and configuring the Model Pricing that converts token counts into cost figures.
Token Metering
Token Metering tracks AI model token consumption across requests, providing visibility into input, output, and total token usage for monitoring, attribution, and cost analysis. Every inference call that passes through the gateway is metered: requests, tokens, and estimated cost are recorded against the caller's quota so you can monitor consumption, attribute usage to the model that generated it, and cap spend before it happens.