AI Inference limits
Review the conditions under which Azion terminates or deprovisions an AI Inference model, and where each remaining limit is stated.
A model in AI Inference meets two kinds of boundary. The first is a condition: Azion terminates a model that consumes too much memory or runs too long, and deprovisions one that is never executed. The second is a stated value that bounds one call, such as a request field or a model’s context length. This page carries the conditions, and names the page that states each value.
Model termination and deprovisioning
Azion applies the conditions below to every model. The first two bound a model while it runs, and Azion publishes no value for either ceiling. The third states a number of days.
| Condition | Value | What happens past it |
|---|---|---|
| Memory a model consumes | Not published | Azion may terminate a model that consumes more than the maximum defined memory. |
| Time a model runs | Not published | Azion may terminate a model that runs for longer than the maximum allowed time. |
| Time since creation without an execution | 3 days | Azion may deprovision a model that is created and not executed for more than three days. |
Limits stated on other pages
The values that bound a model call belong to the model, to the request schema, or to the function that makes the call. Each one is stated on the page that owns it. To size a prompt, read the context length on the page of the model you call.
| Limit | What it bounds | Where it is stated |
|---|---|---|
| Context length | The context a model accepts, in tokens. Each model page states its own value. | AI models |
| Request field bounds | The values temperature, top_p, n, presence_penalty, and frequency_penalty accept. | Model invocation |
| Function limits | The code size, memory, execution time, and outbound requests of the function that calls the model. | Functions limits |