Model invocation
Send a request to an AI Inference model from a function or over HTTP, and read the fields every request body accepts.
You invoke a model with two things: the model id, and an OpenAI-compatible request body. AI Inference accepts that request through two interfaces, and both reach the same models. You call the Azion.AI.run binding inside a function, or you post to the OpenAI-compatible HTTP endpoint that the application you deploy serves. The request fields are the same on both.
The Azion.AI.run binding
Azion.AI.run is the runtime binding a function calls to invoke a model. It takes the model id and the request body this page describes, and the call is asynchronous. The binding belongs to the runtime rather than to AI Inference, so its parameters, the call, and its return value are documented with the other runtime bindings, in AI Inference API.
The HTTP endpoint
The HTTP endpoint accepts a POST request to /v1/chat/completions carrying a JSON body. Azion does not host that endpoint. The host is the application you deploy, which answers at a domain in the form xxxxxxxxxx.map.azionedge.net, or at a custom domain attached to it. An application deployed from the AI Inference Starter Kit serves the endpoint on that domain.
A request to a chat model over HTTP:
Azion issues no credential for a model call, and the endpoint carries no Azion authentication. Authentication is what you add to your own application. When your application requires a header, include it in the request, for example -H "Authorization: Bearer [TOKEN VALUE]".
Request fields
These fields form the body of a chat model request, whether you send it through the binding or over HTTP. Over HTTP, the body also carries a model field holding the model id. Through the binding, the first parameter carries the id and the body does not repeat it. Each chat model names the OpenAI Chat API as the endpoint it is compatible with, so a field behaves as that schema describes; the types, defaults, and bounds below are the ones AI Inference states. A dash in the Default column means the schema sets no default for that field.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
messages | array | Yes | — | The conversation sent to the model. Each entry is a system, user, or assistant message object. |
temperature | number | No | — | Sampling temperature. Minimum 0, maximum 2. |
top_p | number | No | 1 | Top-p sampling value. Minimum 0, maximum 1. |
n | integer | No | 1 | Number of responses to generate. Minimum 1. |
stream | boolean | No | false | Streams the response instead of returning it whole. Accepts true or false. |
max_tokens | integer | No | — | Maximum number of tokens the model generates. Minimum 1. |
presence_penalty | number | No | 0 | Presence penalty applied to the output. Minimum -2, maximum 2. |
frequency_penalty | number | No | 0 | Frequency penalty applied to the output. Minimum -2, maximum 2. |
tools | array | No | — | The tool definitions the model may call, on a model whose Tool calling capability is Yes. The model pages that support it show the shape. |
Context length, tool calling, and the input types a model accepts are per-model values, and they are not part of this schema. For the model you call, refer to AI models.
Message objects
Every entry in the messages array is a message object, and its role field states which of the three objects it is. Each object requires role and content. The role field is an enum holding one value, so the object and the role are the same choice. The content field takes text as a string, or, on a model that accepts images, an array of content parts.
| Role | Schema object | Required fields |
|---|---|---|
system | SystemMessage | role, content |
user | UserMessage | role, content |
assistant | AssistantMessage | role, content |
A messages array that carries a previous turn:
A model that accepts images takes content as an array of content parts instead of a string. Each part states its type, and carries the field that type uses. The model pages of the models that accept images show this form; a text-only model takes the string.
| Content part | Type | Carries |
|---|---|---|
type | string | Which kind of part this is: text or image_url. |
text | string | The text of the part, when type is text. |
image_url | object | The image of the part, when type is image_url. |
image_url.url | string | The URL the model reads the image from. |
A content array carrying text and an image:
To find out whether a model accepts images, read the input types on its page in AI models.
Embedding requests
An embedding model takes a request body the chat schema does not define. input carries the text to embed and is the only required field. The other two shape the vector that comes back, and each accepts a fixed set of values; a value outside the set is not part of the schema.
| Field | Type | Required | Description |
|---|---|---|---|
input | string, or an array | Yes | The text to embed. An array holds strings, integers, or arrays of integers. |
encoding_format | string | No | The encoding of the embedding in the response: float or base64. |
dimensions | integer | No | The size of the vector the model returns: 256, 512, 1024, 2048, or 4096. |
To store the vectors an embedding model returns and query them, refer to Vector search.
Reranking requests
A reranking request does not carry a messages array. It carries one of two input pairs instead, and each pair holds a single value and a collection to score against it. A request uses one pair or the other, and the schema names query and documents as the required pair.
| Field | Type | Carries |
|---|---|---|
query | string | The query the documents are ranked against. |
documents | array of strings | The documents to rank by relevance to query. |
text_1 | string | The first text the model processes. |
text_2 | array of strings | The texts the model scores against text_1. |
top_n | integer | Declared by the schema, which states no default and no bounds. |
max_tokens_per_doc | integer | Declared by the schema, which states no default and no bounds. |
Response fields
A model returns one object, whatever interface carried the request. The shape depends on what the model does, and the three below cover every model in the catalog. The model pages show a complete payload for the models that have one.
A chat model returns a chat.completion object:
| Field | Type | Carries |
|---|---|---|
id | string | The identifier of this completion. |
object | string | chat.completion. |
created | integer | When the completion was created. |
model | string | The id of the model that answered. |
choices | array | The generated responses, one entry per response requested through n. |
choices[].index | integer | The position of the entry in choices. |
choices[].message | object | The generated message. |
choices[].message.role | string | assistant. |
choices[].message.content | string | The generated text. This is where a chat response is read from. |
choices[].message.reasoning_content | string | The reasoning the model returned, where it returns any. |
choices[].message.tool_calls | array | The tools the model chose to call, on a model whose Tool calling capability is Yes. |
choices[].finish_reason | string | Why generation stopped, such as stop or tool_calls. |
choices[].stop_reason | string | The stop condition that ended generation. |
choices[].logprobs | object | The log probabilities, where the model returns them. |
usage | object | The tokens the request and the response consumed. |
usage.prompt_tokens | integer | The tokens read from the request. |
usage.completion_tokens | integer | The tokens generated in the response. |
usage.total_tokens | integer | The two added together. |
usage.prompt_tokens_details | object | The breakdown of the prompt tokens. |
prompt_logprobs | object | The log probabilities of the prompt, where the model returns them. |
An embedding model returns a list object whose data array holds one embedding per input:
| Field | Type | Carries |
|---|---|---|
id | string | The identifier of this response. |
object | string | list. |
created | integer | When the response was created. |
model | string | The id of the model that answered. |
data | array | The embeddings, one entry per input. |
data[].index | integer | The position of the entry in data. |
data[].object | string | embedding. |
data[].embedding | array of numbers | The vector, of the size dimensions selected. |
usage | object | The tokens the request consumed. |
A reranking model returns the ranked results:
| Field | Type | Carries |
|---|---|---|
id | string | The identifier of this response. |
model | string | The id of the model that answered. |
results | array | The ranked entries, highest score first. |
results[].index | integer | The position the entry held in the input collection. |
results[].document | object | The entry that was ranked. |
results[].document.text | string | The text of that entry. |
results[].relevance_score | number | How relevant the model judged the entry. |
usage.total_tokens | integer | The tokens the request consumed. |