Best practices
Secure the endpoint that serves a model, copy the id each model states, and decide what reaches the model before the call is made.
A model call is code you own invoking a model you do not operate. Your code chooses which model answers, what reaches it, and what the caller finally receives. The model reads the request exactly as it arrives, so the decisions that shape an answer are taken before the call.
In AI Inference, the code that makes the call is a function, and the call is Azion.AI.run. Each practice below is one of the decisions that call leaves to you. The values themselves belong elsewhere. The request fields are on Model invocation, and the conditions Azion applies to a model are on AI Inference limits. The path a request crosses is on How AI Inference works.
The sections follow the order in which the decisions arise: the authentication the endpoint does not carry, the model id, and the capabilities a model states. The last two cover the input the function sends and the response it reads.
The authentication the endpoint does not carry
An application deployed from the AI Inference Starter Kit answers on a Workload Domain, in the form xxxxxxxxxx.map.azionedge.net, and serves an OpenAI-compatible endpoint on it. Authentication on that endpoint is entirely yours. Azion neither issues a credential for a model call nor requires one. A request that reaches the domain reaches the function unless your own code stops it.
The consequence is what a deployed application does when nothing is added to it. It answers every request that reaches the domain, and each request the endpoint serves runs a model. A model call consumes Compute Time, the duration of active execution multiplied by the memory allocated. An open endpoint is therefore consumption that anyone can start. For the metric and the rate, refer to Pricing.
The check belongs in the function, before the model call. The function receives the request and preprocesses it, so it is the place a request can be rejected without a model running. Attaching a custom domain in place of the Workload Domain changes where the application answers, not what it checks.
Model ids, copied rather than derived
A model is addressed by its id and by nothing else. The id is the first argument to Azion.AI.run and the model field of an HTTP request body, and no base URL or host accompanies it.
The ids do not share one form. Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 is the model’s repository path unchanged. casperhansen-mistral-small-24b-instruct-2501-awq is the same path with the slash replaced by a hyphen. qwen-qwen25-vl-3b-instruct-awq is lowercased and has lost the dot in the version as well. gpt-oss-20b has dropped the organization entirely.
No rule turns a model name into its id. An id assembled from the name is right for some models in the catalog and wrong for others. The mistake is a string, which is why it reads as correct. Copy the id from the model’s own page in AI models, where each page states the id that reaches that model.
The capabilities a model states for itself
What a model accepts is a property of that model rather than of AI Inference. Each page in AI models states its own context length, the input types it takes, and whether it supports tool calling. The catalog is not uniform on any of the three. Context lengths run from 8k tokens to 131k tokens. Some models take text only, and others take text and images. Tool calling is stated as supported on some pages and unsupported on others.
A request that does not match is a request the model cannot serve. An image sent to a text-only model is one such request. So is a prompt longer than the context the model accepts. Both are decided in your code before the call, which is where they are checkable.
Where a model page states nothing for a capability, nothing about that capability is stated. Read a blank as unknown rather than as a yes or a no, and confirm it before code depends on it.
The input decided before the call
Everything a model reads was assembled by your function, in the same execution that makes the call. Preprocessing sits on one side of Azion.AI.run and response shaping on the other, and both are code you wrote.
The call is awaited, so the time the model spends generating is time the function spends waiting. The length of what you send is part of that. A prompt that carries a whole document where a section would do is read in full, inside an execution measured while it runs.
Narrowing the input is done before the call rather than corrected after it. The response states what the request consumed: usage.prompt_tokens carries the tokens read from the request, and usage.total_tokens the request and the response together. Those two fields are how a change to a prompt is measured. For the response fields, refer to Model invocation.
Optional chaining on the model response
Reading the text a chat model generated means walking four levels into an object: choices, its first entry, that entry’s message, and its content. The worked functions Azion publishes walk them as modelResponse?.choices?.[0]?.message?.content, with optional chaining at every level.
The response is not identical for every model and every call. It carries fields that only some models return, such as the reasoning content and the log probabilities. Reading through the levels with optional chaining is what keeps a missing level from ending the read.
The difference is where a failure lands. With optional chaining the function receives undefined and decides what the caller gets; without it, the read itself fails inside the function. One of Azion’s worked functions chains ?.trim() onto the same read, so the value is normalized where it is taken.