AI Inference quickstart
Deploy the AI Inference Starter Kit template and post a first chat request to a model on your own application domain.
This guide instructs you through getting your first response from an AI Inference model.
- Deploy the AI Inference Starter Kit template as an application.
- Open the Workload Domain that the deployment assigns to the application.
- Post a chat request to a model and read the answer it returns.
- Change the model that the function calls.
The deployment creates three objects, and one request passes through all of them:
- The Workload Domain is the address a client posts to. Azion assigns it to the application, in the form
xxxxxxxxxx.map.azionedge.net. - The application receives the request on that domain, applies its policies, and calls the function.
- The function holds the code that calls the model with
Azion.AI.run, and it returns the model response.
You write no code to reach a first response. The template creates all three objects, and the function it installs already serves an OpenAI-compatible endpoint.
Prerequisites
This guide requires an Azion account. The template creates every other object the guide uses: the application, the function, and the domain.
Deploy the AI Inference Starter Kit template
To create the application that answers model requests:
Access Azion Console and select + Create.
Enter a name for the application. For example: ai-inference-quickstart.
A window shows the deployment logs while the deployment runs. When it finishes, the page shows information about the application and the options to continue.
The application exists, with a function that calls a model and a domain that serves it.
Open the application domain
Azion assigns a Workload Domain to the application during the deployment. The domain has the form xxxxxxxxxx.map.azionedge.net, and it is the address a client sends a model request to. The deployment page shows the link to the application.
To see the application in the browser, select the link.
The domain resolves, and the application answers on it. Copy the domain: the model request goes to it.
Send a request to the model
The application serves an OpenAI-compatible endpoint at /v1/chat/completions on its domain. The request body names the model in its model field and carries the conversation in messages.
The function the template installs calls Qwen/Qwen3-30B-A3B-Instruct-2507-FP8. To send a chat request to it, replace <your-application-domain> with the domain of your application and run:
The model answers with one entry in choices, and the generated text sits at choices[0].message.content:
You have your first model response, served by your own application.
Azion issues no credential for this request, and the endpoint carries no Azion authentication. When you add authentication to your application, include the header it requires, for example -H "Authorization: Bearer [TOKEN VALUE]". For every field a request body accepts, refer to Model invocation.
(Optional) Change the model the function calls
The function that the template deployed holds the call to the model, and the first argument of Azion.AI.run is the model id. The function page carries three tabs:
- Main Settings renames the function and changes its basic settings.
- Code holds the function code.
- Arguments defines the arguments the function uses.
To call a different model:
Access Azion Console > Products menu > Functions.
Select the function that carries the same name as the application.
In the Code tab, change the first argument of Azion.AI.run. The call has this form:
The function calls the model whose id is saved in the Code tab. Model ids differ in form between models, so copy the id from the model’s own page in AI models.