Deploy the AI Inference Starter Kit template
Deploy the AI Inference Starter Kit template from Azion Console, then update the function it installs and manage the application it creates.
You can deploy the AI Inference Starter Kit template and manage what it creates from Azion Console. The template creates an application with a function and a domain, and the function serves an OpenAI-compatible endpoint.
The AI Inference quickstart covers the same deployment and adds a first request to the model.
Prerequisites
The template uses Application Accelerator, Functions, and AI Inference, which can generate usage-related costs. For the rates, refer to Pricing.
Deploy the template
To create the application, the function, and the domain from one template:
Access Azion Console and select + Create.
Select the AI Inference Starter Kit template.
Enter a name for the application.
Select Deploy.
A window shows the deployment logs while the deployment runs. When it finishes, the page shows information about the application and the options to continue.
The application answers on a Workload Domain in the form xxxxxxxxxx.map.azionedge.net/, and it serves an OpenAI-compatible endpoint on that domain. For the path a request takes from the domain to the model, refer to the architecture of Add AI features to existing applications.
Update the function
The function the template installs holds the call to the model, and its page carries three tabs:
- Main Settings renames the function and changes its basic settings.
- Code holds the function code.
- Arguments defines the arguments the function uses.
To change the function:
Access Azion Console > Products menu > Functions.
Select the function that carries the same name as the application. The list is alphabetical, and the search bar filters it by Function Name only.
Open the tab that holds the setting you want to change.
Select Save.
The function has the changes you saved. For every setting a function holds, refer to Functions.
The Code tab opens on a function ready to run. It receives a POST request, calls the model, and returns the response. The call below sets stream to false, so the model returns the response whole:
The model answers with one entry in choices, and the generated text sits at choices[0].message.content:
This example calls the Qwen3 model. To call another model, change the first argument of Azion.AI.run to its id, listed in AI models. The body also accepts stream, which streams the response instead of returning it whole. For every field the body accepts, refer to Model invocation.
Before you change the call, read AI Inference best practices for the request patterns Azion recommends, and AI Inference limits for the conditions under which Azion terminates a model.
Manage the application
The application page holds every setting the deployment created, and you can change any of them. To open it:
Access Azion Console > Products menu > Applications.
Select the application that carries the name you entered during the deployment. The list is alphabetical, and the search bar filters it by Application Name only.
The application page opens on the settings you can adjust. For what each setting does, refer to Applications quickstart.