Call a model on AI Inference from a function
Call a chat model on AI Inference from a function with Azion.AI.run, and return the model's answer to the request that reached the function.
You call a chat model on AI Inference from a function with the Azion.AI.run binding, from Azion Console or the Azion CLI. To turn text into vectors with an embedding model instead, refer to Embed documents into a vector table with AI Inference.
The function sends a model id and a request body, waits for the model, and reads the generated text from the response. Azion issues no credential for the call, so any check on who may reach the model is code in the function, before the call.
Prerequisites
- An application with Application Accelerator turned on, which the Run Function behavior requires. To create one, refer to Applications quickstart.
- The id of the model to call, copied from its page in AI models. Ids do not share one form, so an id derived from the model name can be wrong.
- The Azion CLI installed and authorized, for the CLI procedure.
Under azion dev, Azion.AI is undefined, and a call to Azion.AI.run throws TypeError: Cannot read properties of undefined (reading 'run'). Test the call once the function runs on the application.
The examples call Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 from a function named ai-chat, on POST /api/chat of www.example.com. Replace them with your values.
Create the function that calls the model
Azion.AI.run takes the model id as its first argument and the request body as its second, without a model field. The body requires messages, and each message carries a system, user, or assistant role. max_tokens caps what the model generates, and stream: false returns the response whole.
The function refuses an empty prompt before the call, because every call runs the model. It reads the text with optional chaining at every level, so a response missing a level yields undefined, and the function answers 502 instead of failing inside the read:
To create the function in Azion Console:
Access Azion Console > Products Menu > Libraries > Functions.
Enter ai-chat.
The function is saved and available to instantiate on an application.
The account holds an ai-chat function that sends each prompt to the model and returns the model’s response.
Run the function and read the answer
The function runs once an instance of it is on the application and a rule selects that instance. Create the instance as Functions quickstart describes, named ai-chat. Create the rule as Run a function on one path, and roll it back describes, with the path /api/chat and the method POST. New rules can take a few minutes to propagate.
To send a prompt to the function:
The function returns the model’s chat.completion object. The generated text sits at choices[0].message.content:
usage.prompt_tokens counts the tokens read from the request, and usage.total_tokens adds the generated ones. Read them after a change to the prompt to see what the change consumes. A request with an empty prompt answers 400 and runs no model.
A POST to /api/chat returns an answer that the model generated on AI Inference.