---
name: azion-call-a-model-on-ai-inference-from-a-function
description: >-
  Call a chat model on AI Inference from a function with Azion.AI.run, and return the model's answer to the request that reached the function.
---

# Call a model on AI Inference from a function

You call a chat model on AI Inference from a function with the `Azion.AI.run` binding, from Azion Console or the Azion CLI. To turn text into vectors with an embedding model instead, refer to [Embed documents into a vector table with AI Inference](/en/documentation/guides/ai/agents-and-rag/embed-documents-into-a-vector-table/).

The function sends a model id and a request body, waits for the model, and reads the generated text from the response. Azion issues no credential for the call, so any check on who may reach the model is code in the function, before the call.

---

## Prerequisites

- An application with Application Accelerator turned on, which the **Run Function** behavior requires. To create one, refer to [Applications quickstart](/en/documentation/platform/applications/quickstart/).
- The id of the model to call, copied from its page in [AI models](/en/documentation/platform/ai-inference/models/). Ids do not share one form, so an id derived from the model name can be wrong.
- The [Azion CLI](/en/documentation/devtools/cli/) installed and authorized, for the CLI procedure.

Under `azion dev`, `Azion.AI` is `undefined`, and a call to `Azion.AI.run` throws `TypeError: Cannot read properties of undefined (reading 'run')`. Test the call once the function runs on the application.

The examples call `Qwen/Qwen3-30B-A3B-Instruct-2507-FP8` from a function named `ai-chat`, on `POST /api/chat` of `www.example.com`. Replace them with your values.

---

## Create the function that calls the model

`Azion.AI.run` takes the model id as its first argument and the request body as its second, without a `model` field. The body requires `messages`, and each message carries a `system`, `user`, or `assistant` role. `max_tokens` caps what the model generates, and `stream: false` returns the response whole.

The function refuses an empty prompt before the call, because every call runs the model. It reads the text with optional chaining at every level, so a response missing a level yields `undefined`, and the function answers `502` instead of failing inside the read:

```javascript
const MODEL = "Qwen/Qwen3-30B-A3B-Instruct-2507-FP8";

function json(body, status = 200) {
  return new Response(JSON.stringify(body), {
    status,
    headers: { "Content-Type": "application/json" }
  });
}

export default {
  async fetch(request, env, ctx) {
    if (request.method !== "POST") {
      return new Response("Method not allowed", { status: 405 });
    }
    const { prompt } = await request.json();
    if (typeof prompt !== "string" || prompt.trim() === "") {
      return json({ error: "prompt is required" }, 400);
    }

    const modelResponse = await Azion.AI.run(MODEL, {
      "stream": false,
      "max_tokens": 500,
      "messages": [
        { "role": "system", "content": "You are a helpful assistant." },
        { "role": "user", "content": prompt }
      ]
    });

    const text = modelResponse?.choices?.[0]?.message?.content;
    if (!text) {
      return json({ error: "the model returned no text" }, 502);
    }
    return json(modelResponse);
  },
};
```

**Console**

To create the function in Azion Console:

1. **Open the Functions page**

   Access [Azion Console](https://console.azion.com/) > **Products Menu** > **Libraries** > **Functions**.

2. **Select + Function**

3. **Name the function**

   Enter `ai-chat`.

4. **Paste the code above in the Code tab**

5. **Select Save**

The function is saved and available to instantiate on an application.

**CLI**

The CLI reads the code from a local file. To create the function:

1. **Save the code above to index.js**

2. **Run the create command**

   ```bash
   azion create function --name ai-chat --code ./index.js --active true
   ```

3. **Read the output**

   The command prints the ID of the new function:

   ```text
   Created function with ID 56278
   ```

Record the ID. The instance on the application takes it.

The account holds an `ai-chat` function that sends each prompt to the model and returns the model's response.

---

## Run the function and read the answer

The function runs once an instance of it is on the application and a rule selects that instance. Create the instance as [Functions quickstart](/en/documentation/platform/functions/quickstart/) describes, named `ai-chat`. Create the rule as [Run a function on one path, and roll it back](/en/documentation/guides/application-development/functions-and-runtime/run-a-function-on-one-path-and-roll-it-back/) describes, with the path `/api/chat` and the method `POST`. New rules can take a few minutes to propagate.

To send a prompt to the function:

```bash
curl -s -X POST https://www.example.com/api/chat \
  -H 'Content-Type: application/json' \
  -d '{"prompt":"Name the European capitals."}'
```

The function returns the model's `chat.completion` object. The generated text sits at `choices[0].message.content`:

```json
{
  "id": "chatcmpl-example",
  "object": "chat.completion",
  "created": 1767268800,
  "model": "Qwen/Qwen3-30B-A3B-Instruct-2507-FP8",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "reasoning_content": null,
        "content": "Sure! Here is a list of some European capitals...",
        "tool_calls": []
      },
      "logprobs": null,
      "finish_reason": "stop",
      "stop_reason": null
    }
  ],
  "usage": {
    "prompt_tokens": 9,
    "total_tokens": 527,
    "completion_tokens": 518,
    "prompt_tokens_details": null
  },
  "prompt_logprobs": null
}
```

`usage.prompt_tokens` counts the tokens read from the request, and `usage.total_tokens` adds the generated ones. Read them after a change to the prompt to see what the change consumes. A request with an empty `prompt` answers `400` and runs no model.

A `POST` to `/api/chat` returns an answer that the model generated on AI Inference.

---

## Next steps

- [AI models](/en/documentation/platform/ai-inference/models.md): Copy the id of another model, and check its context length, input types, and tool calling.
- [Model invocation](/en/documentation/platform/ai-inference/model-invocation.md): Every field the request body accepts, and every field of the response.
- [Add AI features to existing applications](/en/documentation/use-cases/build-and-run-ai-workloads/add-ai-features-to-existing-applications.md): A summarization route whose function calls a catalog model and caches each summary.
- [Govern access to multiple AI models](/en/documentation/use-cases/build-and-run-ai-workloads/govern-access-to-multiple-ai-models.md): A gateway function that calls a model on AI Inference first and a provider on failure.
