# AI Inference quickstart

This guide instructs you through getting your first response from an AI Inference model.

- Deploy the AI Inference Starter Kit template as an application.
- Open the Workload Domain that the deployment assigns to the application.
- Post a chat request to a model and read the answer it returns.
- Change the model that the function calls.

The deployment creates three objects, and one request passes through all of them:

1. The **Workload Domain** is the address a client posts to. Azion assigns it to the application, in the form `xxxxxxxxxx.map.azionedge.net`.
2. The **application** receives the request on that domain, applies its policies, and calls the function.
3. The **function** holds the code that calls the model with `Azion.AI.run`, and it returns the model response.

You write no code to reach a first response. The template creates all three objects, and the function it installs already serves an OpenAI-compatible endpoint.

---

## Prerequisites

This guide requires an Azion account. The template creates every other object the guide uses: the application, the function, and the domain.

---

## Deploy the AI Inference Starter Kit template

> **Caution**
>
> The template uses [Application Accelerator](/en/documentation/platform/applications/#application-accelerator), [Functions](/en/documentation/platform/functions/), and [AI Inference](/en/documentation/platform/ai-inference/). These products can generate usage-related costs. For the rates, refer to [Pricing](/en/documentation/fundamentals/pricing/#ai-inference).

To create the application that answers model requests:

1. **Open the templates page**

   Access [Azion Console](https://console.azion.com/) and select **+ Create**.

2. **Select the AI Inference Starter Kit template**

3. **Name the application**

   Enter a name for the application. For example: `ai-inference-quickstart`.

4. **Select Deploy**

A window shows the deployment logs while the deployment runs. When it finishes, the page shows information about the application and the options to continue.

The application exists, with a function that calls a model and a domain that serves it.

---

## Open the application domain

Azion assigns a Workload Domain to the application during the deployment. The domain has the form `xxxxxxxxxx.map.azionedge.net`, and it is the address a client sends a model request to. The deployment page shows the link to the application.

To see the application in the browser, select the link.

> **Note**
>
> The application takes a few minutes to propagate to Azion's data centers. Until it propagates, the domain does not resolve and the page does not open. Wait a few minutes, then select the link again.

The domain resolves, and the application answers on it. Copy the domain: the model request goes to it.

---

## Send a request to the model

The application serves an OpenAI-compatible endpoint at `/v1/chat/completions` on its domain. The request body names the model in its `model` field and carries the conversation in `messages`.

The function the template installs calls `Qwen/Qwen3-30B-A3B-Instruct-2507-FP8`. To send a chat request to it, replace `<your-application-domain>` with the domain of your application and run:

```bash
curl -X POST https://<your-application-domain>/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3-30B-A3B-Instruct-2507-FP8",
    "stream": false,
    "max_tokens": 1024,
    "temperature": 0.7,
    "top_p": 0.9,
    "messages": [
      { "role": "system", "content": "You are a helpful assistant." },
      { "role": "user", "content": "Name the European capitals." }
    ]
  }'
```

The model answers with one entry in `choices`, and the generated text sits at `choices[0].message.content`:

```json
{
  "id": "chatcmpl-example",
  "object": "chat.completion",
  "created": 1767268800,
  "model": "Qwen/Qwen3-30B-A3B-Instruct-2507-FP8",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "reasoning_content": null,
        "content": "Sure! Here is a list of some European capitals...",
        "tool_calls": []
      },
      "logprobs": null,
      "finish_reason": "stop",
      "stop_reason": null
    }
  ],
  "usage": {
    "prompt_tokens": 9,
    "total_tokens": 527,
    "completion_tokens": 518,
    "prompt_tokens_details": null
  },
  "prompt_logprobs": null
}
```

You have your first model response, served by your own application.

Azion issues no credential for this request, and the endpoint carries no Azion authentication. When you add authentication to your application, include the header it requires, for example `-H "Authorization: Bearer [TOKEN VALUE]"`. For every field a request body accepts, refer to [Model invocation](/en/documentation/platform/ai-inference/model-invocation/).

---

## (Optional) Change the model the function calls

The function that the template deployed holds the call to the model, and the first argument of `Azion.AI.run` is the model id. The function page carries three tabs:

- **Main Settings** renames the function and changes its basic settings.
- **Code** holds the function code.
- **Arguments** defines the arguments the function uses.

To call a different model:

1. **Open the Functions page**

   Access [Azion Console](https://console.azion.com/) > **Products menu** > **Functions**.

2. **Select the function**

   Select the function that carries the same name as the application.

3. **In the Code tab, change the model id**

   In the **Code** tab, change the first argument of `Azion.AI.run`. The call has this form:

   ```ts
   const modelResponse = await Azion.AI.run("Qwen/Qwen3-30B-A3B-Instruct-2507-FP8", {
     "stream": true,
     "messages": [
       { "role": "system", "content": "You are a helpful assistant." },
       { "role": "user", "content": "Name the European capitals." }
     ]
   })
   return modelResponse
   ```

4. **Select Save**

The function calls the model whose id is saved in the **Code** tab. Model ids differ in form between models, so copy the id from the model's own page in [AI models](/en/documentation/platform/ai-inference/models/).

---

## Next steps

- [AI models](/en/documentation/platform/ai-inference/models.md): The id to pass for each model, with the context length, input types, and tool calling it states.
- [Model invocation](/en/documentation/platform/ai-inference/model-invocation.md): Every field a request body accepts, through the binding and over HTTP.
- [How AI Inference works](/en/documentation/platform/ai-inference/how-it-works.md): What runs the model, and the path a request takes to reach it.
- [AI Inference Starter Kit](/en/documentation/guides/application-development/frameworks/ai-inference-starter-kit.md): Manage the application the template created, and attach a custom domain to it.
- [AI Inference limits](/en/documentation/platform/ai-inference/limits.md): The conditions under which Azion terminates a model or deprovisions one.
- [Real-Time Events](/en/documentation/platform/real-time-events.md): Read the requests and logs your application produces while it answers model requests.
