# Mistral 3 Small (24B AWQ)

**Mistral 3 Small** is a large language model with 24 billion parameters, built for conversational agents, text generation, and tool calling. It reads text, returns text, and accepts tool definitions in the request. [AI Inference](/en/documentation/platform/ai-inference/) runs it under the id `casperhansen-mistral-small-24b-instruct-2501-awq`.

## Model details

The model id is the string that selects this model, and it is the first argument `Azion.AI.run` takes. The HuggingFace repository holds the model card.

| Detail                     | Value                                                                                                                       |
| -------------------------- | --------------------------------------------------------------------------------------------------------------------------- |
| Model name                 | Mistral 3 Small                                                                                                             |
| Version                    | 24B AWQ                                                                                                                     |
| Model category             | Large Language Model (LLM)                                                                                                  |
| Model id                   | `casperhansen-mistral-small-24b-instruct-2501-awq`                                                                          |
| Size                       | 24B parameters                                                                                                              |
| HuggingFace model          | [casperhansen/mistral-small-24b-instruct-2501-awq](https://huggingface.co/casperhansen/mistral-small-24b-instruct-2501-awq) |
| OpenAI-compatible endpoint | [OpenAI Chat API](https://developers.openai.com/api/reference/overview)                                                     |
| License                    | [Apache 2.0](https://choosealicense.com/licenses/apache-2.0/)                                                               |

## Capabilities

These values belong to the model and are not part of the request body. For the fields every chat request accepts, with their types, defaults, and bounds, refer to [Model invocation](/en/documentation/platform/ai-inference/model-invocation/).

| Capability     | Value      |
| -------------- | ---------- |
| Input data     | Text       |
| Context length | 32k tokens |
| Tool calling   | Yes        |
| Supports LoRA  | No         |

---

## Usage

A [function](/en/documentation/platform/functions/) invokes the model with `Azion.AI.run`, passing the id as the first argument and an OpenAI-compatible request body as the second. The examples below use that binding. To send the same body to the OpenAI-compatible HTTP endpoint instead, refer to [Model invocation](/en/documentation/platform/ai-inference/model-invocation/).

### Chat completion

This request carries a system message and a user message, and caps the length of the answer with `max_tokens`:

```ts
const modelResponse = await Azion.AI.run("casperhansen-mistral-small-24b-instruct-2501-awq", {
  "stream": false,
  "max_tokens": 1024,
  "messages": [
    { "role": "system", "content": "You are a helpful assistant." },
    { "role": "user", "content": "Name the European capitals." }
  ]
})
```

The model answers with one entry in `choices`, and the generated text sits at `choices[0].message.content`:

```json
{
  "id": "chatcmpl-example",
  "object": "chat.completion",
  "created": 1767268800,
  "model": "casperhansen-mistral-small-24b-instruct-2501-awq",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "reasoning_content": null,
        "content": "Sure! Here is a list of some European capitals...",
        "tool_calls": []
      },
      "logprobs": null,
      "finish_reason": "stop",
      "stop_reason": null
    }
  ],
  "usage": {
    "prompt_tokens": 9,
    "total_tokens": 527,
    "completion_tokens": 518,
    "prompt_tokens_details": null
  },
  "prompt_logprobs": null
}
```

### Tool calling

A tool-calling request adds a `tools` array to the body. Each entry sets `type` to `function` and carries a `function` object holding the name the model calls, a description of what it does, and its `parameters` in JSON Schema form:

```ts
const modelResponse = await Azion.AI.run("casperhansen-mistral-small-24b-instruct-2501-awq", {
  "stream": false,
  "max_tokens": 1024,
  "messages": [
    { "role": "system", "content": "You are a helpful assistant with access to tools." },
    { "role": "user", "content": "What is the weather in London?" }
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get the current weather for a location",
        "parameters": {
          "type": "object",
          "properties": {
            "location": {
              "type": "string",
              "description": "The city and state"
            }
          },
          "required": ["location"]
        }
      }
    }
  ]
})
```

When the model selects a tool, it returns `content` as `null`, names the call in `tool_calls` with the arguments it chose, and sets `finish_reason` to `tool_calls`:

```json
{
  "id": "chatcmpl-tool-example",
  "object": "chat.completion",
  "created": 1746821866,
  "model": "casperhansen-mistral-small-24b-instruct-2501-awq",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "reasoning_content": null,
        "content": null,
        "tool_calls": [
          {
            "id": "chatcmpl-tool-0123456789abcdef0123456789abcdef",
            "type": "function",
            "function": {
              "name": "get_weather",
              "arguments": "{\"location\": \"London\"}"
            }
          }
        ]
      },
      "logprobs": null,
      "finish_reason": "tool_calls",
      "stop_reason": null
    }
  ],
  "usage": {
    "prompt_tokens": 293,
    "total_tokens": 313,
    "completion_tokens": 20,
    "prompt_tokens_details": null
  },
  "prompt_logprobs": null
}
```

---

## Related resources

- [Model invocation](/en/documentation/platform/ai-inference/model-invocation.md): Every field a request body accepts, and the HTTP endpoint that takes the same body.
- [AI models](/en/documentation/platform/ai-inference/models.md): The other models AI Inference runs, and the id each one answers to.
- [AI Inference](/en/documentation/platform/ai-inference.md): The product that runs this model.
- [Functions](/en/documentation/platform/functions.md): The product whose runtime holds the binding the examples on this page call.
- [AI Inference limits](/en/documentation/platform/ai-inference/limits.md): The conditions under which Azion terminates or deprovisions a model.
- [Glossary](/en/documentation/platform/ai-inference/glossary.md): Where model id, context length, and the rest of the AI Inference vocabulary are defined.
