# Qwen3 Embedding 4B

**Qwen3 Embedding 4B** is a multilingual embedding model with 4 billion parameters across 36 layers, and it supports instruction-conditioned embeddings. It reads text and returns a vector that represents its meaning, for retrieval over text and code, classification, clustering, and bitext mining. [AI Inference](/en/documentation/platform/ai-inference/) runs it under the id `Qwen/Qwen3-Embedding-4B`.

## Model details

The model id is the string that selects this model, and it is the first argument `Azion.AI.run` takes. The HuggingFace repository holds the model card.

| Detail                     | Value                                                                                         |
| -------------------------- | --------------------------------------------------------------------------------------------- |
| Model name                 | Qwen3 Embedding 4B                                                                            |
| Version                    | Original                                                                                      |
| Model category             | Embedding                                                                                     |
| Model id                   | `Qwen/Qwen3-Embedding-4B`                                                                     |
| Size                       | 4B parameters                                                                                 |
| HuggingFace model          | [Qwen/Qwen3-Embedding-4B](https://huggingface.co/Qwen/Qwen3-Embedding-4B)                     |
| OpenAI-compatible endpoint | [Embeddings](https://developers.openai.com/api/reference/resources/embeddings/methods/create) |
| License                    | [Apache 2.0](https://choosealicense.com/licenses/apache-2.0/)                                 |

## Capabilities

The model's native vector width is 2560, and the `dimensions` request field selects one of the widths listed below. For `dimensions`, `encoding_format`, and the rest of the request body, refer to [Model invocation](/en/documentation/platform/ai-inference/model-invocation/).

| Capability        | Value                      |
| ----------------- | -------------------------- |
| Input data        | Text                       |
| Context length    | 32k tokens                 |
| Output dimensions | 256, 512, 1024, 2048, 4096 |

---

## Usage

A [function](/en/documentation/platform/functions/) invokes the model with `Azion.AI.run`, passing the id as the first argument and an OpenAI-compatible request body as the second. The examples below use that binding. To send the same body to the OpenAI-compatible HTTP endpoint instead, refer to [Model invocation](/en/documentation/platform/ai-inference/model-invocation/).

### Embedding

This request carries the text in `input`, sets `encoding_format` to `float`, and selects a 256-dimension vector with `dimensions`:

```ts
const modelResponse = await Azion.AI.run("Qwen/Qwen3-Embedding-4B", {
  "input": "The food was delicious and the waiter...",
  "encoding_format": "float",
  "dimensions": 256
})
```

The model answers with one entry in `data`, and the vector sits at `data[0].embedding`:

```json
{
  "id": "embd-0123456789abcdef0123456789abcdef",
  "object": "list",
  "created": 1746821207,
  "model": "Qwen/Qwen3-Embedding-4B",
  "data": [
    {
      "index": 0,
      "object": "embedding",
      "embedding": [0.01, ..., 0.005]
    }
  ],
  "usage": {
    "prompt_tokens": 11,
    "total_tokens": 11,
    "completion_tokens": 0,
    "prompt_tokens_details": null
  }
}
```

---

## Related resources

- [Model invocation](/en/documentation/platform/ai-inference/model-invocation.md): Every field a request body accepts, and the HTTP endpoint that takes the same body.
- [Vector search](/en/documentation/platform/sql-database/vector-search.md): Where the vectors this model returns are stored and queried.
- [AI Inference](/en/documentation/platform/ai-inference.md): The product that runs this model.
- [Functions](/en/documentation/platform/functions.md): The product whose runtime holds the binding the examples on this page call.
- [AI Inference limits](/en/documentation/platform/ai-inference/limits.md): The conditions under which Azion terminates or deprovisions a model.
- [Glossary](/en/documentation/platform/ai-inference/glossary.md): Where model id, context length, and the rest of the AI Inference vocabulary are defined.
