Qwen3 Embedding 4B
Invoke Qwen3 Embedding 4B, a 4B-parameter multilingual embedding model, and turn text into vectors for retrieval, classification, and clustering.
Qwen3 Embedding 4B is a multilingual embedding model with 4 billion parameters across 36 layers, and it supports instruction-conditioned embeddings. It reads text and returns a vector that represents its meaning, for retrieval over text and code, classification, clustering, and bitext mining. AI Inference runs it under the id Qwen/Qwen3-Embedding-4B.
Model details
The model id is the string that selects this model, and it is the first argument Azion.AI.run takes. The HuggingFace repository holds the model card.
| Detail | Value |
|---|---|
| Model name | Qwen3 Embedding 4B |
| Version | Original |
| Model category | Embedding |
| Model id | Qwen/Qwen3-Embedding-4B |
| Size | 4B parameters |
| HuggingFace model | Qwen/Qwen3-Embedding-4B |
| OpenAI-compatible endpoint | Embeddings |
| License | Apache 2.0 |
Capabilities
The model’s native vector width is 2560, and the dimensions request field selects one of the widths listed below. For dimensions, encoding_format, and the rest of the request body, refer to Model invocation.
| Capability | Value |
|---|---|
| Input data | Text |
| Context length | 32k tokens |
| Output dimensions | 256, 512, 1024, 2048, 4096 |
Usage
A function invokes the model with Azion.AI.run, passing the id as the first argument and an OpenAI-compatible request body as the second. The examples below use that binding. To send the same body to the OpenAI-compatible HTTP endpoint instead, refer to Model invocation.
Embedding
This request carries the text in input, sets encoding_format to float, and selects a 256-dimension vector with dimensions:
The model answers with one entry in data, and the vector sits at data[0].embedding: