Qwen2.5 VL AWQ 3B
Invoke Qwen2.5 VL AWQ 3B, a 3B-parameter vision-language model with a 32k-token context length, and send it text and images in one request.
Qwen2.5 VL AWQ 3B is a vision-language model with 3 billion parameters. It reads text and images, returns text, and handles visual analysis, agentic reasoning, long video comprehension, visual localization, and structured output generation. AI Inference runs it under the id qwen-qwen25-vl-3b-instruct-awq.
Model details
The id is what a request sends to reach this model, and Azion.AI.run takes it as the first argument. The HuggingFace repository holds the model card.
| Detail | Value |
|---|---|
| Model name | Qwen2.5 VL |
| Version | AWQ 3B |
| Model category | Vision-Language Model (VLM) |
| Model id | qwen-qwen25-vl-3b-instruct-awq |
| Size | 3B parameters |
| HuggingFace model | Qwen/Qwen2.5-VL-3B-Instruct-AWQ |
| OpenAI-compatible endpoint | OpenAI Chat API |
| License | Apache 2.0 |
Capabilities
These capabilities are properties of the model, and a request body does not set them. For every field a chat request accepts, with its type, default, and bounds, refer to Model invocation.
| Capability | Value |
|---|---|
| Input data | Text and image |
| Context length | 32k tokens |
| Tool calling | Yes |
| Supports LoRA | Yes |
Usage
You call this model from a function, through the Azion.AI.run binding. The first argument is the model id, and the second is an OpenAI-compatible request body. Each example on this page uses that binding. For the HTTP form of the same request, refer to Model invocation.
Chat completion
This request carries a system message that sets the behavior of the model and a user message that holds the prompt:
One entry arrives in choices, and the text the model generated sits at choices[0].message.content:
Tool calling
A tools array declares the functions the model may call. Every entry sets type to function and carries a function object with three fields: the name the model returns, a description of what the function does, and the parameters it accepts, written as a JSON Schema object:
A response that selects a tool carries null in content, one entry in tool_calls holding the function name and the arguments the model chose, and tool_calls in finish_reason:
Multimodal input
A user message that carries an image sets content to an array of parts instead of a string. Each part states its type: a text part holds the prompt, and an image_url part holds the URL the model reads the image from. For every content part a message accepts, refer to Message objects.
This request sends one text part and one image part inside the same user message:
The model returns the same response object as a chat completion, and what it read in the image sits at choices[0].message.content.