Qwen2.5 VL AWQ 7B
Invoke Qwen2.5 VL AWQ 7B, a 7B-parameter vision-language model, and send it text and an image in the same request.
Qwen2.5 VL AWQ 7B is a vision-language model with 7 billion parameters, built for visual analysis, agentic reasoning, long video comprehension, visual localization, and structured output generation. The model accepts text and images in the same request, returns text, and takes tool definitions. AI Inference runs it under the id qwen-qwen25-vl-7b-instruct-awq.
Model details
These values identify the model. A request names the model id to reach it, and the HuggingFace repository carries the model card.
| Detail | Value |
|---|---|
| Model name | Qwen2.5 VL |
| Version | AWQ 7B |
| Model category | Vision-Language Model (VLM) |
| Model id | qwen-qwen25-vl-7b-instruct-awq |
| Size | 7B parameters |
| HuggingFace model | Qwen/Qwen2.5-VL-7B-Instruct-AWQ |
| OpenAI-compatible endpoint | OpenAI Chat API |
| License | Apache 2.0 |
Capabilities
Qwen2.5 VL AWQ 7B states these capabilities for itself, and none of them is a request field. To read the fields a chat request carries, with their types, defaults, and bounds, refer to Model invocation.
| Capability | Value |
|---|---|
| Input data | Text and image |
| Context length | 32k tokens |
| Tool calling | Yes |
| Supports LoRA | Yes |
Usage
Every example on this page calls Azion.AI.run from inside a function. The binding takes the model id as its first argument and an OpenAI-compatible request body as its second, so the body never repeats the id. To send the same body to the OpenAI-compatible HTTP endpoint, refer to Model invocation.
Chat completion
A chat request carries the conversation as a messages array, with one system message that sets the behavior and one user message that asks the question:
With stream set to false, the call resolves to one complete response, and modelResponse?.choices?.[0]?.message?.content holds the generated text.
Tool calling
A tool-calling request keeps the same messages and adds a tools array. Every entry declares type as function, then names the function, describes what it does, and states its parameters as a JSON Schema object:
A function reads the model’s answer from the first entry in choices and acts on it, which for this request means calling get_weather.
Multimodal input
A multimodal request puts text and an image in one user message: the content field takes an array of parts instead of a string, and each part states its type. For the fields each part carries, refer to Model invocation. This request sends a question and an image URL together:
The model reads the image from the URL the request gives it, and answers in text at modelResponse?.choices?.[0]?.message?.content.