InternVL3
Invoke InternVL3, a 1B-parameter vision language model with a 16k-token context length, and send it text and images in one request.
InternVL3 is a vision language model with 1 billion parameters, built for tasks that include GUI agents, industrial image analysis, and 3D vision perception. It reads text and images, and returns text. AI Inference runs it under the id opengvlab-internvl3-1b-instruct.
Model details
The model id is the string that selects this model, and it is the first argument Azion.AI.run takes. The HuggingFace repository holds the model card.
| Detail | Value |
|---|---|
| Model name | InternVL3 |
| Version | Instruct 1B |
| Model category | Vision-Language Model (VLM) |
| Model id | opengvlab-internvl3-1b-instruct |
| Size | 1B parameters |
| HuggingFace model | OpenGVLab/InternVL3-1B-Instruct |
| OpenAI-compatible endpoint | OpenAI Chat API |
| License | Apache 2.0 |
Capabilities
These values belong to the model and are not part of the request body. For the fields every chat request accepts, with their types, defaults, and bounds, refer to Model invocation.
| Capability | Value |
|---|---|
| Input data | Text and image |
| Context length | 16k tokens |
| Tool calling | No |
| Supports LoRA | No |
Usage
A function invokes the model with Azion.AI.run, passing the id as the first argument and an OpenAI-compatible request body as the second. The examples below use that binding. To send the same body to the OpenAI-compatible HTTP endpoint instead, refer to Model invocation.
Chat completion
A text request carries a system message and a user message, each holding its content as a string:
The model answers with one entry in choices, and the generated text sits at choices[0].message.content.
Multimodal input
An image goes in the same messages array as the text. The user message takes its content as an array of parts instead of a string, and each part states its type and carries the field that type uses:
The model reads the image at that URL and answers about it in choices[0].message.content. For a field-by-field description of the parts a content array takes, refer to Model invocation.