Nanonets-OCR-s
Invoke Nanonets-OCR-s, an OCR model that reads a document image and returns its text as structured Markdown, with a 32k-token context length.
Nanonets-OCR-s is an optical character recognition (OCR) model that converts a document image into structured Markdown. The output keeps the layout of the source, including headings, lists, tables, and basic tags, so another model can parse it. AI Inference runs it under the id nanonets/Nanonets-OCR-s.
Model details
The model id selects this model in a request, and Azion.AI.run takes it as the first argument.
| Detail | Value |
|---|---|
| Model name | Nanonets-OCR-s |
| Model id | nanonets/Nanonets-OCR-s |
Capabilities
The values here describe the model itself, and none of them is a field in the request body. To look up the fields a request body accepts, with their types, defaults, and bounds, refer to Model invocation.
| Capability | Value |
|---|---|
| Input data | Text and image |
| Context length | 32k tokens |
Usage
Azion.AI.run runs the model from inside a function: the id goes in the first argument, and an OpenAI-compatible request body goes in the second. The example on this page calls it that way. To send the same body over HTTP instead, refer to Model invocation.
OCR
An OCR request carries the document image and the instruction that tells the model how to transcribe it. Both travel in one user message, as a content array whose parts each state a type. For the fields each part carries, refer to Message objects. This request reads a base64-encoded PNG and caps the answer at 500 tokens:
The model returns one entry in choices, and the transcribed text sits at choices[0].message.content: