blade Get your API key

Open Models / Vision, OCR & images

Read images.
Create visuals.

Extract text with OCR, ask questions about images and explore image generation. Put open models to work on the documents and visuals in your product.

Start Building

Document parsing, image understanding and generation.

pixels → readable contentIllustrative example
RECEIPT042
USB-C HUB × 248.00 EUR
page.png
Content your app can use
USB-C HUB × 2
TOTAL 48.00 EUR
Extract content with OCR, then validate and map it to the fields your application needs. Response shape depends on the model.

OCR, image understanding
and generation.

Choose a model for the task: reading a document, interpreting an image or generating something new.

Less retyping.

Extract content from scanned pages, receipts and handwritten notes. Pass it into your review or data-entry workflow.

Context for every image.

Draft product descriptions, explain a screenshot or propose alt text. Review the model’s interpretation before you publish it.

A starting point for design.

Explore image concepts for a campaign, an article or a product screen. Use a generation model, then select and refine the result.

Document OCR

Make the page
usable.

Optical character recognition (OCR) turns text in an image into content your app can process. Send an image to the parsing route and choose a model for general OCR, handwriting, charts or formulas, and use its output in your application.

  1. Set BLADE_API_KEY and BLADE_PARSE_MODEL to an available parsing model for your task.
  2. Save a representative document image as page.png. For PDFs, render the relevant pages to images first.
  3. Run this Python example. Inspect the JSON result before mapping it into your own fields.
Document parsing reference

Python standard library only. For table structure, obtain table regions with /v1/detect and pass table_bboxes.

parse-page.pyPython
import base64
import json
import os
from pathlib import Path
from urllib.request import Request, urlopen

image = base64.b64encode(
    Path("page.png").read_bytes()
).decode("ascii")

request = Request(
    "https://api.models.blade.sh/v1/parse",
    data=json.dumps({
        "model": os.environ["BLADE_PARSE_MODEL"],
        "image": image,
    }).encode(),
    headers={
        "Authorization": "Bearer " + os.environ["BLADE_API_KEY"],
        "Content-Type": "application/json",
    },
    method="POST",
)
with urlopen(request, timeout=120) as response:
    print(json.load(response))

Use a model served by the parsing route. The image is sent as base64; this endpoint does not accept a PDF file directly.

Keep the page readable. Check rotation, resolution and cropping on representative scans.

Validate the output. Check totals, dates and required fields against your application’s rules.

Keep a review path. Show the source image beside extracted content when a person needs to resolve an ambiguity.

Image understanding

Ask about
what’s in view.

Send an image and a question to a vision-capable chat model. The answer can help you describe a product, interpret a screenshot or sort incoming content.

Install openai, set your API key and BLADE_VISION_MODEL, and save a PNG as photo.png. Choose a model that supports image input.

Multimodal chat reference

Image understanding can make mistakes. Use application checks or human review where the answer affects a decision.

describe-image.pyPython
import base64
import os
from pathlib import Path
from openai import OpenAI

client = OpenAI(
    base_url="https://api.models.blade.sh/v1",
    api_key=os.environ["BLADE_API_KEY"],
)
image = base64.b64encode(
    Path("photo.png").read_bytes()
).decode("ascii")

reply = client.chat.completions.create(
    model=os.environ["BLADE_VISION_MODEL"],
    messages=[{"role": "user", "content": [
        {"type": "text", "text": "Describe this product photo."},
        {"type": "image_url", "image_url": {
            "url": "data:image/png;base64," + image,
        }},
    ]}],
    max_tokens=512,
)
print(reply.choices[0].message.content)

Choose the right route.

How is OCR different from image understanding?

OCR reads the content of a document. A vision-capable chat model answers a question about an image. Choose the parsing or chat route according to the model and the output you need.

Will every OCR response have the same fields?

No. The parsing response depends on the model and task. Inspect it, validate it and transform it into your own schema. Table structure also requires table bounding boxes.

Are all visual tasks billed the same way?

No. Check the route and the model: document parsing uses processed-image billing in the API reference, while vision chat uses model token rates. See pricing and the console before estimating a workload.

Contact the team

Loading the form…