Skip to content

Feature request: generate retrieval-oriented VLM descriptions for extracted charts while preserving OCR text #2620

Description

@nathankw

Summary

When a caption model is configured through .caption(...), NeMo Retriever captions entries in the intermediate images column, and can optionally caption infographic entries, but it does not caption entries in the chart column.

For charts, OCR can recover the title, axis labels, tick labels, and category labels without recovering the relationships conveyed visually by the bars or lines. The resulting LanceDB row can therefore contain all of the chart's words while omitting the chart's most important facts. Consequently, the chart row is searchable but not meaningfully queryable for the facts the chart was created to communicate.

Using the repository's data/multimodal_test.pdf, the stored Chart 1 row listed the five product names and the y-axis tick values, but did not associate any product with its bar value or describe the comparisons conveyed by the bars. Manually cropping the detected chart bounding box and sending it to the same hosted Omni VLM with a retrieval-oriented chart prompt produced a standalone description containing the title, axes, categories, approximate values, relative comparisons, and main takeaways. For example:

The chart compares the fictitious costs of five gadgets. The hammer is
the least expensive item, the premium desk fan is the most expensive,
and the power drill is the second most expensive. The Bluetooth speaker
falls between the hammer and the minifridge. Approximate bar values are
$20, $120, $75, just under $100, and $150, respectively.

Requested behavior

Please add opt-in, retrieval-oriented VLM interpretation for extracted chart regions when the caption stage is enabled. One possible API would be caption_charts=True. Keeping this opt-in would avoid adding model cost and latency for existing .caption(...) users.

The chart implementation could follow the existing infographic behavior:

  1. Crop each chart from page_image using its bbox_xyxy_norm.
  2. Preserve the existing chart OCR in chart[i]["text"].
  3. Store the VLM result separately in chart[i]["caption"] instead of overwriting OCR.
  4. Provide a tested, retrieval-oriented default chart prompt, with an optional override. The prompt should capture visible text, relationships, rankings, trends, and approximate values without promising exact numerical digitization or speculating beyond visible evidence.
  5. Allow model-specific reasoning configuration. The successful diagnostic used enable_thinking=True.
  6. Preserve the VLM result through element explosion and LanceDB storage, for example as a chart_caption row or as combined, provenance-aware chart content.

This would preserve verbatim OCR while adding the semantic relationships that are visible only in the chart geometry.

Reproduction environment

Repository commit: 825d8428d8ac0cbbd6757e8c9f06a88bdcff48d2
Python:            3.12.3
uv:                0.11.17
nemo-retriever:    2026.8.31.dev20260831195117
lancedb:           0.34.0
pandas:            2.3.3
Pillow:            12.3.0
Run mode:          inprocess

The inprocess mode was used because this small test host had four CPUs, while the batch Ray execution plan requested at least ten CPUs.

Test document:

data/multimodal_test.pdf
SHA-256: 450dddab418e46717a83b977c9cd20ff2c001f91b59ef36f1fb16e6e121332ce

Authentication was supplied only through the NVIDIA_API_KEY environment variable. No key is embedded below.

1. Ingest the PDF with hosted NIMs and Omni captioning enabled

From the nemo_retriever project directory, install the remote-inference environment and export an NVIDIA API key:

uv sync --no-dev
export NVIDIA_API_KEY=nvapi-...

Save the following as reproduce_chart_caption.py:

import os
from pathlib import Path

from nemo_retriever import create_ingestor

PDF = Path("../data/multimodal_test.pdf").resolve()
LANCEDB_URI = Path("./lancedb-chart-caption-repro").resolve()
TABLE = "chart-caption-repro"

api_key = os.environ["NVIDIA_API_KEY"]

pipeline = (
    create_ingestor(run_mode="inprocess")
    .files([str(PDF)])
    .extract(
        extract_text=True,
        extract_images=True,
        extract_tables=True,
        extract_charts=True,
        extract_infographics=True,
        page_elements_invoke_url=(
            "https://ai.api.nvidia.com/v1/cv/nvidia/"
            "nemotron-page-elements-v3"
        ),
        ocr_invoke_url=(
            "https://ai.api.nvidia.com/v1/cv/nvidia/nemotron-ocr-v2"
        ),
        table_structure_invoke_url=(
            "https://ai.api.nvidia.com/v1/cv/nvidia/"
            "nemotron-table-structure-v1"
        ),
        api_key=api_key,
    )
    .caption(
        endpoint_url="https://integrate.api.nvidia.com/v1/chat/completions",
        model_name="nvidia/nemotron-3-nano-omni-30b-a3b-reasoning",
        api_key=api_key,
    )
    .embed(
        embed_invoke_url="https://integrate.api.nvidia.com/v1/embeddings",
        model_name="nvidia/llama-nemotron-embed-vl-1b-v2",
        embed_modality="text",
        api_key=api_key,
    )
    .vdb_upload(
        vdb_op="lancedb",
        vdb_kwargs={
            "uri": str(LANCEDB_URI),
            "table_name": TABLE,
        },
    )
)

result = pipeline.ingest()
print(f"Wrote {len(result)} rows")
print("LanceDB URI:", LANCEDB_URI)
print("table:", TABLE)

Run it:

uv run python reproduce_chart_caption.py

2. Inspect the resulting LanceDB rows

Save the following as inspect_chart_rows.py:

import json
from pathlib import Path

import lancedb

LANCEDB_URI = Path("./lancedb-chart-caption-repro").resolve()
TABLE = "chart-caption-repro"

table = lancedb.connect(str(LANCEDB_URI)).open_table(TABLE)
df = table.to_pandas()

print("columns:", list(df.columns))
print("row count:", len(df))

for index, row in df.iterrows():
    metadata = json.loads(row["metadata"])
    if metadata.get("type") == "chart":
        print(f"\nchart row {index}")
        print("metadata:", json.dumps(metadata, indent=2))
        print("text:\n", row["text"])

Run it:

uv run python inspect_chart_rows.py

The final LanceDB schema was:

['vector', 'text', 'metadata', 'source', 'id']

The intermediate chart, table, infographic, and images columns are exploded before storage. Their type is retained in metadata["type"].

The relevant Chart 1 metadata was:

{
  "page_number": 1,
  "page_elements_v3_num_detections": 9,
  "page_elements_v3_counts_by_label": {
    "table": 1,
    "chart": 1,
    "title": 3,
    "text": 4
  },
  "ocr_table_detections": 1,
  "ocr_chart_detections": 1,
  "ocr_infographic_detections": 0,
  "type": "chart",
  "fidelity": "ocr",
  "bbox_xyxy_norm": [
    0.08263888955116272,
    0.4644775390625,
    0.9312499761581421,
    0.8343505859375
  ]
}

Its exact stored text was:

Chart 1 This chart shows some gadgets, and some very fictitious costs. T                 e t            eeee ry iiiiiit oooo Gadgets and their cost $160.00 $140.00 $120.00 $100.00 Dollart $80.00 $60.00 $40.00 $20.00 $- Minifridge Bluetooth speaker Premium desk fan Powerdrill Hammer Cost Cost

This output contains the category labels and y-axis ticks, but it does not say which value belongs to each category. Enabling .caption(...) did not add a chart_caption row or augment this chart content.

The OCR output is useful for discovering that the document contains a chart about gadget costs and for retrieving it by terms such as “Powerdrill” or “Minifridge.” However, it does not preserve the relationships encoded by the bar geometry. A downstream query such as “How much does the Powerdrill cost?”, “Which gadget is most expensive?”, or “Rank the gadgets by cost” cannot be answered from this stored text alone. In this example, chart OCR provides lexical searchability but not a usable interpretation of the chart's quantitative content. A VLM-generated chart interpretation is needed to make those relationships available to retrieval and downstream question answering.

3. Manually crop the detected chart and request a retrieval-oriented description from the same Omni VLM

The following script reproduces the diagnostic experiment. It stops after extraction so that the intermediate page_image, chart, and normalized bounding-box data remain available. It then performs the crop that a chart-caption path could perform and submits the crop through Retriever's caption helper.

Save as probe_chart_with_omni.py:

import base64
import io
import os
from pathlib import Path

import pandas as pd
from PIL import Image

from nemo_retriever import create_ingestor
from nemo_retriever.operators.extract.caption.caption import caption_images

PDF = Path("../data/multimodal_test.pdf").resolve()
api_key = os.environ["NVIDIA_API_KEY"]

extracted = (
    create_ingestor(run_mode="inprocess")
    .files([str(PDF)])
    .extract(
        extract_text=True,
        extract_images=True,
        extract_tables=True,
        extract_charts=True,
        extract_infographics=True,
        page_elements_invoke_url=(
            "https://ai.api.nvidia.com/v1/cv/nvidia/"
            "nemotron-page-elements-v3"
        ),
        ocr_invoke_url=(
            "https://ai.api.nvidia.com/v1/cv/nvidia/nemotron-ocr-v2"
        ),
        table_structure_invoke_url=(
            "https://ai.api.nvidia.com/v1/cv/nvidia/"
            "nemotron-table-structure-v1"
        ),
        api_key=api_key,
    )
    .ingest()
)

page = extracted.loc[extracted["page_number"] == 1].iloc[0]
chart = page["chart"][0]
x1, y1, x2, y2 = chart["bbox_xyxy_norm"]

page_bytes = base64.b64decode(page["page_image"]["image_b64"])
with Image.open(io.BytesIO(page_bytes)) as image:
    image = image.convert("RGB")
    width, height = image.size
    crop = image.crop(
        (
            round(x1 * width),
            round(y1 * height),
            round(x2 * width),
            round(y2 * height),
        )
    )
    buffer = io.BytesIO()
    crop.save(buffer, format="JPEG", quality=95)

crop_b64 = base64.b64encode(buffer.getvalue()).decode("ascii")
probe = pd.DataFrame(
    {
        "images": [
            [
                {
                    "image_b64": crop_b64,
                    "text": "",
                    "bbox_xyxy_norm": [0.0, 0.0, 1.0, 1.0],
                }
            ]
        ]
    }
)

captioned = caption_images(
    probe,
    endpoint_url="https://integrate.api.nvidia.com/v1/chat/completions",
    model_name="nvidia/nemotron-3-nano-omni-30b-a3b-reasoning",
    api_key=api_key,
    prompt=(
        "Create a retrieval-ready description of this chart. Explain in plain language "
        "what the chart represents and the main comparisons, patterns, or takeaways it "
        "shows. Include the title, axes, categories, and legend when relevant so the "
        "description stands alone without the image. Prefer qualitative statements and "
        "approximate comparisons. Do not manufacture exact values when they are not "
        "explicitly labeled on the data marks. Do not speculate beyond information shown "
        "in the chart, including whether the data is realistic or what it implies outside "
        "the displayed comparison."
    ),
    system_prompt=(
        "You are an expert chart analyst. Reason carefully from the visual evidence, "
        "distinguish directly visible text from approximate visual interpretation, and "
        "do not add unsupported real-world conclusions."
    ),
    temperature=0.0,
    max_tokens=1024,
    extra_body={
        "chat_template_kwargs": {
            "enable_thinking": True,
        }
    },
)

print(captioned.iloc[0]["images"][0]["text"])

Observed response:

**Chart 1** displays a bar graph titled **"Gadgets and their cost."** A caption
below the main title states, "This chart shows some gadgets, and some very
fictitious costs."

The vertical axis (Y-axis) is labeled **"Dollars"** and ranges from $0 to
$160.00 in increments of $20.00. The horizontal axis (X-axis) lists five
categories of items: **Hammer**, **Powerdrill**, **Bluetooth speaker**,
**Minifridge**, and **Premium desk fan**. A legend indicates that the dark blue
bars represent **"Cost."**

The chart illustrates a wide disparity in pricing among the listed items. The
**Hammer** has the lowest cost, appearing to be around $20.00. The
**Powerdrill** is significantly more expensive, reaching approximately $120.00.
The **Bluetooth speaker** costs roughly $75.00. The **Minifridge** is priced
near $100.00. Finally, the **Premium desk fan** is the most expensive item on
the chart, with a cost estimated around $150.00.

This output is substantially more useful for retrieval than the OCR row. It supports questions such as “Which gadget is most expensive?”, “Is the Bluetooth speaker more expensive than the hammer?”, “Which products cost approximately $100 or more?”, and “What is the chart comparing?” It also produced reasonable approximate values without requiring exact data-point extraction.

This diagnostic demonstrates that the source crop contains semantic information missing from OCR and that the configured VLM can produce a useful, standalone representation when Retriever sends the chart region with an appropriate prompt. The proposal is therefore chart interpretation for retrieval, not guaranteed exact chart digitization.

Current implementation behavior

The caption operator currently documents and implements the ordinary-image path as “caption only when images[i]["text"] is empty.” It also has a dedicated opt-in caption_infographics=True path that crops infographic regions and writes the VLM result to a separate caption field while preserving OCR text.

There is no corresponding chart path or caption_charts parameter. Separately, the existing content-explosion code already recognizes a caption field and assigns <content type>_caption, so populating chart[i]["caption"] appears consistent with the current structured-content representation.

Expected result

When chart captioning is enabled, the stored representation should retain both:

  • OCR text, including verbatim title, axis labels, tick labels, and category labels.
  • A grounded, retrieval-oriented VLM description that records visible relationships, rankings, trends, and approximate values without claiming guaranteed exact digitization.

The VLM output should have distinct provenance from OCR and should not silently overwrite the OCR text.

As an acceptance check, after ingesting with chart captioning enabled, a retrievable LanceDB row with metadata["type"] == "chart_caption" (or an equivalently documented provenance-aware representation) should contain the VLM interpretation while the original metadata["type"] == "chart" OCR row remains available.

Actual result

Only the OCR-derived chart row was stored. It contained labels and tick values but omitted the category-to-value relationships. The configured Omni caption endpoint was not invoked for the chart region.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions