@molecule/api-ocr-teleocr

Provider bond · ocr · API (Node) · v1.0.0 · Apache-2.0

TeleOCR provider for molecule.dev — extracts text from images with the open TeleOCR document-parsing model served by vLLM

npm install @molecule/api-ocr-teleocr

npm · Source on GitHub · Implements @molecule/api-ocr

How it works

@molecule/api-ocr-teleocr is a provider bond on the API (Node) side: it implements the ocr core interface (@molecule/api-ocr) with a concrete vendor or library behind it.

Your code calls the core; you wire this provider once at startup. Swapping vendors later is one line in that wiring, not a rewrite.

import { createProvider as createLocalAi } from '@molecule/api-ai-local'
import { requireProvider, setProvider } from '@molecule/api-ocr'
import { createProvider } from '@molecule/api-ocr-teleocr'

// vllm serve StarDoc-AI/TeleOCR (with the TeleOCR_vllm plugin installed)
const ai = createLocalAi({ baseUrl: 'http://localhost:8000/v1' })
setProvider(createProvider({ ai }))

const result = await requireProvider().recognize({
  data: new Uint8Array(imageBytes),
  mimeType: 'image/png',
})
console.log(result.text)

Works with: @molecule/api-ai, @molecule/api-ocr

Secrets: TELEOCR_MODEL (optional)

Reference

TeleOCR provider for molecule.dev.

Implements the @molecule/api-ocr contract with TeleOCR, an open (Apache-2.0 weights), ~1.2B-parameter vision-language model built for document parsing — scanned and photographed pages, tables, formulas, code. It runs on your own GPU behind vLLM's OpenAI-compatible server, reached through @molecule/api-ai-local, so there is no per-call vendor cost.

Quick Start

import { createProvider as createLocalAi } from '@molecule/api-ai-local'
import { requireProvider, setProvider } from '@molecule/api-ocr'
import { createProvider } from '@molecule/api-ocr-teleocr'

// vllm serve StarDoc-AI/TeleOCR (with the TeleOCR_vllm plugin installed)
const ai = createLocalAi({ baseUrl: 'http://localhost:8000/v1' })
setProvider(createProvider({ ai }))

const result = await requireProvider().recognize({
  data: new Uint8Array(imageBytes),
  mimeType: 'image/png',
})
console.log(result.text)

Type

provider

Installation

npm install @molecule/api-ocr-teleocr @molecule/api-ai @molecule/api-ocr

API

Interfaces

TeleOcrConfig

Options for createProvider. ai falls back to the bonded ai provider; model falls back to TELEOCR_MODEL, then DEFAULT_TELEOCR_MODEL.

interface TeleOcrConfig {
  /**
   * The AI provider that reaches the TeleOCR server. Defaults to the provider
   * bonded as `ai` (`@molecule/api-ai` `requireProvider()`) — typically
   * `@molecule/api-ai-local` pointed at the vLLM server
   * (`LOCAL_AI_BASE_URL=http://localhost:8000/v1`).
   */
  ai?: AIProvider
  /**
   * The served model name (what `vllm serve` was started with). Defaults to
   * the `TELEOCR_MODEL` env var, read on each call, then `'StarDoc-AI/TeleOCR'`.
   */
  model?: string
  /**
   * Output token ceiling per call. Defaults to 4096, the reference
   * implementation's `max_new_tokens`. A page that hits it throws
   * `TeleOcrTruncatedError` instead of returning cut-off text.
   */
  maxOutputTokens?: number
}

TeleOcrEnv

Environment variables the provider reads (each on every call, never at import).

interface TeleOcrEnv {
  /** The served model name; overrides `DEFAULT_TELEOCR_MODEL`. */
  TELEOCR_MODEL?: string
}

Types

TeleOcrTask

One of TeleOCR's tasks — a key of TELEOCR_PROMPTS.

type TeleOcrTask = keyof typeof TELEOCR_PROMPTS

Classes

TeleOcrTruncatedError

The model used its whole output budget, so the transcription is cut off. Carries the partial text so a caller can keep it knowingly — never returned silently as if it were the full page. Dense pages hit the reference budget of 4096 tokens; raise maxOutputTokens (within the server's --max-model-len) or split the image.

Functions

convertOtslToHtml(otsl)

Converts TeleOCR's OTSL table output (<fcel>, <nl>, … tokens) into an HTML <table> with rowspan/colspan for merged cells. Cell text is HTML-escaped. Output that is already a <table>…</table> is returned as-is.

function convertOtslToHtml(otsl: string): string
  • otsl — The model's raw answer to the table or scientificFigure prompt.

Returns: An HTML table, or '' when the input holds no cells.

createProvider(config)

Creates a TeleOCR provider.

function createProvider(config?: TeleOcrConfig): OcrProvider
  • config — Which AI provider reaches the TeleOCR server, and the served model name.

Returns: An OcrProvider backed by TeleOCR.

formatFormula(latex)

Normalizes TeleOCR's formula answer the way the reference does: strips a surrounding \[ … \] and wraps the LaTeX in $$ … $$ unless it is already $-delimited.

function formatFormula(latex: string): string
  • latex — The model's raw answer to the formula prompt.

Returns: Display-math LaTeX.

Constants

DEFAULT_TELEOCR_MAX_OUTPUT_TOKENS

The reference implementation's output budget (max_new_tokens).

const DEFAULT_TELEOCR_MAX_OUTPUT_TOKENS: 4096

DEFAULT_TELEOCR_MODEL

The served model name used when neither config nor TELEOCR_MODEL sets one.

const DEFAULT_TELEOCR_MODEL: 'StarDoc-AI/TeleOCR'

OTSL_TOKENS

Every OTSL token, in the order the reference splits on them.

const OTSL_TOKENS: readonly string[]

provider

The provider implementation.

const provider: OcrProvider

TELEOCR_LAYOUT_IMAGE_SIZE

Side length, in pixels, the layout tasks expect the image resized to.

const TELEOCR_LAYOUT_IMAGE_SIZE: 1036

TELEOCR_PROMPTS

The task prompts. text is what recognize() sends; the rest are exported for callers that drive the model directly through @molecule/api-ai.

  • text → plain text.
  • table → OTSL tokens; convert with convertOtslToHtml.
  • formula → LaTeX; normalize with formatFormula.
  • code → the code snippet's text.
  • layout / distortedLayout → layout blocks. The image must first be resized to 1036×1036 (bicubic), and the raw serialization of the blocks is not documented upstream — parse it only after checking it yourself.
  • scientificFigure → the table a chart implies, as OTSL.
const TELEOCR_PROMPTS: {
  readonly text: 'Please output the text content from the image.'
  readonly table: 'This is the image of a table. Please output the table in OTSL format.'
  readonly formula: 'Please write out the expression of the formula in the image using LaTeX format.'
  readonly code: 'The image contains a code snippet, please output the parsing result.'
  readonly layout: 'Analyze the image layout.'
  readonly distortedLayout: '\nMulti-point Layout Segmentation Analysis.'
  readonly scientificFigure: 'This is a scientific figure. Please extract the table implied by this figure.'
}

TELEOCR_SYSTEM_PROMPT

The system prompt the reference implementation sends with every task.

const TELEOCR_SYSTEM_PROMPT: 'You are a helpful assistant.'

Core Interface

Implements @molecule/api-ocr interface.

Bond Wiring

Setup function to register this provider with the core interface:

import { setProvider } from '@molecule/api-ocr'
import { provider } from '@molecule/api-ocr-teleocr'

export function setupOcrTeleocr(): void {
  setProvider(provider)
}

Injection Notes

Requirements

Peer dependencies:

  • @molecule/api-ai >=1.0.0
  • @molecule/api-ocr >=1.0.0

Environment Variables

  • TELEOCR_MODEL (optional) — TeleOCR served model name
    • Setup: The model name your vLLM server was started with (vllm serve <name>, with the TeleOCR_vllm plugin from github.com/caipeng328/TeleOCR installed). The server URL is set on the AI provider this bond uses, e.g. LOCAL_AI_BASE_URL for @molecule/api-ai-local. Defaults to StarDoc-AI/TeleOCR.
    • Get it here: https://huggingface.co/XingChen-AGI/TeleOCR
    • Example: StarDoc-AI/TeleOCR

Runtime Dependencies

  • @molecule/api-ai

  • @molecule/api-ocr

  • Serve it with the TeleOCR_vllm plugin, never stock vLLM alone. The model reuses the architecture name Qwen2_5_VLForConditionalGeneration, so stock vLLM may load it with the wrong model code instead of failing. pip install -e . in the TeleOCR repo registers the plugin; it pins vLLM 0.11.x, torch 2.8 and transformers 4.57 and needs a CUDA GPU. The weights load with trust_remote_code, so pin a model revision.

  • The task prompts are load-bearing — do not rewrite them. TeleOCR is trained on fixed strings (TELEOCR_PROMPTS); a paraphrase, an extra instruction, or @molecule/api-ocr-llm's generic OCR prompt is off-distribution. The distortedLayout prompt starts with a newline on purpose.

  • Chinese and English only. language is ignored — there is no slot for a language hint in the prompt, and appending one hurts accuracy. For other scripts bond @molecule/api-ocr-tesseract or @molecule/api-ocr-llm.

  • recognize() returns plain text only. Tables in the table / scientificFigure tasks come back as OTSL tokens (<fcel>, <nl>, …) — never show those to users; pass them through convertOtslToHtml(). Formulas come back as LaTeX; formatFormula() wraps them in $$…$$.

  • The layout tasks need the image resized to 1036×1036 first, and their raw output format is not documented upstream — do not parse it on a guess.

  • One image per call, no confidence, no PDFs. Rasterize PDF pages yourself; pages[0].confidence is always unset.

  • Output is capped at 4096 tokens by default (the reference budget). A page that fills it throws TeleOcrTruncatedError (with partialText) rather than returning cut-off text; raise maxOutputTokens within the server's --max-model-len, or split the image.

  • The served model name defaults to TELEOCR_MODEL, then StarDoc-AI/TeleOCR; it must match the name vllm serve was started with. The endpoint and any --api-key belong to the injected AI provider (LOCAL_AI_BASE_URL, LOCAL_AI_API_KEY), not to this package.