@molecule/api-ocr-teleocr
Provider bond · ocr · API (Node) · v1.0.0 · Apache-2.0
TeleOCR provider for molecule.dev — extracts text from images with the open TeleOCR document-parsing model served by vLLM
npm install @molecule/api-ocr-teleocrnpm · Source on GitHub · Implements @molecule/api-ocr
How it works
@molecule/api-ocr-teleocr is a provider bond on the API (Node) side: it implements the ocr core interface (@molecule/api-ocr) with a concrete vendor or library behind it.
Your code calls the core; you wire this provider once at startup. Swapping vendors later is one line in that wiring, not a rewrite.
import { createProvider as createLocalAi } from '@molecule/api-ai-local'
import { requireProvider, setProvider } from '@molecule/api-ocr'
import { createProvider } from '@molecule/api-ocr-teleocr'
// vllm serve StarDoc-AI/TeleOCR (with the TeleOCR_vllm plugin installed)
const ai = createLocalAi({ baseUrl: 'http://localhost:8000/v1' })
setProvider(createProvider({ ai }))
const result = await requireProvider().recognize({
data: new Uint8Array(imageBytes),
mimeType: 'image/png',
})
console.log(result.text)Works with: @molecule/api-ai, @molecule/api-ocr
Secrets: TELEOCR_MODEL (optional)
Reference
TeleOCR provider for molecule.dev.
Implements the @molecule/api-ocr contract with TeleOCR, an open
(Apache-2.0 weights), ~1.2B-parameter vision-language model built for
document parsing — scanned and photographed pages, tables, formulas, code.
It runs on your own GPU behind vLLM's OpenAI-compatible server, reached
through @molecule/api-ai-local, so there is no per-call vendor cost.
Quick Start
import { createProvider as createLocalAi } from '@molecule/api-ai-local'
import { requireProvider, setProvider } from '@molecule/api-ocr'
import { createProvider } from '@molecule/api-ocr-teleocr'
// vllm serve StarDoc-AI/TeleOCR (with the TeleOCR_vllm plugin installed)
const ai = createLocalAi({ baseUrl: 'http://localhost:8000/v1' })
setProvider(createProvider({ ai }))
const result = await requireProvider().recognize({
data: new Uint8Array(imageBytes),
mimeType: 'image/png',
})
console.log(result.text)
Type
provider
Installation
npm install @molecule/api-ocr-teleocr @molecule/api-ai @molecule/api-ocr
API
Interfaces
TeleOcrConfig
Options for createProvider. ai falls back to the bonded ai
provider; model falls back to TELEOCR_MODEL, then DEFAULT_TELEOCR_MODEL.
interface TeleOcrConfig {
/**
* The AI provider that reaches the TeleOCR server. Defaults to the provider
* bonded as `ai` (`@molecule/api-ai` `requireProvider()`) — typically
* `@molecule/api-ai-local` pointed at the vLLM server
* (`LOCAL_AI_BASE_URL=http://localhost:8000/v1`).
*/
ai?: AIProvider
/**
* The served model name (what `vllm serve` was started with). Defaults to
* the `TELEOCR_MODEL` env var, read on each call, then `'StarDoc-AI/TeleOCR'`.
*/
model?: string
/**
* Output token ceiling per call. Defaults to 4096, the reference
* implementation's `max_new_tokens`. A page that hits it throws
* `TeleOcrTruncatedError` instead of returning cut-off text.
*/
maxOutputTokens?: number
}
TeleOcrEnv
Environment variables the provider reads (each on every call, never at import).
interface TeleOcrEnv {
/** The served model name; overrides `DEFAULT_TELEOCR_MODEL`. */
TELEOCR_MODEL?: string
}
Types
TeleOcrTask
One of TeleOCR's tasks — a key of TELEOCR_PROMPTS.
type TeleOcrTask = keyof typeof TELEOCR_PROMPTS
Classes
TeleOcrTruncatedError
The model used its whole output budget, so the transcription is cut off.
Carries the partial text so a caller can keep it knowingly — never returned
silently as if it were the full page. Dense pages hit the reference budget of
4096 tokens; raise maxOutputTokens (within the server's --max-model-len)
or split the image.
Functions
convertOtslToHtml(otsl)
Converts TeleOCR's OTSL table output (<fcel>, <nl>, … tokens) into an
HTML <table> with rowspan/colspan for merged cells. Cell text is
HTML-escaped. Output that is already a <table>…</table> is returned as-is.
function convertOtslToHtml(otsl: string): string
otsl— The model's raw answer to thetableorscientificFigureprompt.
Returns: An HTML table, or '' when the input holds no cells.
createProvider(config)
Creates a TeleOCR provider.
function createProvider(config?: TeleOcrConfig): OcrProvider
config— Which AI provider reaches the TeleOCR server, and the served model name.
Returns: An OcrProvider backed by TeleOCR.
formatFormula(latex)
Normalizes TeleOCR's formula answer the way the reference does: strips a
surrounding \[ … \] and wraps the LaTeX in $$ … $$ unless it is already
$-delimited.
function formatFormula(latex: string): string
latex— The model's raw answer to theformulaprompt.
Returns: Display-math LaTeX.
Constants
DEFAULT_TELEOCR_MAX_OUTPUT_TOKENS
The reference implementation's output budget (max_new_tokens).
const DEFAULT_TELEOCR_MAX_OUTPUT_TOKENS: 4096
DEFAULT_TELEOCR_MODEL
The served model name used when neither config nor TELEOCR_MODEL sets one.
const DEFAULT_TELEOCR_MODEL: 'StarDoc-AI/TeleOCR'
OTSL_TOKENS
Every OTSL token, in the order the reference splits on them.
const OTSL_TOKENS: readonly string[]
provider
The provider implementation.
const provider: OcrProvider
TELEOCR_LAYOUT_IMAGE_SIZE
Side length, in pixels, the layout tasks expect the image resized to.
const TELEOCR_LAYOUT_IMAGE_SIZE: 1036
TELEOCR_PROMPTS
The task prompts. text is what recognize() sends; the rest are exported
for callers that drive the model directly through @molecule/api-ai.
text→ plain text.table→ OTSL tokens; convert withconvertOtslToHtml.formula→ LaTeX; normalize withformatFormula.code→ the code snippet's text.layout/distortedLayout→ layout blocks. The image must first be resized to 1036×1036 (bicubic), and the raw serialization of the blocks is not documented upstream — parse it only after checking it yourself.scientificFigure→ the table a chart implies, as OTSL.
const TELEOCR_PROMPTS: {
readonly text: 'Please output the text content from the image.'
readonly table: 'This is the image of a table. Please output the table in OTSL format.'
readonly formula: 'Please write out the expression of the formula in the image using LaTeX format.'
readonly code: 'The image contains a code snippet, please output the parsing result.'
readonly layout: 'Analyze the image layout.'
readonly distortedLayout: '\nMulti-point Layout Segmentation Analysis.'
readonly scientificFigure: 'This is a scientific figure. Please extract the table implied by this figure.'
}
TELEOCR_SYSTEM_PROMPT
The system prompt the reference implementation sends with every task.
const TELEOCR_SYSTEM_PROMPT: 'You are a helpful assistant.'
Core Interface
Implements @molecule/api-ocr interface.
Bond Wiring
Setup function to register this provider with the core interface:
import { setProvider } from '@molecule/api-ocr'
import { provider } from '@molecule/api-ocr-teleocr'
export function setupOcrTeleocr(): void {
setProvider(provider)
}
Injection Notes
Requirements
Peer dependencies:
@molecule/api-ai>=1.0.0@molecule/api-ocr>=1.0.0
Environment Variables
TELEOCR_MODEL(optional) — TeleOCR served model name- Setup: The model name your vLLM server was started with (vllm serve <name>, with the TeleOCR_vllm plugin from github.com/caipeng328/TeleOCR installed). The server URL is set on the AI provider this bond uses, e.g. LOCAL_AI_BASE_URL for @molecule/api-ai-local. Defaults to StarDoc-AI/TeleOCR.
- Get it here: https://huggingface.co/XingChen-AGI/TeleOCR
- Example:
StarDoc-AI/TeleOCR
Runtime Dependencies
-
@molecule/api-ai -
@molecule/api-ocr -
Serve it with the
TeleOCR_vllmplugin, never stock vLLM alone. The model reuses the architecture nameQwen2_5_VLForConditionalGeneration, so stock vLLM may load it with the wrong model code instead of failing.pip install -e .in the TeleOCR repo registers the plugin; it pins vLLM 0.11.x, torch 2.8 and transformers 4.57 and needs a CUDA GPU. The weights load withtrust_remote_code, so pin a model revision. -
The task prompts are load-bearing — do not rewrite them. TeleOCR is trained on fixed strings (
TELEOCR_PROMPTS); a paraphrase, an extra instruction, or@molecule/api-ocr-llm's generic OCR prompt is off-distribution. ThedistortedLayoutprompt starts with a newline on purpose. -
Chinese and English only.
languageis ignored — there is no slot for a language hint in the prompt, and appending one hurts accuracy. For other scripts bond@molecule/api-ocr-tesseractor@molecule/api-ocr-llm. -
recognize()returns plain text only. Tables in thetable/scientificFiguretasks come back as OTSL tokens (<fcel>,<nl>, …) — never show those to users; pass them throughconvertOtslToHtml(). Formulas come back as LaTeX;formatFormula()wraps them in$$…$$. -
The layout tasks need the image resized to 1036×1036 first, and their raw output format is not documented upstream — do not parse it on a guess.
-
One image per call, no confidence, no PDFs. Rasterize PDF pages yourself;
pages[0].confidenceis always unset. -
Output is capped at 4096 tokens by default (the reference budget). A page that fills it throws
TeleOcrTruncatedError(withpartialText) rather than returning cut-off text; raisemaxOutputTokenswithin the server's--max-model-len, or split the image. -
The served model name defaults to
TELEOCR_MODEL, thenStarDoc-AI/TeleOCR; it must match the namevllm servewas started with. The endpoint and any--api-keybelong to the injected AI provider (LOCAL_AI_BASE_URL,LOCAL_AI_API_KEY), not to this package.