@molecule/api-ocr
Core interface · ocr · API (Node) · v1.1.0 · Apache-2.0
Abstract OCR (text extraction from images) interface — bond a vision model, Tesseract, or molecule.dev's hosted service
npm install @molecule/api-ocrHow it works
@molecule/api-ocr is the ocr core interface on the API (Node) side: the API your app calls, with no vendor inside.
Choose the implementation by bonding one of its 4 providers: @molecule/api-ocr-llm, @molecule/api-ocr-molecule, @molecule/api-ocr-teleocr, @molecule/api-ocr-tesseract.
import { setProvider, requireProvider } from '@molecule/api-ocr'
import { provider as ocr } from '@molecule/api-ocr-molecule' // or -llm / -tesseract
setProvider(ocr)
// Use anywhere in the app
const result = await requireProvider().recognize({
data: new Uint8Array(imageBytes),
mimeType: 'image/png',
})
if (result.text.trim()) {
console.log('Recognized:', result.text)
}Providers (4): @molecule/api-ocr-llm, @molecule/api-ocr-molecule, @molecule/api-ocr-teleocr, @molecule/api-ocr-tesseract
Works with: @molecule/api-bond, @molecule/api-i18n
Reference
OCR core interface for molecule.dev.
Defines the abstract contract for extracting text from images. Bond a
concrete provider to enable OCR in your application — a vision language
model (@molecule/api-ocr-llm), self-hosted Tesseract
(@molecule/api-ocr-tesseract), or molecule.dev's hosted service
(@molecule/api-ocr-molecule).
Quick Start
import { setProvider, requireProvider } from '@molecule/api-ocr'
import { provider as ocr } from '@molecule/api-ocr-molecule' // or -llm / -tesseract
setProvider(ocr)
// Use anywhere in the app
const result = await requireProvider().recognize({
data: new Uint8Array(imageBytes),
mimeType: 'image/png',
})
if (result.text.trim()) {
console.log('Recognized:', result.text)
}
Type
core
Installation
npm install @molecule/api-ocr @molecule/api-bond @molecule/api-i18n
API
Interfaces
OcrConfig
Configuration for the OCR provider.
interface OcrConfig {
/** Default language hint when a call does not pass one. */
language?: string
}
OcrInput
An image to recognize text in.
interface OcrInput {
/** The image bytes. */
data: Uint8Array
/** The image's MIME type, e.g. `image/png`. */
mimeType: string
}
OcrOptions
Options for one OCR pass.
interface OcrOptions {
/**
* Language hint. Providers interpret the code their own way — a vision
* model takes any language name or BCP-47 tag (`de`, `zh-TW`), Tesseract
* takes its traineddata codes (`eng`, `deu`). Omit for auto-detection.
*/
language?: string
}
OcrPage
Text recognized on one page.
interface OcrPage {
/** 1-based page number (a single-image input always yields page 1). */
pageNumber: number
/** The text recognized on this page. */
text: string
/** The provider's mean confidence for this page, 0–1. Absent when the provider cannot score itself. */
confidence?: number
}
OcrProvider
OCR provider interface.
Implement this interface in a bond package to extract text from images —
with a vision language model (@molecule/api-ocr-llm), self-hosted
Tesseract (@molecule/api-ocr-tesseract), or a hosted service
(@molecule/api-ocr-molecule).
interface OcrProvider {
/** Provider name (e.g. 'llm', 'tesseract', 'molecule'). */
readonly name: string
/**
* Recognize text in one image.
*
* @param input - The image bytes and MIME type.
* @param options - Language hint.
* @returns The recognized text, per page.
*/
recognize(input: OcrInput, options?: OcrOptions): Promise<OcrResult>
/**
* Release provider resources (e.g. Tesseract's worker threads). Optional —
* only providers that hold resources outside the JS heap implement it.
*/
dispose?(): Promise<void>
}
OcrResult
Result of recognizing text in one image.
interface OcrResult {
/** All recognized text, pages joined with a blank line. */
text: string
/** Per-page breakdown (single-image input yields exactly one page). */
pages: OcrPage[]
}
Functions
getProvider()
Retrieves the bonded OCR provider, or null if none is bonded.
function getProvider(): OcrProvider | null
Returns: The bonded provider, or null.
hasProvider()
Checks whether an OCR provider is currently bonded.
function hasProvider(): boolean
Returns: true if a provider is bonded.
requireProvider()
Retrieves the bonded OCR provider, throwing if none is bonded. Use this when OCR functionality is required.
function requireProvider(): OcrProvider
Returns: The bonded OCR provider.
setProvider(provider)
Registers an OCR provider.
function setProvider(provider: OcrProvider): void
provider— The OCR provider to bond.
Available Providers
| Provider | Package |
|---|---|
| LLM (vision model) | @molecule/api-ocr-llm |
| Molecule (hosted) | @molecule/api-ocr-molecule |
| TeleOCR (self-hosted) | @molecule/api-ocr-teleocr |
| Tesseract.js | @molecule/api-ocr-tesseract |
Injection Notes
Requirements
Peer dependencies:
@molecule/api-bond^1.0.1@molecule/api-i18n^1.0.1
Runtime Dependencies
-
@molecule/api-bond -
@molecule/api-i18n -
Like most cores there are NO module-level convenience delegates. Call methods on
requireProvider()(throws when unbonded). NotegetProvider()returnsnullrather than throwing. -
languageis a HINT, and providers interpret it differently. A vision model takes any language name or BCP-47 tag; Tesseract needs its own traineddata codes (eng, noten). Passing Tesseract a BCP-47 tag fails at recognition time — use the code the bonded provider documents. -
One image per call. Multi-page documents (PDF, multi-page TIFF) are not in this contract yet — recognize a rendered page image at a time and join the results yourself.
-
Image size is unbounded here, bounded by the provider. The core never resizes or re-encodes; each bond documents (and enforces) its own maximum. Downscale server-side before recognizing a scan you only need the words of.
-
confidenceis the provider scoring ITSELF. A vision model returns no confidence at all — do not gate on it unless your bond fills it (Tesseract does), and never treat a high score as proof the words are right. -
OCR output is untrusted input. Recognized text came from an image anyone could have crafted — treat it like user content: escape before rendering, moderate before publishing.