@molecule/api-ocr

Core interface · ocr · API (Node) · v1.1.0 · Apache-2.0

Abstract OCR (text extraction from images) interface — bond a vision model, Tesseract, or molecule.dev's hosted service

npm install @molecule/api-ocr

npm · Source on GitHub

How it works

@molecule/api-ocr is the ocr core interface on the API (Node) side: the API your app calls, with no vendor inside.

Choose the implementation by bonding one of its 4 providers: @molecule/api-ocr-llm, @molecule/api-ocr-molecule, @molecule/api-ocr-teleocr, @molecule/api-ocr-tesseract.

import { setProvider, requireProvider } from '@molecule/api-ocr'
import { provider as ocr } from '@molecule/api-ocr-molecule' // or -llm / -tesseract

setProvider(ocr)

// Use anywhere in the app
const result = await requireProvider().recognize({
  data: new Uint8Array(imageBytes),
  mimeType: 'image/png',
})
if (result.text.trim()) {
  console.log('Recognized:', result.text)
}

Providers (4): @molecule/api-ocr-llm, @molecule/api-ocr-molecule, @molecule/api-ocr-teleocr, @molecule/api-ocr-tesseract

Works with: @molecule/api-bond, @molecule/api-i18n

Reference

OCR core interface for molecule.dev.

Defines the abstract contract for extracting text from images. Bond a concrete provider to enable OCR in your application — a vision language model (@molecule/api-ocr-llm), self-hosted Tesseract (@molecule/api-ocr-tesseract), or molecule.dev's hosted service (@molecule/api-ocr-molecule).

Quick Start

import { setProvider, requireProvider } from '@molecule/api-ocr'
import { provider as ocr } from '@molecule/api-ocr-molecule' // or -llm / -tesseract

setProvider(ocr)

// Use anywhere in the app
const result = await requireProvider().recognize({
  data: new Uint8Array(imageBytes),
  mimeType: 'image/png',
})
if (result.text.trim()) {
  console.log('Recognized:', result.text)
}

Type

core

Installation

npm install @molecule/api-ocr @molecule/api-bond @molecule/api-i18n

API

Interfaces

OcrConfig

Configuration for the OCR provider.

interface OcrConfig {
  /** Default language hint when a call does not pass one. */
  language?: string
}

OcrInput

An image to recognize text in.

interface OcrInput {
  /** The image bytes. */
  data: Uint8Array
  /** The image's MIME type, e.g. `image/png`. */
  mimeType: string
}

OcrOptions

Options for one OCR pass.

interface OcrOptions {
  /**
   * Language hint. Providers interpret the code their own way — a vision
   * model takes any language name or BCP-47 tag (`de`, `zh-TW`), Tesseract
   * takes its traineddata codes (`eng`, `deu`). Omit for auto-detection.
   */
  language?: string
}

OcrPage

Text recognized on one page.

interface OcrPage {
  /** 1-based page number (a single-image input always yields page 1). */
  pageNumber: number
  /** The text recognized on this page. */
  text: string
  /** The provider's mean confidence for this page, 0–1. Absent when the provider cannot score itself. */
  confidence?: number
}

OcrProvider

OCR provider interface.

Implement this interface in a bond package to extract text from images — with a vision language model (@molecule/api-ocr-llm), self-hosted Tesseract (@molecule/api-ocr-tesseract), or a hosted service (@molecule/api-ocr-molecule).

interface OcrProvider {
  /** Provider name (e.g. 'llm', 'tesseract', 'molecule'). */
  readonly name: string

  /**
   * Recognize text in one image.
   *
   * @param input - The image bytes and MIME type.
   * @param options - Language hint.
   * @returns The recognized text, per page.
   */
  recognize(input: OcrInput, options?: OcrOptions): Promise<OcrResult>

  /**
   * Release provider resources (e.g. Tesseract's worker threads). Optional —
   * only providers that hold resources outside the JS heap implement it.
   */
  dispose?(): Promise<void>
}

OcrResult

Result of recognizing text in one image.

interface OcrResult {
  /** All recognized text, pages joined with a blank line. */
  text: string
  /** Per-page breakdown (single-image input yields exactly one page). */
  pages: OcrPage[]
}

Functions

getProvider()

Retrieves the bonded OCR provider, or null if none is bonded.

function getProvider(): OcrProvider | null

Returns: The bonded provider, or null.

hasProvider()

Checks whether an OCR provider is currently bonded.

function hasProvider(): boolean

Returns: true if a provider is bonded.

requireProvider()

Retrieves the bonded OCR provider, throwing if none is bonded. Use this when OCR functionality is required.

function requireProvider(): OcrProvider

Returns: The bonded OCR provider.

setProvider(provider)

Registers an OCR provider.

function setProvider(provider: OcrProvider): void
  • provider — The OCR provider to bond.

Available Providers

ProviderPackage
LLM (vision model)@molecule/api-ocr-llm
Molecule (hosted)@molecule/api-ocr-molecule
TeleOCR (self-hosted)@molecule/api-ocr-teleocr
Tesseract.js@molecule/api-ocr-tesseract

Injection Notes

Requirements

Peer dependencies:

  • @molecule/api-bond ^1.0.1
  • @molecule/api-i18n ^1.0.1

Runtime Dependencies

  • @molecule/api-bond

  • @molecule/api-i18n

  • Like most cores there are NO module-level convenience delegates. Call methods on requireProvider() (throws when unbonded). Note getProvider() returns null rather than throwing.

  • language is a HINT, and providers interpret it differently. A vision model takes any language name or BCP-47 tag; Tesseract needs its own traineddata codes (eng, not en). Passing Tesseract a BCP-47 tag fails at recognition time — use the code the bonded provider documents.

  • One image per call. Multi-page documents (PDF, multi-page TIFF) are not in this contract yet — recognize a rendered page image at a time and join the results yourself.

  • Image size is unbounded here, bounded by the provider. The core never resizes or re-encodes; each bond documents (and enforces) its own maximum. Downscale server-side before recognizing a scan you only need the words of.

  • confidence is the provider scoring ITSELF. A vision model returns no confidence at all — do not gate on it unless your bond fills it (Tesseract does), and never treat a high score as proof the words are right.

  • OCR output is untrusted input. Recognized text came from an image anyone could have crafted — treat it like user content: escape before rendering, moderate before publishing.