Self-host first · Open source · MIT

Composable OCR pipelines, fully under your control.

Drag OCR, layout-detection, table-extraction, and structuring models onto a visual canvas. Wire their typed inputs and outputs together, and run the whole flow — interactively or headless via the API.

Canvas ↔ API parity Typed wires Docling & Surya runners
Built on Docling Surya FastAPI React Flow Next.js 16 Celery + Redis PostgreSQL

Most document-AI tooling is either a closed SaaS or a pile of one-off scripts.

OCRFlow wraps best-in-class open models behind a typed, composable interface you own and run on your own infrastructure. The guiding idea is canvas ↔ API parity: anything you build on the canvas runs the same way headless via the API/SDK — so the pipeline you prototype visually is the same one that ships to production.

Everything a real pipeline needs

A composition layer over inference — not another wrapper. Typed, atomic, and designed to run entirely on your own hardware.

01

Visual pipeline builder

Compose OCR stages as a node graph powered by React Flow — drag, drop, and wire.

02

Typed wires

Nodes only connect when an output type satisfies the next input — invalid pipelines are caught before they run.

03

Atomic tasks

One HTTP endpoint = one task. Each Docling/Surya task gets its own API, schemas, validation, and tests.

04

Adapter, don't fork

Model runners wrap upstream inference (Docling, Surya, …) instead of reimplementing it.

05

Lazy model loading

Weights load on first request (or explicit warm-up) and are shared as a per-process singleton.

06

Self-host first

The default model catalog is Apache / MIT-licensed and runs fully on-prem. No data leaves your box.

Canvas ↔ API parity

Everything you build on the canvas is runnable headless via the API/SDK. Prototype visually, ship the exact same graph to production — no rewrite, no drift between what you designed and what runs.

A typed backend, a canvas frontend

A FastAPI service exposing one endpoint per model task, driven by a canvas built on React Flow. Long jobs go to Celery; state lives in PostgreSQL.

Frontend

The canvas

A Next.js App Router app where the pipeline canvas is built on React Flow — styled with Tailwind v4 and shadcn/ui, animated with Framer Motion.

FRAMEWORKNext.js 16 · React 19 · TypeScript
CANVASReact Flow (@xyflow/react)
UITailwind v4 · shadcn/ui · Framer Motion
Backend

The engine

A model registry describes each model; a ModelRunner protocol (load / run / health) standardizes inference behind a cached, lazy-loaded runner.

APIFastAPI · Pydantic · JWT auth
JOBSCelery workers · Redis broker
DATAPostgreSQL · async SQLAlchemy 2 · Alembic
Canvas
React Flow
→
REST API
FastAPI
→
Runners
Docling · Surya
→
Workers + DB
Celery · Postgres

Run it locally in three steps

Python 3.11+, Node 20+, PostgreSQL 14+, and Redis 6+. Then clone, boot the backend, and start the canvas.

1

Clone the repo

Grab the source and drop into the project root.

2

Boot the backend

Create a venv, install deps, run alembic upgrade head, then start Uvicorn + a Celery worker.

3

Start the canvas

Install frontend deps, point NEXT_PUBLIC_API_URL at the backend, and run the dev server.

$ git clone https://github.com/baselhusam/OCRFlow.git
$ cd OCRFlow
$ cd backend
$ python -m venv .venv && source .venv/bin/activate
$ pip install -r requirements.txt   # + requirements-docling.txt / -surya.txt
$ cp .env.example .env             # set DATABASE_URL, REDIS_URL, secrets
$ alembic upgrade head             # run migrations
$ uvicorn app.main:app --reload   # API → localhost:8000

# in a second terminal — background jobs
$ celery -A app.worker worker --loglevel=info
$ cd frontend
$ npm install
$ cp .env.local.example .env.local  # point NEXT_PUBLIC_API_URL at backend
$ npm run dev                      # app → localhost:3000

ℹ️ Exact module paths (app.main, app.worker) may differ — check backend/app/ and backend/docs/ for current entrypoints.

Models land one at a time

Each model is implemented and validated on its own — with its own API, schemas, validation, and tests — before the next is added.

Foundations

Shipped

Typed model registry, the ModelRunner protocol, canvas scaffolding, and the atomic-task API pattern.

Docling

In progress

First runner online — OCR, layout, and table tasks wired end-to-end from canvas to API.

Surya

Next

Per-task Surya runners — detection, recognition, and layout — added one endpoint at a time.

The rest of the catalog

Planned

Additional structuring and extraction models, plus SDK ergonomics for headless runs.

Own your document pipeline.

Clone the repo, run it on your own hardware, and start composing. Contributions and issues are welcome.