Drag OCR, layout-detection, table-extraction, and structuring models onto a visual canvas. Wire their typed inputs and outputs together, and run the whole flow — interactively or headless via the API.
OCRFlow wraps best-in-class open models behind a typed, composable interface you own and run on your own infrastructure. The guiding idea is canvas ↔ API parity: anything you build on the canvas runs the same way headless via the API/SDK — so the pipeline you prototype visually is the same one that ships to production.
A composition layer over inference — not another wrapper. Typed, atomic, and designed to run entirely on your own hardware.
Compose OCR stages as a node graph powered by React Flow — drag, drop, and wire.
Nodes only connect when an output type satisfies the next input — invalid pipelines are caught before they run.
One HTTP endpoint = one task. Each Docling/Surya task gets its own API, schemas, validation, and tests.
Model runners wrap upstream inference (Docling, Surya, …) instead of reimplementing it.
Weights load on first request (or explicit warm-up) and are shared as a per-process singleton.
The default model catalog is Apache / MIT-licensed and runs fully on-prem. No data leaves your box.
Everything you build on the canvas is runnable headless via the API/SDK. Prototype visually, ship the exact same graph to production — no rewrite, no drift between what you designed and what runs.
A FastAPI service exposing one endpoint per model task, driven by a canvas built on React Flow. Long jobs go to Celery; state lives in PostgreSQL.
A Next.js App Router app where the pipeline canvas is built on React Flow — styled with Tailwind v4 and shadcn/ui, animated with Framer Motion.
A model registry describes each model; a ModelRunner protocol (load / run / health) standardizes inference behind a cached, lazy-loaded runner.
Python 3.11+, Node 20+, PostgreSQL 14+, and Redis 6+. Then clone, boot the backend, and start the canvas.
Grab the source and drop into the project root.
Create a venv, install deps, run alembic upgrade head, then start Uvicorn + a Celery worker.
Install frontend deps, point NEXT_PUBLIC_API_URL at the backend, and run the dev server.
$ git clone https://github.com/baselhusam/OCRFlow.git $ cd OCRFlow
$ cd backend $ python -m venv .venv && source .venv/bin/activate $ pip install -r requirements.txt # + requirements-docling.txt / -surya.txt $ cp .env.example .env # set DATABASE_URL, REDIS_URL, secrets $ alembic upgrade head # run migrations $ uvicorn app.main:app --reload # API → localhost:8000 # in a second terminal — background jobs $ celery -A app.worker worker --loglevel=info
$ cd frontend $ npm install $ cp .env.local.example .env.local # point NEXT_PUBLIC_API_URL at backend $ npm run dev # app → localhost:3000
ℹ️ Exact module paths (app.main, app.worker) may differ — check backend/app/ and backend/docs/ for current entrypoints.
Each model is implemented and validated on its own — with its own API, schemas, validation, and tests — before the next is added.
Typed model registry, the ModelRunner protocol, canvas scaffolding, and the atomic-task API pattern.
First runner online — OCR, layout, and table tasks wired end-to-end from canvas to API.
Per-task Surya runners — detection, recognition, and layout — added one endpoint at a time.
Additional structuring and extraction models, plus SDK ergonomics for headless runs.
Clone the repo, run it on your own hardware, and start composing. Contributions and issues are welcome.