Shadow Swarm & LeanDoc-1G v11 — Hybrid OCR

An OCR-capable document pipeline that keeps native parsing for digital text. V11 combines LeanDoc's PDF/DOCX ingestion with separately installed pretrained Tesseract OCR for image-only PDF pages, then reconstructs layout and tables from recognized word coordinates. It runs on CPU.

This repository contains software, tests, benchmark labels/results and synthetic DOCX fixtures. It does not contain a newly trained Shadow Swarm neural checkpoint. Pixel recognition uses Tesseract's pretrained language data, installed separately. The tested English traineddata file is approximately 4 MiB. No model weights or tokenizer are bundled, and there is no Transformers from_pretrained interface. The Python package retains version 0.3.0.dev1; v11 identifies this release/experiment series.

Digital-PDF grouped-table update

V11 now includes an opt-in native geometry repair for leading row groups: stream_ruled_pdf(path, grouped=True). Read expanded group associations through block.table_data.logical_rows. This adds no weights or dependencies; default PDF, OCR and DOCX routing remains unchanged.

On 49 selected digital-PDF rows, the original strict score stays 37/49. A separate, post-hoc semantic diagnostic improves 44/49 → 49/49 after repairing five group associations. This is not 100% general document accuracy. The diagnostic explicitly accepts listed source-label aliases and exact means within uncertainty cells. Saved Docling outputs score 34/49 under those rules.

The 26-page run took 2.937 seconds versus 2.916 before, with approximately 156 MiB peak process-tree RSS in both. A 527-page run took 53.347 seconds and peaked at 268.63 MiB; observed peaks vary between runs. Source and citation checks remained unchanged, and all 86 tests pass. Only the intended table changed in that corpus; other pages have no new accuracy claim.

See the full grouped-table report for before/after results, limitations, provenance checks and reproduction steps.

What improved in OCR

The prior local OCR preview sometimes attached a prose note below a scanned table to its final row. V11 keeps that note outside the table while preserving the recognized text and evidence offsets. Native digital-PDF and DOCX behavior is retained. The change adds no runtime dependencies or neural weights.

Scanned-table check Previous local OCR preview V11
Original two pages: complete rows 37/40 39/40
Original two pages: cells 208/220 219/220
Additional two pages: complete rows 38/40 40/40
Additional two pages: cells 209/220 220/220

The additional pages were first scored after freezing the rule but use the same synthetic templates. These results do not establish general scanned-table accuracy. One original numeric value remains misrecognized. The improvement is in table structure; recognition itself is unchanged. The previous public v10 repository had native DOCX/PDF parsing but no OCR path; the “previous preview” column refers to the intervening local OCR experiment.

General comparison

The full comparison report covers the same 46 input cases for the previous OCR preview, v11, standalone Tesseract and official Docling 2.126.0. It reports each of these separately:

  • 26 digital PDF pages: selected text, headings and complete table rows.
  • Ten clean image-only PDF pages: selected excerpt word-error rate.
  • Two synthetically degraded scan pages: selected excerpt word-error rate.
  • Four scanned table pages: 80 data rows and 440 cells.
  • One two-page mixed scan/digital PDF: routing and output checks.
  • Three DOCX documents: 180 rows, 960 cells, headings and paragraphs.
  • Conversion time and sampled process-tree resident RAM by input category.

Read the report's measurement scopes before comparing RAM or throughput. Each engine runs separately and sequentially. Models retained from previous categories remain in the measured process; Docling retains digital/OCR pipeline instances. These are harness footprints, not isolated minimum memory requirements for each format. The first relevant conversion can include lazy model loading. In the broad run, v11 clean scans averaged 1.049 seconds per page, versus 1.051 for the previous preview and 1.027 for standalone Tesseract. V11's peak in that category was 135 MiB, including memory retained after digital-PDF cases. V11 and Tesseract both had 7/578 selected clean-scan word errors; Docling had 9/578 with the tested configuration. Native DOCX retained 960/960 cells, matching Docling; digital-PDF selected rows remained 37/49, versus Docling's 29/49.

Docling's scanned-table results were strongly configuration-dependent: the main Tesseract/scale-3 setup scored 0/80 complete rows, while a separate full-page/scale-2 diagnostic recovered 34/80. V11 recovered 79/80. These small synthetic tests do not establish Docling's best achievable accuracy; other Docling OCR backends were not tested. Both configurations are disclosed in the full report rather than presenting the worst configuration alone.

Standalone Tesseract has no direct DOCX or semantic table/heading parser; those metrics are N/A rather than misleading zero-accuracy scores.

The preceding low-resolution OCR pilot measured about 1.05 seconds per clean scanned page and 101 MiB peak process-tree RSS, with 7/578 selected word errors (1.21% WER). The broad v11 report provides updated same-input results; the pilot is a historical baseline, not a universal performance guarantee. All scan evaluations here are synthetic rasterizations, not real camera photos.

Installation and usage

Requires Python 3.10+, Poppler utilities and Tesseract with the desired language data. For example, install the OS packages poppler-utils, tesseract-ocr and tesseract-ocr-eng on Ubuntu/Debian. No GPU is required.

hf download psikosen/shadowswarm-leandoc-v11 --local-dir shadowswarm-leandoc-v11
cd shadowswarm-leandoc-v11
python -m pip install -e '.[test]'
python -m pytest -q

Use the explicit hybrid entry point for PDFs that may contain scans:

from swarmkg.leandoc.ocr_experiment import stream_hybrid_pdf

for result in stream_hybrid_pdf("mixed.pdf"):
    print(result.route)  # "native" or "ocr"
    print(result.page.markdown)
    for span in result.page.evidence_spans:
        print(span.extraction_method, span.quoted_text)

Use native ingestion for DOCX, Markdown, text or ordinary digital PDFs:

from swarmkg.ingest.docling_adapter import DoclingIngestionAdapter

adapter = DoclingIngestionAdapter()
for chunk in adapter.stream_chunks("benchmarks/docx_fixtures/ledger.docx"):
    print(chunk.text, chunk.table_rows)

The historical DoclingIngestionAdapter name does not mean official Docling is required. Our default adapter remains native; select stream_hybrid_pdf explicitly for OCR. The public native DOCX fixtures and scorer are runnable without Poppler/Tesseract. Optional official-Docling comparisons require a separate installation and are not runtime dependencies.

The OCR function accepts tesseract, language, env, max_pages, max_dimension, dpi, force_ocr and trim_table_notes. Defaults are English, one Tesseract CPU thread and a 2,400-pixel longest raster edge. Poppler's -scale-to controls raster dimensions; effective DPI varies with physical page size. dpi supplies the coordinate conversion scale, not a guarantee of exact rendered DPI. Higher resolution is not automatically better: the pilot's 3,600-pixel profile slightly improved selected text but harmed these tables.

Supported scope and limitations

  • Automatic OCR targets pages with fewer than 60 embedded-text characters and at least one raster image. Substantial digital text mixed with a scanned region can bypass OCR; region-level detection/deduplication and bad existing OCR-layer detection are not implemented. force_ocr=True forces page OCR.
  • Input to the hybrid path is PDF. Photos embedded in PDFs can be read, but direct JPEG/PNG ingestion, perspective correction, handwriting and real camera-photo accuracy are not established. The tested language is English.
  • The trailing-note rule treats content well below the final numeric anchor as prose. A legitimately wrapped last cell can require a stronger boundary detector. It is enabled only for OCR table reconstruction.
  • Processing is sequential with one temporary grayscale page image and one Tesseract invocation per scanned page. No persistent OCR worker is claimed. Raster dimensions and subprocess durations are bounded; there is no enforced universal RAM ceiling for arbitrary complex PDFs or very large tables.
  • OCR evidence quotes recognized text, with offsets into the emitted source stream joined by form feeds and a SHA-256 hash of the original PDF. Exact offsets establish internal consistency, not correct recognition of pixels. Tesseract confidence is not a calibrated probability of correctness.
  • Native DOCX reads conventional transitional OOXML body paragraphs/tables, including inherited heading levels and horizontal/vertical merged cells. Nested tables and strict OOXML namespaces are rejected; headers, footers, notes, drawings, automatic list numbering and Word pagination are not rendered. DOCX evidence uses canonical extracted-text offsets and no invented page boxes.
  • Benchmarks use selected development labels and synthetic fixtures. They do not exhaustively score reading order, false-positive precision, multilingual OCR, or every document structure. There is no single overall accuracy score and no general claim of superiority over Docling.

Validation and included artifacts

86 unit/regression tests pass, including DOCX extraction, OCR routing, offset rebasing, stream cleanup and preservation of prose below scanned tables. Historical PDF tests can depend on local PDFs and do not exercise those paths when the files are absent.

PDF benchmark reproduction needs the source documents and locally generated page clips/scans at the recorded paths. Third-party PDFs and full extracted paper outputs are not bundled. The DOCX fixture data is included. This release is a snapshot of the current workspace, including changes beyond an earlier local Git commit.

License and attribution

Maintainer: psikosen. The source workspace specifies no project license grant; public visibility alone does not grant an open-source license. Dependencies and pretrained OCR data retain their respective licenses. Tesseract and Docling are external projects; no affiliation or endorsement is implied. No new OCR training or fine-tuning was performed for v11.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support