yugeshkarunamurthy commited on
Commit
d0fefbc
·
verified ·
1 Parent(s): bd7831a

Create README.md

Browse files
Files changed (1) hide show
  1. README.md +318 -0
README.md ADDED
@@ -0,0 +1,318 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: ATH-MaaS/OvisOCR2
4
+ language:
5
+ - multilingual
6
+ library_name: mlx
7
+ pipeline_tag: image-text-to-text
8
+ tags:
9
+ - ocr
10
+ - document-parsing
11
+ - markdown
12
+ - tables
13
+ - formulas
14
+ - multimodal
15
+ - vision-language
16
+ - qwen3.5
17
+ - mlx
18
+ - apple-silicon
19
+ - oq
20
+ - oqe4
21
+ - quantized
22
+ ---
23
+
24
+ # OvisOCR2-oQe4
25
+
26
+ > **Apple Silicon Optimized oQe4 MLX Quantized Release**
27
+
28
+ This repository contains an **oQe4 mixed-precision MLX quantized** version of **OvisOCR2**, optimized for fast and memory-efficient document understanding on Apple Silicon.
29
+
30
+ The original **OvisOCR2** model was developed by **ATH-MaaS**. This repository provides an optimized **MLX/oQe4 conversion only** and **does not include any additional training or fine-tuning**.
31
+
32
+ Using **oQe4 sensitivity-aware mixed-precision quantization**, this release preserves the excellent OCR and document parsing capabilities of the original model while significantly reducing memory usage and improving inference efficiency on Apple Silicon. OvisOCR2 is a compact 0.8B end-to-end document parser that converts document images directly into structured Markdown and achieves state-of-the-art results on OmniDocBench v1.6 and PureDocBench. 0
33
+
34
+ ---
35
+
36
+ # About OvisOCR2
37
+
38
+ OvisOCR2 is a compact **Vision Language Model (VLM)** specialized for **end-to-end document parsing**.
39
+
40
+ Unlike traditional OCR pipelines that separately detect layouts, recognize text, and reconstruct documents, OvisOCR2 directly converts a document page into structured Markdown while preserving natural reading order.
41
+
42
+ The model is capable of extracting:
43
+
44
+ - 📄 Plain text
45
+ - 📐 Mathematical formulas (LaTeX)
46
+ - 📊 Tables (HTML)
47
+ - 🖼 Images and visual regions
48
+ - 📚 Complex layouts
49
+ - 📰 Multi-column documents
50
+ - 📑 Scientific papers
51
+ - 📋 Forms
52
+ - 📖 Books
53
+ - 📃 Receipts
54
+ - 📜 Scanned documents
55
+
56
+ The model is based on **Qwen3.5-0.8B** and is trained using a combination of:
57
+
58
+ - Real-world document data
59
+ - Synthetic document generation
60
+ - Supervised Fine-Tuning (SFT)
61
+ - Reinforcement Learning (RL)
62
+ - On-Policy Distillation (OPD)
63
+
64
+ OvisOCR2 achieves an **overall score of 96.58 on OmniDocBench v1.6**, becoming the first end-to-end model to top that benchmark, and also leads PureDocBench with an Avg3 score of 75.06. 1
65
+
66
+ ---
67
+
68
+ # Quantization
69
+
70
+ This release uses **oQe4 mixed-precision quantization**.
71
+
72
+ ## Specifications
73
+
74
+ - **Format:** MLX
75
+ - **Quantization:** oQe4
76
+ - **Method:** Sensitivity-Aware Mixed Precision
77
+ - **Target Platform:** Apple Silicon
78
+ - **Inference Engine:** MLX / oMLX
79
+
80
+ Unlike traditional fixed-bit quantization, **oQe4 automatically assigns precision according to layer sensitivity**, preserving critical components while compressing less sensitive regions.
81
+
82
+ Benefits include:
83
+
84
+ - Better OCR accuracy retention
85
+ - Lower memory usage
86
+ - Faster inference
87
+ - Higher throughput
88
+ - Excellent Apple Silicon optimization
89
+
90
+ ---
91
+
92
+ # Recommended Settings
93
+
94
+ For the best OCR quality:
95
+
96
+ ```yaml
97
+ temp: 0.0
98
+ top_p: 1.0
99
+ top_k: 1
100
+ max_tokens: 4096
101
+ ```
102
+
103
+ For difficult or noisy scans:
104
+
105
+ ```yaml
106
+ temp: 0.1
107
+ top_p: 0.95
108
+ top_k: 20
109
+ max_tokens: 4096
110
+ ```
111
+
112
+ OCR is generally deterministic, so greedy or near-greedy decoding is recommended for maximum transcription accuracy.
113
+
114
+ ---
115
+
116
+ # Example Usage
117
+
118
+ ```python
119
+ from mlx_vlm import load, generate
120
+ from PIL import Image
121
+
122
+ model, processor = load("yugeshkarunamurthy/OvisOCR2-oQe4")
123
+
124
+ image = Image.open("document.png")
125
+
126
+ prompt = """
127
+ Extract all readable content from the document.
128
+ Output Markdown preserving reading order.
129
+ Render tables as HTML.
130
+ Render formulas using LaTeX.
131
+ """
132
+
133
+ response = generate(
134
+ model=model,
135
+ processor=processor,
136
+ image=image,
137
+ prompt=prompt,
138
+ temperature=0.0,
139
+ )
140
+
141
+ print(response)
142
+ ```
143
+
144
+ ---
145
+
146
+ # Optimized For
147
+
148
+ This release is optimized for:
149
+
150
+ - Apple M1
151
+ - Apple M2
152
+ - Apple M3
153
+ - Apple M4
154
+
155
+ Compatible with:
156
+
157
+ - MLX
158
+ - MLX-VLM
159
+ - oMLX
160
+ - Local OCR Applications
161
+ - Document Processing Pipelines
162
+
163
+ ---
164
+
165
+ # Model Highlights
166
+
167
+ - End-to-End OCR
168
+ - Page-level Document Parsing
169
+ - Markdown Generation
170
+ - HTML Table Extraction
171
+ - LaTeX Formula Recognition
172
+ - Scientific Paper Parsing
173
+ - Forms and Receipts
174
+ - Complex Multi-column Documents
175
+ - Reading Order Preservation
176
+ - High OCR Accuracy
177
+
178
+ ---
179
+
180
+ # Intended Use
181
+
182
+ OvisOCR2-oQe4 is well suited for:
183
+
184
+ - OCR
185
+ - PDF Digitization
186
+ - Scientific Paper Extraction
187
+ - Book Digitization
188
+ - Invoice Processing
189
+ - Receipt Processing
190
+ - Form Parsing
191
+ - Research Automation
192
+ - Knowledge Base Construction
193
+ - RAG Preprocessing
194
+ - Markdown Conversion
195
+ - Digital Archives
196
+
197
+ ---
198
+
199
+ # Hardware Recommendations
200
+
201
+ Recommended systems:
202
+
203
+ - Apple M1 Pro / Max / Ultra
204
+ - Apple M2 Pro / Max / Ultra
205
+ - Apple M3 Series
206
+ - Apple M4 Series
207
+
208
+ The compact 0.8B model is lightweight and runs comfortably on most Apple Silicon devices while benefiting from additional memory for larger documents and higher throughput.
209
+
210
+ ---
211
+
212
+ # About oQe4 Quantization
213
+
214
+ oQe4 is a sensitivity-aware mixed-precision quantization technique designed to preserve model quality while significantly reducing memory requirements.
215
+
216
+ Rather than assigning identical precision to every layer, oQe4 analyzes the sensitivity of individual modules and allocates higher precision only where it has the greatest impact.
217
+
218
+ Benefits include:
219
+
220
+ - Better OCR accuracy retention
221
+ - Improved document layout understanding
222
+ - Lower RAM usage
223
+ - Faster inference
224
+ - Excellent Apple Silicon performance
225
+
226
+ ---
227
+
228
+ # Original Model
229
+
230
+ The original **OvisOCR2** is an end-to-end document parsing model built upon **Qwen3.5-0.8B**.
231
+
232
+ Notable features include:
233
+
234
+ - State-of-the-art document OCR
235
+ - Markdown document generation
236
+ - HTML table extraction
237
+ - LaTeX formula recognition
238
+ - Natural reading order
239
+ - End-to-end architecture
240
+ - Compact 0.8B model
241
+ - Apache-2.0 license
242
+
243
+ For benchmark results, technical details, and training methodology, please visit the original repository. 2
244
+
245
+ ---
246
+
247
+ # Credits
248
+
249
+ ## Original Model
250
+
251
+ All credit for the original model, datasets, training methodology, evaluation, benchmarks, and research belongs entirely to:
252
+
253
+ **ATH-MaaS**
254
+
255
+ Original Repository:
256
+
257
+ https://huggingface.co/ATH-MaaS/OvisOCR2
258
+
259
+ Technical Report:
260
+
261
+ https://arxiv.org/abs/2607.13639
262
+
263
+ ---
264
+
265
+ ## oQe4 MLX Quantized Release
266
+
267
+ This repository provides an Apple Silicon optimized **oQe4 MLX quantized** version of the original model.
268
+
269
+ No additional fine-tuning has been performed.
270
+
271
+ ---
272
+
273
+ # Acknowledgements
274
+
275
+ - ATH-MaaS
276
+ - Qwen Team
277
+ - Apple MLX
278
+ - Hugging Face
279
+ - Transformers
280
+ - vLLM
281
+ - SGLang
282
+ - MLX-VLM
283
+ - oMLX
284
+ - OptiQ Quantization
285
+
286
+ ---
287
+
288
+ # Citation
289
+
290
+ If you use this model in research, please cite the original OvisOCR2 paper:
291
+
292
+ ```bibtex
293
+ @misc{lu2026ovisocr2,
294
+ title = {OvisOCR2 Technical Report},
295
+ author = {Shiyin Lu and Yinglun Li and Yu Xia and Yuhui Chen and An-Yang Ji and Jun-Peng Jiang and Qing-Guo Chen and Jianshan Zhao and En Lin and Haijun Li and Cheng Qin and Zhao Xu and Weihua Luo},
296
+ year = {2026},
297
+ eprint = {2607.13639},
298
+ archivePrefix= {arXiv},
299
+ primaryClass = {cs.CV}
300
+ }
301
+ ```
302
+
303
+ ---
304
+
305
+ # License
306
+
307
+ This release inherits the **Apache-2.0** license from the original model.
308
+
309
+ Please refer to the original repository for complete licensing information.
310
+
311
+ ---
312
+
313
+ # Disclaimer
314
+
315
+ This repository contains an optimized **oQe4 MLX quantized conversion** intended for efficient local inference on Apple Silicon.
316
+
317
+ All original model architecture, datasets, training methodology, benchmarks, evaluations, and research remain entirely the work of the original ATH-MaaS team.
318
+ ````3