@marc.kaz · Marc Kaz
Saved 2026-07-02 · Posted 2026-06-30 · Status: New
olmOCR turns messy PDFs, scans, and images into clean, structured Markdown that models can actually understand.
Ships with:
• Handles tables, equations, handwriting & multi-column layouts
• Preserves natural reading order
• Works on old scans, figures, insets, headers & footers
• Fixes the “first mile” problem in RAG/document AI
Research papers, legal filings, financial reports — no more broken text.
👉 https://github.com/allenai/olmocr
Who’s fixing their document pipeline today? Drop a 🔥
Content ideas (0)
No ideas generated yet. Run /instagram-sync ideate from Claude Code to create some.
Comments (14)
(Optical)
Be a real one and tell us required specs from now on.
optical not optimal 😉
🚨 THE OCR TOOL BUILT FOR THE LLM ERA JUST DROPPED
olmOCR turns messy PDFs, scans, and images into clean, structured Markdown that models can actually understand.
Ships with:
• Handles tables, equations, handwriting & multi-column layouts
• Preserves natural reading order
• Works on old scans, figures, insets, headers & footers
• Fixes the “first mile” problem in RAG/document AI
Research papers, legal filings, financial reports — no more broken text.
👉 https://github.com/allenai/olmocr
Who’s fixing their document pipeline today? Drop a 🔥
There are so many of these, Microsoft has one as well. What one is the fastest and least token consuming/generating
I made a 99.99% fidelity OCR for complex technical documents in chemical engineering, physics, etc,that I use for my owm AI training.
Yoooooo idk where you get this gold but I really appreciate you sharing it 🙌." We need more people like you 👏. Open source is the key.
🔥
🔥 Getting rid of Google drive for obsidian, this will help
My gemma 4 would love this. Thanks for sharing.
how is this better than Google vision?
I build something like this but hand writing is a nightmare. You gotta have thousands of training hours to make it work. However if you're working with printing fonts, it's perfect. Otherwise, you can have a maximum accuracy up to 10-20% of the samples.
PaddleOCR is just better
How is it better than docling