RadarTopicsBuildersWeeklyReads
Open Source Radar
NanoNets/

docext

GitHubWebsite

docext is an on-premises document information extraction and markdown conversion toolkit that supports PDF/image to markdown and benchmarking. It exposes templates, on-prem deployment, and a REST API, with ongoing releases and leaderboard integration.

2.0kstars
148forks
22issues
Apache-2.0license
2025since
Star historydaily snapshots by VibeCrowd

Collecting history — the radar snapshots this repo daily. The trend line appears after 3 days of data (1 so far).

Alternatives & relatedmatched by topic overlap
Reviewgenerated from repository data · Aug 5, 2026

What it is

docext is an on-premises document information extraction and benchmarking toolkit powered by vision-language models. It provides three core capabilities: PDF & Image to Markdown Conversion, Document Information Extraction, and an Intelligent Document Processing Leaderboard.

How it works

  • PDF and Image to Markdown: converts documents to markdown with content recognition and semantic tagging, including LaTeX equation recognition, image descriptions, signature and watermark tagging, page number tagging, and conversion of form controls to Unicode symbols. Table data is converted to HTML tables.
  • Intelligent Document Processing Leaderboard: benchmarks performance across seven tasks such as KIE, VQA, OCR, document classification, long document processing, table extraction, and confidence score calibration.
  • Docext module: offers flexible extraction with custom fields or pre-built templates, table extraction, confidence scoring, on-premises deployment, multi-page support, a REST API, and pre-built templates for common document types (invoices, passports, etc.).

Getting started

Installation and usage details are referenced in the feature guide and EXT_README.md within the repository. New model release highlights show support for Nanonets-OCR-s model integration.

Recent releases

  • v0.1.14 (2025-06-30): Integrate Ollama; dev/benchmark updates.
  • v0.1.7 (2025-04-08): Add vendor-hosted models (OpenAI, Anthropic, OpenRouter); fix for webp images.
  • v0.1.2 (2025-04-05): Base release with custom & pre-built extraction templates, table + field data extraction, Gradio-powered web interface, on-prem deployment.

Traction

Stars: 2032

Behind the repo

Developer activity indicates contributors include @mandalsouvik3333 with early contributions.

Caveats

License: Apache-2.0. Created 2025-03-25; last_push 2026-03-17. Open issues: 22. Language: Python. On-prem deployment supported for Linux and MacOS.

SharePost on XLinkedIn
All trending reposRevenue-verified startups →