A Repo For Document AI
-
Updated
Sep 2, 2026 - Python
A Repo For Document AI
Complex data extraction and orchestration framework designed for processing unstructured documents. It integrates AI-powered document pipelines (GenAI, LLM, VLLM) into your applications, supporting various tasks such as document cleanup, optical character recognition (OCR), classification, splitting, named entity recognition, and form processing
[ICDAR 2024] Multi-Cell Decoder and Mutual Learning for Table Structure and Character Recognition [ICDAR 2026] Revisiting Structural Dependency in Autoregressive Multi-task Table Recognition via Order-Independent Cell-Level Representations
To associate your repository with the pubtabnet topic, visit your repo's landing page and select "manage topics."