Open Internet by MindsNet
opendataloader-project/opendataloader-pdf
OpenDataLoader PDF is an open-source tool for extracting AI-ready data from PDFs and automating PDF accessibility compliance. It can parse digital, scanned, and tagged PDFs, producing structured data in Markdown, JSON, HTML, and Tagged PDF formats. Best for: Developers and organizations seeking to extract data from PDFs and ensure accessibility compliance Use cases: Converting scanned PDFs into structured data for machine learning models; Automating PDF accessibility compliance for regulatory requirements; Extracting data from digital PDFs for data analysis and processing Highlights: Benchmark #1 PDF parser with 0.907 overall extraction accuracy and 0.928 table extraction accuracy; Deterministic output with bounding boxes for every element and XY-Cut++ reading order; First open-source PDF auto-tagging to Tagged PDF, with PDF Association and Dual Lab (veraPDF) collaboration
Computing & Technology, Computer Science, Artificial Intelligence