Open Internet by MindsNet
Extracting Tables from PDFs with VLM Models
The extraction of tables from PDFs, particularly those with borderless tables or more than 5-6 columns, remains a significant challenge. Current open-source solutions like docling, graphite-docling, and marker have limitations. The challenge is to find an effective open-source method for converting PDFs to Markdown, especially for financial data.
Computing & Technology, Computer Science, Machine Learning