lupAI
pesquisa

Datalab launches omniextractbench to address bias and transparency in document extraction benchmarks

DatalabSource: MarkTechPost02/10/2026, 17:45
Datalab has introduced OmniExtractBench, an open-source benchmark designed to enhance the fairness and transparency of structured document extraction. The benchmark evaluates how accurately systems can populate JSON schemas from PDFs, drawing from 620 documents across four existing benchmarks. A deterministic scorer grades each submission and explains its decisions, aiming to provide a standardized evaluation framework amid the proliferation of vendor-specific leaderboards. OmniExtractBench employs content-based row alignment using the Hungarian algorithm, addressing challenges posed by tables where positional mismatches can drastically affect scores. The benchmark also introduces a verdict layer to refine accuracy metrics, distinguishing between matched values, precision, and recall. Datalab reports that its accurate mode achieved 93.85 accuracy, closely followed by other systems like Reducto deep_extract v2. The benchmark is deployable via PyPI, with data hosted on Hugging Face and code available on GitHub. Datalab emphasizes that empty fields and placeholders do not count as valid matches, preventing potential exploitation through schema padding. The release, published on September 27, 2026, aims to set a new standard for fairness and transparency in document extraction benchmarks.
Datalab launches omniextractbench to address bias and transparency in document extraction benchmarks — lupAI