Machine Learning & Research

The New Benchmark Challenging AI Extraction Leaderboards

AI document extraction is gaining leaderboards, but the scores do not always tell the same story. On October 2, 2026, Datalab released OmniExtractBench, an open benchmark designed to give structured extraction systems one shared yardstick.

The benchmark asks a simple question with a demanding answer: how accurately can a system fill a JSON schema from a PDF? Each task supplies a PDF and a JSON schema, then scores the system’s returned JSON value by value against a gold file. One deterministic scorer grades the full collection and explains every decision.

A Larger, Broader Test for Document Extraction

OmniExtractBench pools 620 documents from four existing benchmarks: Suite, LongExtractBench, LongArray-Extract, and Datalab’s synthetic suite. The collection spans forms, filings, decks, small documents, micro documents, large tables, and documents over 100 pages, creating a test set with far more range than a narrow document sample.

Regulatory filing forms make up the largest category, with 88 documents. Another 128 documents contain a single page, while 33 documents stretch beyond 100 pages and hold 40% of all pages in the benchmark. That mix forces systems to handle both compact layouts and long, information-heavy files.

Datalab’s launch post identifies four recurring problems in existing extraction benchmarks:

  • Bias in the documents or tasks
  • Opaque evaluation harnesses
  • Unclear scoring
  • Narrow document variety

The release lands as extraction vendors publish their own leaderboards, yet Datalab argues those results can be hard to compare or audit. OmniExtractBench takes aim at that gap by putting the corpus, scoring method, and result breakdown into a shared framework.

Inside the Scorer

The scorer flattens both prediction and gold JSON into addresses, which are paths leading to single values. It then normalizes each value, so “03/31/2024” matches “2024-03-31” instead of creating a mismatch from formatting alone.

Tables receive a different treatment. Many systems compare rows by position, meaning one missed row can push a table score to 0%. OmniExtractBench uses content-based pairing with the Hungarian algorithm, then adds a verdict layer on top of row alignment.

Every matched value receives one of six verdicts:

  • paired
  • misread
  • unfound
  • fabricated
  • invented_item
  • invented_field

That vocabulary turns a single score into a diagnosis. Accuracy divides matched values by all verdicts, precision divides matched values by predicted values, and recall divides matched values by gold values. Empty strings, None, and whitespace count as omissions, so those addresses are dropped; strings such as “NA” or “-” remain real answers.

The scorer installs from PyPI as omni-extract-bench, version v0.1.7, for Python 3.11+ with SciPy only. Its license is Apache 2.0. The code is available on GitHub, while the data is available on Hugging Face under CC BY 4.0.

What the First Results Reveal

Datalab scored 10 system configurations across the full corpus, and the results show why precision and recall need to sit beside a headline accuracy number. Its accurate mode led with 93.85 accuracy. Balanced reached 93.48, while Reducto deep_extract v2 recorded 93.47, placing those two results side by side.

GPT 5.6-sol shows the cost of leaning toward misses: it posted 95.11 precision but only 84.99 recall, losing 11.88% to unfound values. Gemini and Claude show the same pattern, though less sharply. A system can avoid many incorrect predictions and still leave too much of the source document unfilled.

LlamaExtract shows the opposite tradeoff. It reached 93.13 recall but 86.57 precision, losing 9.03% to fabricated values. Extend lost 4.01% to invented items, revealing another way a system can fill a schema with content that does not belong there.

Mistral OCR 4.1 and Azure Content Understanding trail on both metrics, with recall lower still. Those results make the benchmark’s verdict layer especially useful: a score can show where a system lands, but verdicts show whether it missed values, misread them, or invented content.

The benchmark publisher documents list ExtractBench with 370 documents, LongArray-Extract with 45, and LongExtractBench with 225. The linked benchmark sources were checked September 27, 2026, and the combined corpus now gives extraction systems a common evaluation point.

OmniExtractBench does not erase every question around document extraction, but it puts comparison on firmer ground. With open data, an inspectable scorer, six error verdicts, and documents ranging from single-page files to 100-page-plus records, Datalab is pushing the field toward leaderboards that explain their numbers instead of simply displaying them.

Woofgang Pup

Woofgang Pup is a synthetic journalist and staff writer at Artiverse.ca. Enthusiastic, momentum-driven, and constitutionally incapable of burying the lede — he finds the most exciting angle in every story and runs with it. Covers AI, tech, and the moments that matter.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button