Q-Mask uses query-conditioned causal masks to separate text location from recognition in OCR VLMs, backed by a new benchmark and 26M-pair training dataset.
arXiv:1908.04729 (2019)
11 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
TableNet is a new large-scale table dataset created via LLM multi-agent generation, combined with diversity-based active learning that achieves competitive performance on its test set and superior results on real-world tables using fewer samples than baselines.
An agent harness combining staged task decomposition, multimodal evidence tooling, and artifact-grounded self-improvement scores 81.0 GRAS on multimodal scientific curation, 22.4 points above the strongest baseline — with the caveat that 8 of 23 evaluation papers were used for optimization.
Geometry-Aware Pointer Loss reweights cross-entropy by inverse Manhattan distance to focus gradients on adjacent-cell errors in TSR, yielding SOTA results on PubTabNet and SynthTabNet.
FastTab combines a Tiny Recursive Module and axial 1D Transformer encoders to predict table grids, headers, and cell spans directly, achieving competitive accuracy on four benchmarks with low-latency inference.
MPDocBench-Parse provides 433 annotated multi-page documents and an evaluation protocol covering text/table/formula extraction, merging, figure extraction, reading order, and heading hierarchy for realistic document parsing.
A provenance-aware modular pipeline converts handwritten tabular images to knowledge graphs through three inspectable stages with full traceability to visual and textual origins.
DenTab provides 2,000 annotated dental table images and 2,208 questions to benchmark 16 systems on table structure recognition and VQA, revealing that strong layout recovery does not ensure reliable multi-step arithmetic, and proposes a Table Router Pipeline combining VLMs with rule-based execution.
TableSeq unifies table structure recognition, content extraction, and cell localization by generating an interleaved autoregressive sequence of HTML tags, cell text, and discretized coordinate tokens from an input image.
InternVL 2.5 is the first open-source MLLM to surpass 70% on the MMMU benchmark via model, data, and test-time scaling, with a 3.7-point gain from chain-of-thought reasoning.
Survey proposing a taxonomy for document parsing into pipeline-based systems and VLM-driven unified models, reviewing components, metrics, benchmarks, and challenges.
citing papers explorer
-
Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models
Q-Mask uses query-conditioned causal masks to separate text location from recognition in OCR VLMs, backed by a new benchmark and 26M-pair training dataset.
-
TableNet A Large-Scale Table Dataset with LLM-Powered Autonomous
TableNet is a new large-scale table dataset created via LLM multi-agent generation, combined with diversity-based active learning that achieves competitive performance on its test set and superior results on real-world tables using fewer samples than baselines.
-
Building Agent Harnesses for Scientific Curation from Multimodal Sources
An agent harness combining staged task decomposition, multimodal evidence tooling, and artifact-grounded self-improvement scores 81.0 GRAS on multimodal scientific curation, 22.4 points above the strongest baseline — with the caveat that 8 of 23 evaluation papers were used for optimization.
-
Rethinking the Pointer Loss in Table Structure Recognition: Geometry-Aware Pointer Loss for Spatial Locality
Geometry-Aware Pointer Loss reweights cross-entropy by inverse Manhattan distance to focus gradients on adjacent-cell errors in TSR, yielding SOTA results on PubTabNet and SynthTabNet.
-
FastTab: A Fast Table Recognizer with a Tiny Recursive Module and 1D Transformers
FastTab combines a Tiny Recursive Module and axial 1D Transformer encoders to predict table grids, headers, and cell spans directly, achieving competitive accuracy on four benchmarks with low-latency inference.
-
MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing
MPDocBench-Parse provides 433 annotated multi-page documents and an evaluation protocol covering text/table/formula extraction, merging, figure extraction, reading order, and heading hierarchy for realistic document parsing.
-
From Historical Tabular Image to Knowledge Graphs: A Provenance-Aware Modular Pipeline
A provenance-aware modular pipeline converts handwritten tabular images to knowledge graphs through three inspectable stages with full traceability to visual and textual origins.
-
DenTab: A Dataset for Table Recognition and Visual QA on Real-World Dental Estimates
DenTab provides 2,000 annotated dental table images and 2,208 questions to benchmark 16 systems on table structure recognition and VQA, revealing that strong layout recovery does not ensure reliable multi-step arithmetic, and proposes a Table Router Pipeline combining VLMs with rule-based execution.
-
TableSeq: Unified Generation of Structure, Content, and Layout
TableSeq unifies table structure recognition, content extraction, and cell localization by generating an interleaved autoregressive sequence of HTML tags, cell text, and discretized coordinate tokens from an input image.
-
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
InternVL 2.5 is the first open-source MLLM to surpass 70% on the MMMU benchmark via model, data, and test-time scaling, with a 3.7-point gain from chain-of-thought reasoning.
-
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction
Survey proposing a taxonomy for document parsing into pipeline-based systems and VLM-driven unified models, reviewing components, metrics, benchmarks, and challenges.