REVIEW 7 cited by
Syntax-Aware Network for Handwritten Mathematical Expression Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Handwritten mathematical expression recognition (HMER) is a challenging task that has many potential applications. Recent methods for HMER have achieved outstanding performance with an encoder-decoder architecture. However, these methods adhere to the paradigm that the prediction is made "from one character to another", which inevitably yields prediction errors due to the complicated structures of mathematical expressions or crabbed handwritings. In this paper, we propose a simple and efficient method for HMER, which is the first to incorporate syntax information into an encoder-decoder network. Specifically, we present a set of grammar rules for converting the LaTeX markup sequence of each expression into a parsing tree; then, we model the markup sequence prediction as a tree traverse process with a deep neural network. In this way, the proposed method can effectively describe the syntax context of expressions, alleviating the structure prediction errors of HMER. Experiments on three benchmark datasets demonstrate that our method achieves better recognition performance than prior arts. To further validate the effectiveness of our method, we create a large-scale dataset consisting of 100k handwritten mathematical expression images acquired from ten thousand writers. The source code, new dataset, and pre-trained models of this work will be publicly available.
Forward citations
Cited by 7 Pith papers
-
CoMemo: LVLMs Need Image Context with Image Memory
CoMemo adds a cross-attention image-memory path and thumbnail-anchored position encoding to reduce visual neglect in long-context and multi-image LVLM tasks.
-
FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
A 1.5B unified multimodal model trained with discrete flow matching and metric-induced probability paths matches autoregressive baselines of similar size on generation and understanding benchmarks.
-
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
HoVLE is a monolithic VLM whose holistic embedding module maps images and text into one shared space, letting a frozen LLM reach near-compositional performance.
-
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
PVC unifies image and video token compression in VLMs by repeating images as static videos and using causal temporal attention with adaptive compression, achieving strong benchmark results at 64 tokens per frame.
-
DOGR: Towards Versatile Visual Document Grounding and Referring
The authors build a data-generation engine, a seven-task grounding/referring benchmark, and a model that localizes and reads text in document images better than existing MLLMs on that benchmark.
-
Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts
Splitting the vision encoder of an MLLM into domain-specific experts with a lightweight router yields small benchmark improvements at near-zero extra inference cost.
-
DocFusion: A Unified Framework for Document Parsing Tasks
A 289M-parameter generative model with a Gaussian-kernel cross-entropy loss jointly handles layout analysis, OCR, math expression recognition, and table recognition, with competitive but partially overstated benchmark gains.
Discussion (0). Continue with ORCID to comment.