Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T01:46:00.580035Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 10 inbound Pith citation observations for arXiv:2502.16161.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T01:46:00.580035Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T07:02:44.498255Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T12:15:01.137692Z
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 13e65ec5-eb4e-4459-9780-9c13d1e65a6a · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Platypus: A generalized specialist model for reading text in various forms
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3c266900-febe-4838-8167-411660fe3c23 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Rico: A mobile app dataset for building data-driven design applications
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 93565460-fbf8-4697-97ac-7e5a24a20fb8 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Icdar 2013 robust reading competition
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 57ebaf9e-277f-4023-87e1-633d5d82aff7 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Towards end-to-end unified scene text detection and layout analysis
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 210b4449-ce46-401d-85b4-767576d4ae98 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Textocr: Towards large-scale end-to-end reasoning for arbitrary- shaped scene text
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 751405ab-7c29-4c9e-8b38-2ded99a6337d · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Abcnet v2: Adaptive bezier-curve network for real-time end-to-end text spotting
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b19aa20e-9a73-4309-b68d-b1054d2b7f06 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Open images v5 text annotation and yet another mask text spotter
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0be2cda8-8f54-49a1-a482-df1179e01030 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Doclaynet: A large human-annotated dataset for document-layout segmentation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f4e6febe-c698-42c8-9c04-6c94d882ab9e · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Icdar2017 robust reading chal- lenge on multi-lingual scene text detection and script identification- rrc-mlt
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e6fccc06-10fc-477d-a5e2-7c92fd14288c · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Coco- text: Dataset and benchmark for text detection and recognition in natural images
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f6477e2c-4157-4a71-863e-560eacbf0db6 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Vision grid transformer for document layout analysis
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 18a860fa-498a-47d2-8ee4-2b242c3a20a7 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Ocr-free document understanding transformer
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9038b828-7e2e-47ae-b98c-4c98f97df9b7 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Publaynet: largest dataset ever for document layout analysis
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b8f1e4f2-29e6-4bc8-830f-c8b649fec6f7 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0c303b5f-53ce-487f-b7f4-737d6dd13490 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Conditional text image generation with diffusion models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b79dadf4-a3d8-439a-97da-672bae7a6d53 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Towards unified scene text spotting based on sequence generation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a323e862-da0d-448e-91a0-349e9e29b5f0 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Challenges in end-to-end neural scientific table recognition
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4b29dbe5-aecb-43c7-a22b-35ba7d679a91 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Tablebank: Table benchmark for image-based table detection and recognition
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5d0f5c36-bd14-40ce-be5d-8218542424be · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models An open approach towards the benchmarking of table structure recognition systems
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b12533ee-63d3-49ce-9193-e7f8bbdc9081 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Icdar 2019 competition on table detection and recognition (ctdar)
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e9776811-ce66-4bc5-aa16-9c4055d19ad2 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Parsing table structures in the wild
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d0b25758-1fa5-481c-8a05-2d66df6560ef · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Visual understanding of complex table structures from document images
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e3d7c24c-885e-418b-9a3d-522d468b3c32 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Icdar 2013 table competition
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e2cbbbd0-0ab7-4a5c-8f6a-12f1370ce05e · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Com- plicated table structure recognition
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d3b52e2a-10dd-4821-8851-b8108b94d53b · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Pubtables-1m: Towards comprehensive table extraction from unstructured documents
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9b65f55a-32a5-4e6b-8fcc-f75339006c80 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Image-based table recognition: data, model, and evaluation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bb386544-2e6a-4f95-9fe7-77924694580c · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Global table extractor (gte): A framework for joint table identifica- tion and cell structure recognition using visual context
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 85653bd0-aa54-4f63-8088-807f7dcc162e · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Pingan- vcgroup’s solution for icdar 2021 competition on scientific literature parsing task b: table recognition to html
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a10ba39d-4254-4358-a6e8-d031c44d6f7a · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Improving table structure recognition with visual- alignment sequential coordinate modeling
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3557634e-ba91-40fd-a72d-6117be7c09e6 · outbound
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 857f6a40-f154-4e3a-9f17-8ddba77c1603 · inbound
Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1082df2f-537e-4cea-aac1-077ec8bce7d1 · inbound
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f966fed9-3b47-4c38-b93f-9d4146b8e943 · inbound
MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ff975bca-8771-4def-a02b-a6391413d32d · inbound
MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa7e96e1-cd41-4ed5-856a-ee7c982473f9 · inbound
InstructTable: Improving Table Structure Recognition Through Instructions OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 35b5a14a-efa8-437f-b74a-ebd5a49636e3 · inbound
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 31c2d01d-f65f-4b92-bab1-50c8e0dbed3b · inbound
AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8dfae63a-20a8-434a-a161-58a7e4fbb871 · inbound
Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
Reference 127
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c4e55fcd-f82a-4e91-b124-51c1dfa09309 · inbound
Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
Reference 127
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 286a4907-071d-463b-a99f-d7ca1b40faae · inbound
StepGuard: Guarding Web Navigation via Single-Step Calibration OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.