Pith. sign in

Paper Citation Record · LEDGER

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

As of 6 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 10 inbound Pith citation observations for arXiv:2502.16161.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.16161 v2

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T01:46:00.580035Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:02:44.498255Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 13e65ec5-eb4e-4459-9780-9c13d1e65a6a · outbound

This paper cites Platypus: A generalized specialist model for reading text in various forms.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Platypus: A generalized specialist model for reading text in various forms

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.608542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:7cade7f8cabec75fbd80b12cc53c885f63074c9bd772ddfbd0b989356f586e40

Observation 3c266900-febe-4838-8167-411660fe3c23 · outbound

This paper cites Rico: A mobile app dataset for building data-driven design applications.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Rico: A mobile app dataset for building data-driven design applications

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.606082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:477fa3ce87b38e09b6fabc9fd4ffd778779681c242177bfcd31a5dc5a0df2fae

Observation 93565460-fbf8-4697-97ac-7e5a24a20fb8 · outbound

This paper cites Icdar 2013 robust reading competition.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Icdar 2013 robust reading competition

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.573637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:5c000a3a5429442ac7ed42fd8a77fc7a2a0fbc75b0141de87b84169b71625649

Observation 57ebaf9e-277f-4023-87e1-633d5d82aff7 · outbound

This paper cites Towards end-to-end unified scene text detection and layout analysis.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Towards end-to-end unified scene text detection and layout analysis

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.554970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:5eb6a0773586f4bdfa4bac95793c139e07dc65557b2289bff590ef0f1658e4b4

Observation 210b4449-ce46-401d-85b4-767576d4ae98 · outbound

This paper cites Textocr: Towards large-scale end-to-end reasoning for arbitrary- shaped scene text.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Textocr: Towards large-scale end-to-end reasoning for arbitrary- shaped scene text

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.591873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:a2f1dc57aa2b90fdb210d163520edbde17bf557f27e1682de7fc3ad68423d33b

Observation 751405ab-7c29-4c9e-8b38-2ded99a6337d · outbound

This paper cites Abcnet v2: Adaptive bezier-curve network for real-time end-to-end text spotting.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Abcnet v2: Adaptive bezier-curve network for real-time end-to-end text spotting

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.561641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:d19474aa1606bc6be44cf5214a1c4a71106c9ffe98b8c257dd4b53e3514aa809

Observation b19aa20e-9a73-4309-b68d-b1054d2b7f06 · outbound

This paper cites Open images v5 text annotation and yet another mask text spotter.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Open images v5 text annotation and yet another mask text spotter

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.589552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:4dd52841fd39bcab7ac4582364aa6267f74e8e5782c9e18e56b14a08efb309a3

Observation 0be2cda8-8f54-49a1-a482-df1179e01030 · outbound

This paper cites Doclaynet: A large human-annotated dataset for document-layout segmentation.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Doclaynet: A large human-annotated dataset for document-layout segmentation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.619159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:16582a37bf61f68aabf57d34efc06bd2c4916839deb958e4219753627c9e521c

Observation f4e6febe-c698-42c8-9c04-6c94d882ab9e · outbound

This paper cites Icdar2017 robust reading chal- lenge on multi-lingual scene text detection and script identification- rrc-mlt.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Icdar2017 robust reading chal- lenge on multi-lingual scene text detection and script identification- rrc-mlt

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.564412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:4d56034303959d20cd4793bb5b51abaea26bee5f0d68516fe3ffe221703484d7

Observation e6fccc06-10fc-477d-a5e2-7c92fd14288c · outbound

This paper cites Coco- text: Dataset and benchmark for text detection and recognition in natural images.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Coco- text: Dataset and benchmark for text detection and recognition in natural images

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.578499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:4e1490aa45eb4e06b9740ef419370d8a11655d6f042b5d5132d9dd8201868b77

Observation f6477e2c-4157-4a71-863e-560eacbf0db6 · outbound

This paper cites Vision grid transformer for document layout analysis.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Vision grid transformer for document layout analysis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.583106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:99e71e0f452c5612e71382fe723fa7e8be964ea7b76f06ffb9c1671c3ff6c183

Observation 18a860fa-498a-47d2-8ee4-2b242c3a20a7 · outbound

This paper cites Ocr-free document understanding transformer.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Ocr-free document understanding transformer

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.576089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:2c2e7ab827b2bfeacc1ae83d268ab805487b97a77c4e0effb7454c590e180280

Observation 9038b828-7e2e-47ae-b98c-4c98f97df9b7 · outbound

This paper cites Publaynet: largest dataset ever for document layout analysis.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Publaynet: largest dataset ever for document layout analysis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.598747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:36e23bca0fe08552b9fc32937cdbd2928aeff01debef62228a7bfa9fab07a333

Observation b8f1e4f2-29e6-4bc8-830f-c8b649fec6f7 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.566818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:4645de082c1ba0386064154d2bd90a4dc14b4f4f7d63cdcaf24c1b5c6b2fec27

Observation 0c303b5f-53ce-487f-b7f4-737d6dd13490 · outbound

This paper cites Conditional text image generation with diffusion models.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Conditional text image generation with diffusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.596170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:e98d0236daf908e1192c28ff0e5edc6afad8a94dc6aa64b3baa14df1c3137f35

Observation b79dadf4-a3d8-439a-97da-672bae7a6d53 · outbound

This paper cites Towards unified scene text spotting based on sequence generation.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Towards unified scene text spotting based on sequence generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.614070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:be2bd761c953b90a4bc028272899317ad0f359da6bf48c007d865ddfb8de9492

Observation a323e862-da0d-448e-91a0-349e9e29b5f0 · outbound

This paper cites Challenges in end-to-end neural scientific table recognition.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Challenges in end-to-end neural scientific table recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.593840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:d2f1e0f39bcf52f6cf2e495bf0308d6fa93d6807ccf9e025360c84921e0d046e

Observation 4b29dbe5-aecb-43c7-a22b-35ba7d679a91 · outbound

This paper cites Tablebank: Table benchmark for image-based table detection and recognition.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Tablebank: Table benchmark for image-based table detection and recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.557251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:4cbd08a531d6360f70c983cd821b376c268f8e4fd6635b62b961a08342f82bed

Observation 5d0f5c36-bd14-40ce-be5d-8218542424be · outbound

This paper cites An open approach towards the benchmarking of table structure recognition systems.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models An open approach towards the benchmarking of table structure recognition systems

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.559559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:1c63cffe8af7eac7d7b45624513c127ef32fe3455c68aff61273d71609b4974e

Observation b12533ee-63d3-49ce-9193-e7f8bbdc9081 · outbound

This paper cites Icdar 2019 competition on table detection and recognition (ctdar).

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Icdar 2019 competition on table detection and recognition (ctdar)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.587593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:455a42d28bbaf13b1612abfefc210f2add68a4d8bc74e5ce94e5f9fd0185dc46

Observation e9776811-ce66-4bc5-aa16-9c4055d19ad2 · outbound

This paper cites Parsing table structures in the wild.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Parsing table structures in the wild

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.616495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:1c4774a2f4d4bb998c5567b87fff3bbc420c4b3ba8d5d4a3c61027de949a443f

Observation d0b25758-1fa5-481c-8a05-2d66df6560ef · outbound

This paper cites Visual understanding of complex table structures from document images.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Visual understanding of complex table structures from document images

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.601327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:a1ee0987f6242bc1fc2c49d9b97ebd7d72e4a2717c929100e04c0ea406119d38

Observation e3d7c24c-885e-418b-9a3d-522d468b3c32 · outbound

This paper cites Icdar 2013 table competition.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Icdar 2013 table competition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.580948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:7de0c9061534add88e338105f2859673f28812d6eebd9ee669583f1a8daebe31

Observation e2cbbbd0-0ab7-4a5c-8f6a-12f1370ce05e · outbound

This paper cites Com- plicated table structure recognition.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Com- plicated table structure recognition

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.603614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:bebe0be67df05f2e1194fd8876c9346db63cd1e35f22f9452c385a4ffba7bf22

Observation d3b52e2a-10dd-4821-8851-b8108b94d53b · outbound

This paper cites Pubtables-1m: Towards comprehensive table extraction from unstructured documents.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Pubtables-1m: Towards comprehensive table extraction from unstructured documents

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.611610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:817e448a3f8ad835f4fa61551eb34f9ded19dd800328a0f16fdd80317370a881

Observation 9b65f55a-32a5-4e6b-8fcc-f75339006c80 · outbound

This paper cites Image-based table recognition: data, model, and evaluation.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Image-based table recognition: data, model, and evaluation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.552622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:6f6aff9e619382445decb0dba37aa79deaf19abe5bd0627f5bba1a1db45fddb9

Observation bb386544-2e6a-4f95-9fe7-77924694580c · outbound

This paper cites Global table extractor (gte): A framework for joint table identifica- tion and cell structure recognition using visual context.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Global table extractor (gte): A framework for joint table identifica- tion and cell structure recognition using visual context

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.571447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:71f011556d4769d86c01dcf0a00fb1fe25bd6cd3fe7aa487780fc73de84ce774

Observation 85653bd0-aa54-4f63-8088-807f7dcc162e · outbound

This paper cites Pingan- vcgroup’s solution for icdar 2021 competition on scientific literature parsing task b: table recognition to html.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Pingan- vcgroup’s solution for icdar 2021 competition on scientific literature parsing task b: table recognition to html

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.585188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:e516003edf0f510a6938f44633ff7b51b3e14550c2b1c51de8e331c09963ffc3

Observation a10ba39d-4254-4358-a6e8-d031c44d6f7a · outbound

This paper cites Improving table structure recognition with visual- alignment sequential coordinate modeling.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Improving table structure recognition with visual- alignment sequential coordinate modeling

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.569158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:d905afc8da9e114c673c071c3ae25ea962333a2b80deb00c82b8c1dcf2b56240

Observation 3557634e-ba91-40fd-a72d-6117be7c09e6 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:47:22.623354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:46:00.580035Z digest=sha256:3e8b5b672d3ef37d0c14d6f3233b07d1ea404a76299f093f6a7eb08d1b03d30f

Pith citing papers

Observation 857f6a40-f154-4e3a-9f17-8ddba77c1603 · inbound

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning cites this paper.

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T07:02:44.498255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:02:44.498255Z digest=sha256:d150802c66e472c31c201572076fcca3966bf7a6db178362e116afd066054852

Observation 1082df2f-537e-4cea-aac1-077ec8bce7d1 · inbound

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding cites this paper.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:25.554263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:25.554263Z digest=sha256:89c339913dd119162f794c3cd97c98e081bbbb7103eb66fc0449595bb03e95d1

Observation f966fed9-3b47-4c38-b93f-9d4146b8e943 · inbound

MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents cites this paper.

MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:01:20.564324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T22:59:16.413568Z digest=sha256:c8384f9fd3154a38337d3e9107ec64bc60a212c3b729afdfb9a0bc8d6daf1697

Observation ff975bca-8771-4def-a02b-a6391413d32d · inbound

MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents cites this paper.

MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T16:38:37.451063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:38:37.451063Z digest=sha256:faa042e7625aebe6ca4e9d6ae9a14cdce7b93161941adcf0c937cad5d9ec7013

Observation aa7e96e1-cd41-4ed5-856a-ee7c982473f9 · inbound

InstructTable: Improving Table Structure Recognition Through Instructions cites this paper.

InstructTable: Improving Table Structure Recognition Through Instructions OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:58:16.054688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T20:53:57.029294Z digest=sha256:3a74447c705251edf588f83a9e9722b3c4580040c32e55e0c4772b0196ead115

Observation 35b5a14a-efa8-437f-b74a-ebd5a49636e3 · inbound

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web cites this paper.

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:40:58.655472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:00:34.401698Z digest=sha256:da3af47aeb39aaba475236de0f4d86dc38a0e97fba0323edbdfe1b77a6bec5b0

Observation 31c2d01d-f65f-4b92-bab1-50c8e0dbed3b · inbound

AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding cites this paper.

AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-05-09T06:20:38.679110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T18:37:48.993080Z digest=sha256:83cf382156912f7604c59f12dd2c2d36414521ca143acdf18d6d1480fbef3bc2

Observation 8dfae63a-20a8-434a-a161-58a7e4fbb871 · inbound

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment cites this paper.

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

Reference 127

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:58:28.772144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T01:58:19.295247Z digest=sha256:dc5118fd29defd1879cbf06b238b022ba79719ac232f293d9f08fd35952a00f0

Observation c4e55fcd-f82a-4e91-b124-51c1dfa09309 · inbound

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment cites this paper.

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

Reference 127

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:47:40.420945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T16:45:16.963802Z digest=sha256:6c88d23df82b31dbd10ca5aa596721066c3392997d256a01b5b4240af4817af9

Observation 286a4907-071d-463b-a99f-d7ca1b40faae · inbound

StepGuard: Guarding Web Navigation via Single-Step Calibration cites this paper.

StepGuard: Guarding Web Navigation via Single-Step Calibration OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-27T01:40:21.154131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T01:34:10.325378Z digest=sha256:d0947b4e97a41cc56269d0408bc52c2ea666975fce5ef8999d97580efd1cfd47