Pith. sign in

Paper Citation Record · LEDGER

UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2308.11592.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.11592 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:25.553617Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.680740Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0e2b5b13-105c-4794-aadd-32dc8d11ce79 · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:25.553617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:25.553617Z digest=sha256:5184d26bcf747033cee40c815ddcdd9e767f03715e385fd73fa2ebc6e02697df

Observation 42e9637c-7a60-4daa-90e6-0a6a96d37501 · inbound

Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning cites this paper.

Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:17.222816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:17.222816Z digest=sha256:6d236aa14cbcc81b70526d2f32e186f61ea4f9e71a481c9f8abf23f125bfbbbf

Observation 491b1aae-a70a-43c4-8641-0e21201a8dae · inbound

SETransformer: A Hybrid Attention-Based Architecture for Robust Human Activity Recognition cites this paper.

SETransformer: A Hybrid Attention-Based Architecture for Robust Human Activity Recognition UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:20.539328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:20.539328Z digest=sha256:13e462f03ed75471b3105af490f4f3fe6459caa31c400ba8f120d71b4a315456

Observation bfe7478d-1672-4ab9-b39b-a3d2e9c34a67 · inbound

Multimodal Tabular Reasoning with Privileged Structured Information cites this paper.

Multimodal Tabular Reasoning with Privileged Structured Information UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:58.839998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:58.839998Z digest=sha256:c8c0d606c90bcff4b8917e5430bce2ad3e286f6ceea66db6b52cfb2b85c389c1

Observation d4ed33a4-f251-4927-860b-22ed8db00373 · inbound

DREAM: Document Reconstruction via End-to-end Autoregressive Model cites this paper.

DREAM: Document Reconstruction via End-to-end Autoregressive Model UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.203139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.203139Z digest=sha256:cb58da5d15739d3b1235af6c96867d9f94bb1abd46303da901a3ba2710891dac

Observation 9fc47ca3-2c32-4c31-ad41-99cd785f1d70 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.332555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:842dabbb86d037565b3ac6991010e718883b33e0f7d916a0ad86e875899e0029

Observation 18f6e641-c3e9-4bdb-8b07-e38c0bab17d3 · inbound

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data cites this paper.

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:50.176177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:50.176177Z digest=sha256:b981a196297bcef9fc004cf05d99983cd62f1a47b7ce63c26901c2512aa3c890

Observation 66d31a73-531b-4a3a-b04b-5cc2dc5084f7 · inbound

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method cites this paper.

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:42.195991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:42.195991Z digest=sha256:a804c7c3a789b0d28796ad1cd304e79e928afdf9b2d1e44f170a2216c2c1d5a9

Observation 65d1ca06-3817-414e-9a67-08313bfafc78 · inbound

Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction cites this paper.

Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:03:12.272128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T20:01:57.029143Z digest=sha256:9dc706276260d0a420a1c2880e03bbe33db84cb2772216e7fcff1023be8c5c07

Observation 3ea70519-a7a2-4a26-b304-bb5115688556 · inbound

Feature Perturbation Pool-based Fusion Network for Unified Multi-Class Industrial Defect Detection cites this paper.

Feature Perturbation Pool-based Fusion Network for Unified Multi-Class Industrial Defect Detection UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:29:21.633982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T03:28:11.747826Z digest=sha256:297ea740b65fb5635b7a67c6bb6a84c0cab217bebd48667a7a85e4cdf361d084

Observation 40e68cff-df95-4c4b-a0cb-0a2a6d77cd14 · inbound

Lightweight Real-Time Rendering Parameter Optimization via XGBoost-Driven Lookup Tables cites this paper.

Lightweight Real-Time Rendering Parameter Optimization via XGBoost-Driven Lookup Tables UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:21:14.845234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T17:25:18.561710Z digest=sha256:171c54602aedddda5c93609b20522044682929819526acdadfa30aa8d5e93d3c

Observation 173571e1-59a4-4d9d-bcfd-0229e46e3375 · inbound

Image Classification via Random Dilated Convolution with Multi-Branch Feature Extraction and Context Excitation cites this paper.

Image Classification via Random Dilated Convolution with Multi-Branch Feature Extraction and Context Excitation UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:22:02.896411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T17:07:55.177114Z digest=sha256:9363c64220097d44ba44f393a66ff1381b3cc421cb5cb7972c5f215caa3a1854

Observation 74eb008a-5670-42f5-a095-8928d9f4af5d · inbound

Gait Recognition via Deep Residual Networks and Multi-Branch Feature Fusion cites this paper.

Gait Recognition via Deep Residual Networks and Multi-Branch Feature Fusion UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:01:27.637705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T08:33:09.327681Z digest=sha256:0553159729e5ed900c04b8066def755d7bd064b52ca93e9d0ef854fdfd6c1726

Observation d5f1614e-d80a-4f87-a011-61fffda2d4b3 · inbound

Multi-Branch Non-Homogeneous Image Dehazing via Concentration Partitioning and Image Fusion cites this paper.

Multi-Branch Non-Homogeneous Image Dehazing via Concentration Partitioning and Image Fusion UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:08.682659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T20:03:36.551943Z digest=sha256:9b2de95abe1c380c194a9617ae09916c065f7146f6fda9ccb4abd685ab2fb27b

Observation a254b20e-b1bc-468f-9d04-297da62be12c · inbound

Local Spatiotemporal Convolutional Network for Robust Gait Recognition cites this paper.

Local Spatiotemporal Convolutional Network for Robust Gait Recognition UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:08:29.320024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:07:25.917495Z digest=sha256:dd63ff1d24b6c928bd854f8ab331c9ff63f0660d97459f3837f8c6a3affcb60c

Observation 9921cc2d-3f12-4dae-8d28-12f01bd57dd3 · inbound

Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting cites this paper.

Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:33:14.584828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:28:51.711892Z digest=sha256:f38733f0cb6e14e4646223f1c0a929b6162a2a734a242296bced35f233423ab8

Observation e39545d6-7c74-40c9-9b46-2e5e934fec8a · inbound

An LMM for Precisely Grounding Elements in Documents cites this paper.

An LMM for Precisely Grounding Elements in Documents UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.682949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T01:51:50.043216Z digest=sha256:32c6bcf7c26867d90ca226cff4e70b8bd5c2d48d0ed67b81698f78ed13e6aaa7

Observation 01858738-d7f0-413c-ae03-36fc32022231 · inbound

DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning cites this paper.

DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:04:21.403127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:01:20.803078Z digest=sha256:02f987ef4b40b2596ac4fda5e012abe335d5bd0ddecf14a10a035ae4d5c949f6

Observation fb14e358-776c-4cd9-92c6-922475fe23a9 · inbound

ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators cites this paper.

ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:33.653477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T15:23:48.748933Z digest=sha256:eb662111a4d1d5fca2570c10df3e9e9fc8c68676e4d98edb15e0c6732230e3d2

Observation 9db67247-d520-4e54-8ab4-e626c6140e15 · inbound

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth cites this paper.

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 217

Resolution
unresolved
no resolver link, observed 2026-08-02T14:55:42.466530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:55:42.466530Z digest=sha256:c26c0a1456fac877d96303a0f0d7a92016bc07ae6738593f4bc9c71f26cb04df