Pith. sign in

Paper Citation Record · LEDGER

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis

As of 11 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 0 inbound Pith citation observations for arXiv:2607.04636.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04636 v1

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T16:02:00.920066Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

96 of 96 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved96
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 06d27319-2041-4dea-9113-3b50573c4fe3 · outbound

This paper cites UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:72b6943438cdcec857f54151aac0505d8e20154ce57d1185c794a2ce332963f5

Observation 280541b3-5625-42a8-b733-cb1285de1ec7 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:9d431938822b54ead55e461ce02dc66cf9fb028af7bb93287d67d3ef00d2a8ce

Observation 621d4859-753b-49e9-9882-b07955f8714f · outbound

This paper cites Proceedings of the 14th ACM international conference on Information and knowledge management , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 14th ACM international conference on Information and knowledge management , pages=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:cf5135d323a90936161b5217b0a663a7d9f6ec440ad05c506059fbc8b37077eb

Observation 223b5939-5f87-46c7-9651-d6eda948ddb0 · outbound

This paper cites ACM Computing Surveys , volume=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis ACM Computing Surveys , volume=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:252a1fdc4831f023c459c141ac3f3e085d6ad3de54a2b187b59f5ed492ade0df

Observation 9956dcd8-1188-40af-a70b-3e3b7750ffd0 · outbound

This paper cites Artificial Intelligence Review , volume=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Artificial Intelligence Review , volume=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:8b66c83ede33127e8f308820903d77e5caa7be413f2142fbbf7ae3ba8e963fe9

Observation f5ac9e35-cfa7-4522-98f5-dc6b98b0a828 · outbound

This paper cites Spatial Dual-Modality Graph Reasoning for Key Information Extraction.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Spatial Dual-Modality Graph Reasoning for Key Information Extraction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:b7b6e5c3d6c4f1fd406c1cb391c151be0bd1cd178cd2f6a5e7c769eaa50f1984

Observation 413b357d-4933-4422-bd3a-4d74dee709b1 · outbound

This paper cites Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:46cc23fdc8555616ac333bb8e6fbc32e12b903f2aa24701c4ebb8ea8e7d4902d

Observation eed51a40-4270-40f3-9e1d-745b7f50d52f · outbound

This paper cites 2019 International Conference on Document Analysis and Recognition , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis 2019 International Conference on Document Analysis and Recognition , pages=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:5599bd738f6714b948abba297087a25b21efc39e22d5d469d4c4fdb9860c3330

Observation 18963410-65f3-4049-80f0-aa3e7a05fe89 · outbound

This paper cites Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:066cc3c8e03a4054212d6d55477331e3e9ac01aed59d7fd57d01942e15566258

Observation 33bb58e6-3d5e-4a29-ab02-125053bb9ebd · outbound

This paper cites Proceedings of the 30th ACM international conference on multimedia , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 30th ACM international conference on multimedia , pages=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:ab056a9fa1efba036e8d439b3d03c03f10298ba0378049e645e75a6ff7e31e9b

Observation 63dcfff2-98eb-46d7-951a-29dfa9d5e82e · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:8cc6be0380bae32aac6fcd3d25f884ab2cf8a47f0bab87b5514edc9426a60507

Observation 603affde-05ca-4a17-94e3-92d081d62d26 · outbound

This paper cites Artificial Intelligence Review , year=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Artificial Intelligence Review , year=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:6113f1f5615e12960ee13ab4ee95f6c352652685b3a0ee2e43c547ce26c456b4

Observation 71d01a8c-8c17-4286-a810-793821e7ce03 · outbound

This paper cites 2020 25th International conference on pattern recognition (ICPR) , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis 2020 25th International conference on pattern recognition (ICPR) , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:cfd55fbad8f1962cf4ae87c8ac76b74aae37bf1c6eada413c4995d038e843d4d

Observation 134ccc23-5d5d-47e3-8fa8-2638f88e134d · outbound

This paper cites Lang2Act: Fine-Grained Visual Reasoning through Self-Emergent Linguistic Toolchains.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Lang2Act: Fine-Grained Visual Reasoning through Self-Emergent Linguistic Toolchains

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:0c7969770bd8263460975f7acf7014b923672f40363845ba94a86c17d54d7bef

Observation 38891b53-678b-4dd5-b109-8fa058a86bfe · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:1b1ee4b81c6a0e5ae3d700d835d7ad4c56178bdc887f79f89782cb698ffcce8f

Observation 8f4d180f-4297-4370-86a0-2ba6247cd0a9 · outbound

This paper cites MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:24b777d19fc6fa572264ff503600dfcfae56c6603b05cf6e265ec925153e3c68

Observation bb841e23-0e1c-448e-9519-cdee7432d478 · outbound

This paper cites VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:1f00a6123a56310d654599d7eb91024a63573e2e59d61d686aff7c46338d6627

Observation 61b58683-faff-46e7-95fe-8c7f347cf367 · outbound

This paper cites Proceedings of the 30th ACM international conference on multimedia , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 30th ACM international conference on multimedia , pages=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:8742b5fe610f6425e9c06098bce36d8c8550d3d27a07cb59a7457e5c7826c4f0

Observation d4f3cdf5-27c1-4949-a424-d004e54ce2d5 · outbound

This paper cites Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:616f0daa9d734f6e03ebf1a65394af6005b37736f7b750e8ef3da2ca10d80a84

Observation 9f70ea1f-30b4-475a-a974-03ea6c2751a7 · outbound

This paper cites Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:ec408e91fcbfc1a487b8312d4bd9299fb37e2f035908660481294e0fb17b1230

Observation 66d38330-e94f-4b95-bc86-05e12f800d78 · outbound

This paper cites Qwen2.5-VL Technical Report.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Qwen2.5-VL Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:6407bdb6adcad5c7bb37218abc6228a16bf63b911a2c8168023a83e781fdf557

Observation 3121aab9-e59d-4b4c-90c1-044f41bc10be · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:62c356ac3154d08d6f5663c8919b139677599927d6abfc8bd9f9a65d15668088

Observation 1c1c0523-36c4-4316-8999-261c0eefe40f · outbound

This paper cites Proceedings of the 2024 conference on empirical methods in natural language processing , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 2024 conference on empirical methods in natural language processing , pages=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:81abcfea3861f6c07059c13404d0d6c50918568a73be8512d076bd50d9a61f1b

Observation 8058a3e9-3cae-47c6-b388-6c65a2872a1f · outbound

This paper cites Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer , booktitle =.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer , booktitle =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:32c3df880607c55a6332ad79a42ed82143c125489b1a0b2b0ecb9195fcee7baf

Observation 493c7f8b-5557-4012-a4b9-2ed035e1cc38 · outbound

This paper cites Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , pages=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:d0762db145341c8f793cbbf1b2540c243a0860fe572e74fdfb915b98ef6e0e16

Observation f70dc14e-c0be-47ba-8d0b-2673a20e556e · outbound

This paper cites Advances in neural information processing systems , volume=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Advances in neural information processing systems , volume=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:c4883467938c1c36f4fc7a01c2f0eec98fc228ac41ae50429bebc0add3423da9

Observation 7cc4a75e-4de3-402c-91f5-74a6f119ffdd · outbound

This paper cites an unresolved cited work.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:d8664cdacf729310bd6d9d750c26060164b8d7ee915125bff4ce333c03ca4657

Observation f874afa0-3754-46cb-baf0-d41c07837246 · outbound

This paper cites Science China Information Sciences , volume=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Science China Information Sciences , volume=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:e82f498914da743f6c22db547f4be5067f291b5f2b80882f756c01657a024bf2

Observation f47cf745-683e-4aeb-885e-26686afa4415 · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:b7c9b69b95281f286d14da3dfe3640870e5a32262a808a1c15fd7656e8799cde

Observation 3349c432-dd51-4aec-8b49-d46246e7641e · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Findings of the Association for Computational Linguistics: ACL 2024 , pages=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:f88c8d1d74854856c862d9b76898808d559fb44cf59ef972f056f48cb5897d5d

Observation 617aa287-d0c1-4577-8c7c-cf697a74090d · outbound

This paper cites What is the value of templates?.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis What is the value of templates?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:4a9fe5e1a9064a9973ab5100f940443a992dbc5e991a83798009719e0234d0f9

Observation e096066f-35a8-42d0-acc9-f6365e740d76 · outbound

This paper cites Proceedings of the 31st International Conference on Computational Linguistics , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 31st International Conference on Computational Linguistics , pages=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:b2bcd2c3bd2aeeed18c0d5b24ecab7a1e336dfef17c554988a7f5192bf11b07b

Observation 0a0b6f79-f7de-4eb5-a1c7-1a7a563765ab · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:64db3ce622578dee26bad04273914bf08b92416eac0eeffeb700c9f2c677bd48

Observation 39446f5e-caf7-45f2-aa57-bfa51104db5b · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:b563ee2dc42f2aff3377f726002b2709aaad028979d966474b342d675642c10d

Observation ab9818dd-9e45-4d69-97e6-bea0d7c99a2b · outbound

This paper cites proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:b16a40f06ca7f4875aefb40b43938adadbc22f753cda2e37e5639671a21a7420

Observation f98e27b1-3b68-4450-b015-796f12fdaf20 · outbound

This paper cites PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:4d8babae9b6506fdd98ffa96828ebaeae888603209a643865779d849530e9cec

Observation 7a8c586f-4968-4739-9513-665768efc30f · outbound

This paper cites Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:4365b80860c5ef5d76083ddc9d56d97a0232185d00414a50b4dcddc64ac8b6b0

Observation 0c2e35bb-ceda-4876-9000-5ee155ec70e5 · outbound

This paper cites ACM Transactions on Information Systems , volume=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis ACM Transactions on Information Systems , volume=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:7557cfb7c711e3cc4c134522cb0172ae4d23cc6660cb08f2e4e07bf8ee32ff9b

Observation ebc4903d-7c39-4c93-8029-f1dd11cf11de · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:caf691f7ad485ab31ee72f77c13327f0a644043ecbc57b03cee590a75147a785

Observation 39dfd1b2-5311-4f44-9747-3c1e16c3412d · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:61a8ab613c862118b50bfd38914f800c47d4f2297feba02da4ae4bbba3332e7b

Observation 3c3fcf04-7ac1-4163-b3c5-6a284f37f954 · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence , year=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis IEEE Transactions on Pattern Analysis and Machine Intelligence , year=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:98753532926f7a98aee826e9c05ef90f185a2d91b891c6f0349a76836959410f

Observation 5ee13aa9-d741-4473-b09c-e8cecc6c92a0 · outbound

This paper cites Neurocomputing , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Neurocomputing , pages=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:9a66c167fb79565cb428499a5afd98a0aee57260f377816a7db9099d4da8e18e

Observation 9fd0cc8f-eb94-4236-b34e-e085fcfb0a4d · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:a8a7a4c13fd185a0b8488e5f46b25ab2b3849fe8611e2cc60762292ae5821f55

Observation d9bbcf12-fca7-4257-8deb-b1fafe75e9ac · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Transactions of the Association for Computational Linguistics , volume=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:6947d63649a020cc18929eac0aa7089e53078d17dcddce0934ed21916b0b55d1

Observation 1858c4f1-d680-4c94-9da8-e312a78c829b · outbound

This paper cites arXiv preprint arXiv:2603.02789 , year=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis arXiv preprint arXiv:2603.02789 , year=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:8b900caf63afb68032a614d8533ef8ea0d7b97588d9ecc55531d9a19a0d89bf6

Observation f5ae56c4-d086-4170-8f15-54d0b6499a65 · outbound

This paper cites LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:864bec32621bcb3961027d03c93e59b1dfd7113f1fba4690dc74aaf0664ed88e

Observation a3d02ead-cb32-43d5-a40a-9d398a885f1b · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track) , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track) , pages=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:f9e62473bca4cbc4a5def534f63395c00b58b386df535b4c783fdb9fa9761134

Observation da2978a3-9010-44a8-a288-eb37c5ab59e7 · outbound

This paper cites Proceedings of the 31st International Conference on Computational Linguistics: Industry Track , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 31st International Conference on Computational Linguistics: Industry Track , pages=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:cfefa430c8ee58a5fd28c91e9f007b67beaaaee829e3f8333775da5253a0eb9e

Observation c5ce143a-9bbc-498e-9314-6309e84a9d51 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:8dc5abeb2928745999dfe91b0d8b0b554c776727f588f1a0f448106762a9718d

Observation 7f213cfe-e56b-4b8e-9a4c-bae28eeec14e · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2022 , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Findings of the Association for Computational Linguistics: EMNLP 2022 , pages=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:7d5f8a46b47245f3a1cb818320f048869c59daa3f6eaddd211aed5a2a3e6f8b9

Observation 27a59bed-e4b7-4dc1-9436-f8313f2ea1f7 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:67b5068e10ef3e9591e2303754bfacadf8c5eb90ac773926a5fdad568f4c3f66

Observation 3a93b484-37be-4d36-8508-c37503183b53 · outbound

This paper cites Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:405a5285b77535d569836e9258a1fa483a161b7aa14135ae5c7f93299e92f54b

Observation 9bc0b35d-3d41-4456-a97f-82808d6a7688 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:52acb69425addfb8388a494148b4ccc5f3aab6ba5b09f0395c2daac2391a7ad8

Observation 14e1b65e-6702-46bd-a881-e1522b53810a · outbound

This paper cites Proceedings of the 29th ACM international conference on multimedia , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 29th ACM international conference on multimedia , pages=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:ab2a1cffdec33f36896abc3f214208218146b803e7ad3c8f8903c416b1bba050

Observation 061a7170-8950-414a-b9aa-6f0c3db2d18a · outbound

This paper cites Proceedings of the IEEE/CVF international conference on computer vision , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the IEEE/CVF international conference on computer vision , pages=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:fd75e27842f82f2727fd08186a8b3f70f82ac7de51c2fb1c505dc1ec5cbd94c3

Observation 50c72099-510a-4ce6-8a79-ddcab52b9d24 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:9fbe6ef7a1d3104b49e94c54d12d4874a0e671693b03d79283d14e356f8bcccd

Observation 721ae9f2-287b-40ef-97b0-1ccda520ce2e · outbound

This paper cites International Conference of the Cross-Language Evaluation Forum for European Languages , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis International Conference of the Cross-Language Evaluation Forum for European Languages , pages=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:fa58032c7a43327578ef6c1bd9e7770cd9a28e8bfdfee61e322eb55517814b3a

Observation 574db71b-915d-4116-8bca-c02c3615638a · outbound

This paper cites Proceedings of the 28th ACM International Conference on Multimedia , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 28th ACM International Conference on Multimedia , pages=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:c8f88e80721406fc8ee92d5f52be28c1f2e7590bc7d6fc1d567e58870a847259

Observation 49d3d45c-2b53-4677-b1ed-39fe38c06088 · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:69f4b0da0ec18f862fe962effcb0c187f47dad58ef9dfdcc2d6db6736bdd2e29

Observation 7929a79c-a7b2-431a-8876-70ec2b407549 · outbound

This paper cites Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:267592a31fb4d21ac5499f3994e5248d2ece20f2c8015939233eeacffce9aab8

Observation 21b2f045-e8d7-46e9-bf68-28a254ea589a · outbound

This paper cites 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR) , volume=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR) , volume=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:83a613a10bd03e7e62abb8bc74057fe55f5996567cfc27aa18a46ef88445347e

Observation 5e94695e-e322-4c66-9669-c3e1a6881d4b · outbound

This paper cites Document AI: Benchmarks, Models and Applications.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Document AI: Benchmarks, Models and Applications

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:e3c3490fbcf95336e3563b683ad85515500933a8340d7b986df5c769f9eae664

Observation fd1d6e38-e3eb-4d3f-b386-7b889e8d2535 · outbound

This paper cites Proceedings of the 2010 ACM SIGMOD International Conference on Management of data , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 2010 ACM SIGMOD International Conference on Management of data , pages=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:a81a77c06e409f3d920a1b1eb0c0da0bb55adb2ae26652fbec5b19c253fd4c89

Observation 9ae07564-4d51-4af9-ac07-407a1c013a86 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:78019b8350ec6ba01f082823f39ec2edd88e00fab5d21f85719f3142bf162463

Observation 0ee47245-4372-4cd8-ad73-9736739c4b8a · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:c6a7764962edb159b4384b78bba52537d74bff37d24ea7ff32a256971a42b566

Observation f97b6b3f-7211-40cc-8c80-e08592c4dda4 · outbound

This paper cites Gemma 3 Technical Report.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Gemma 3 Technical Report

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:30135bb4c598ab2c105c8cf2b5f03ec99b424d3eabf09d605dac220c569c36b2

Observation 3bedee8f-71a7-4441-89a7-50208326a777 · outbound

This paper cites Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Industry Papers) , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Industry Papers) , pages=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:3b4a508815d01ac5ceb12d8fb432a89a340b0de3bf38582e136c91f1f070b65f

Observation 1d9315d9-99ba-434e-97ae-f6e63a56742d · outbound

This paper cites Knowledge and Information Systems , volume=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Knowledge and Information Systems , volume=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:a63ac2dbbf06c804d9d2030f3bb0fda9c4348fe2e1fcbfcb4ae5c03320734b69

Observation 3fd51263-f79e-4719-92f8-854b1a890db9 · outbound

This paper cites International summer school on information extraction , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis International summer school on information extraction , pages=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:60de577999d3c8ccac5783c4e369f0d6c71c30f5d7e5d652b4471bce76bc6ecd

Observation 2a438e36-0e83-4d7c-8070-751240fcdc1f · outbound

This paper cites Scaling Synthetic Data Creation with 1,000,000,000 Personas.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Scaling Synthetic Data Creation with 1,000,000,000 Personas

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:0a8a0927aba8d5ac4db75af50f8f82acd732683cc237d38aa0324eff5248a311

Observation b5804c0e-2c5b-4c75-a234-8e016fbdd720 · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis SmolVLM: Redefining small and efficient multimodal models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:c76b26405ad74a1f6f5046bd1544419b7e51f3c01d171ef4eb3247d57fb25f48

Observation 06b41558-71e1-4941-8da9-1bcf0e6f2b8b · outbound

This paper cites Ministral 3.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Ministral 3

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:ebfede4db06d3ac00239735722b8f9ddb3af73178a38c7b9f2ab89d72212c182

Observation 97cce45f-4119-4c8b-930a-ea8867f2d211 · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:f74c9ea84ed52e1d8fcbe3e636e65a9122d7a20d9e801dfa9c2e3d1417414cd7

Observation b7922922-8c34-402f-a0c2-46c49115aad3 · outbound

This paper cites Ovis2.5 Technical Report.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Ovis2.5 Technical Report

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:cc90807be9b2c0a9657dc9f55a746576186ab7107b3fc9ee5945d6eb36fe0b4a

Observation 0b978ad5-3159-474f-a6c1-7a0e74e04f8b · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:f23a5f85a0d7b749d60bf059ab43a52e091d95796c6ca44fdc277aec04717b5e

Observation d75b02fd-d8ad-4683-a33d-658a27bf25e1 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:c7f8a11aea130379b88e89e3e907efb674a84fb7d0f63a0c971818720452e351

Observation 45dcfd5a-c49d-4bb1-b805-1baccebc06cb · outbound

This paper cites MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:2805b5380b2bb97a04fa8c175231ed36559aaffa43895b383de940e9881b8855

Observation 73a63745-09c4-41f7-82a3-190adfbab1c9 · outbound

This paper cites Qwen3-VL Technical Report.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Qwen3-VL Technical Report

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:30df6900131a890489bbe913f2d653f0ba339d94ba2c29b01d1d5128136638bf

Observation 7f95a342-630d-439e-a2d3-761865503734 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:a643746c152a7f2d77abc01cb920a1f9a35e7f79d35f955bb94c4153ad55f535

Observation 177c8529-0cf6-41d0-8df9-73ecd4283e13 · outbound

This paper cites Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages=

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:f24d3512de864f2e133fda6490b82599087757c03a5567d9a66f5d63b49e546a

Observation 9d589030-9f6a-4dbe-b3e4-0323bc62c9eb · outbound

This paper cites Journal of Pathology Informatics , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Journal of Pathology Informatics , pages=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:8580a0b3fc9e247ad50fe0dd88590bbb471e936db3e7f78d50bd3eefa4f5b719

Observation 3512e7eb-727d-4cd3-a658-5fa24d3f93dd · outbound

This paper cites CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:d81dcfbab6d1effc61c82267e5d3a41a8f44cc9c06d72676364216ac92d61692

Observation 46bd0a45-2bba-4652-864a-47c828d10b57 · outbound

This paper cites Findings of the association for computational linguistics: ACL 2022 , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Findings of the association for computational linguistics: ACL 2022 , pages=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:5f70058bde37e2539cce68e14b8fea1d5d3029410c0b1c620287168613cd59c0

Observation 124af223-04e4-4b4e-89fa-723d0d2eef33 · outbound

This paper cites International Conference on Document Analysis and Recognition , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis International Conference on Document Analysis and Recognition , pages=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:bbd30aa1dc48b80266405a2e99695db01987cf8eeb558ec11bd28f963bd27f2a

Observation 97fdfa06-0c64-4e70-bd86-28b2249fe0c2 · outbound

This paper cites Workshop on Document Intelligence at NeurIPS 2019 , year=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Workshop on Document Intelligence at NeurIPS 2019 , year=

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:d386c48339badfe34cf51f7b34e93b3a1ed81d7a8ee36b3e088abe82bf1f78fd

Observation c9693407-6e6c-4aa3-b561-3ce0e37890d5 · outbound

This paper cites Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining , pages=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining , pages=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:a8e54ec14786b0ee5aa99a4bc4e642dc1b8a4a67745b343ae53cc6915c5d82fc

Observation cc362501-73a8-4172-b17c-eddfad69ac7d · outbound

This paper cites Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:bd51f9f6e27fc3eb224dc54607f0dc4120ce3cc09cfe00f721ff8d6eda4e616b

Observation e18c5deb-f0e2-4abf-85f8-f618d039274a · outbound

This paper cites arXiv preprint arXiv:2510.19817 , year=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis arXiv preprint arXiv:2510.19817 , year=

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:1298891053514a0c0dcb7d6ae43e979b86b55825000620cc94d782bb9066b18b

Observation 60305a1d-27b9-434a-8d52-3db4582ce57f · outbound

This paper cites Advances in neural information processing systems , volume=.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Advances in neural information processing systems , volume=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:5a3d27e19da315cd889d28741f930d3019e7a1af1bb3172f88d1f16bfa48117e

Observation 853aebbd-9bce-4d20-a075-9f4988e918ad · outbound

This paper cites 2018 IEEE International Conference on Cloud Engineering , pages =.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis 2018 IEEE International Conference on Cloud Engineering , pages =

Reference 90

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:5521edbbc08dbb4f3e6b553f9a3bc16d51e7c4f47bbdc17636b6697092098664

Observation a6a4146e-2871-480d-a0b6-75f64e2e9633 · outbound

This paper cites Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems , pages =.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems , pages =

Reference 91

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:2581637762549b1eb59ab77ec137b06f95a254842136e87c0414ceacd6662bb5

Observation eeac0cbd-e93e-4932-a5fa-0cef14505ce1 · outbound

This paper cites Key Information Extraction From Documents: Evaluation And Generator.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Key Information Extraction From Documents: Evaluation And Generator

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:0a77b564bc18e8924074a21709e0d5d6c59b877365ac65d25520e29acfe37c28

Observation 5dc9a41b-5099-442e-8556-3bb6023f390a · outbound

This paper cites DocILE Benchmark for Document Information Localization and Extraction.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis DocILE Benchmark for Document Information Localization and Extraction

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:9e2c29a4e488a862e9b5c7d54e85835c418f83ec065a8cd7a2496f6c91a53d75

Observation 44b164d9-fe27-4182-b946-20d0e1c30eaa · outbound

This paper cites 2025 , address =.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis 2025 , address =

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:dc950b31f76e259b84309e25374b014fdca477e29739ed89ae68eba0af97c0a8

Observation cc60c917-5587-41fc-9b2b-bc38eeb90a4f · outbound

This paper cites MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining

Reference 95

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:8806641bc034f23d3b733682e4b6bbb1a3e43ed3d2c98380865cdaa53c1dc524

Observation eb800356-a325-42de-aa51-2fc6318ca1ff · outbound

This paper cites ChatGLM: A Family of Large Language Models from.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis ChatGLM: A Family of Large Language Models from

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:6aaccb440e825e77935e0609c10e348df81acaded680902b94cb999d60260c9f

Pith citing papers

No inbound Pith citation observations are available.