Pith. sign in

Paper Citation Record · LEDGER

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers

As of 11 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 2 inbound Pith citation observations for arXiv:2501.16297.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16297 v2

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T13:38:18.919283Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T23:09:32.594194Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T23:14:01.665133Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9cecb0cd-669b-499c-b5d1-611b54e4602a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.660902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.660902Z digest=sha256:e99ce2a5287d926b1ce49e79679badc9808f897e4a090b56644fcaa1f198649e

Observation 2930fe71-38c5-4dd1-96c0-cf4eb91f3fa8 · outbound

This paper cites Lion: Empowering multimodal large language model with dual-level visual knowledge.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Lion: Empowering multimodal large language model with dual-level visual knowledge

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.650731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.665980Z digest=sha256:1e6e32b6ca45e37639c8f7b35fda5049c20634a29727fa17495ede3f9bcf0043

Observation 00b8b4a7-9f6e-4a5b-9383-1862af359e04 · outbound

This paper cites Less is more: Empowering gui agent with context- aware simplification.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Less is more: Empowering gui agent with context- aware simplification

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.639220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.670578Z digest=sha256:c3e0ebdf2ad4eac4388936e44f0aa39e41a88f3c78da41f907afb7f5721cb83d

Observation 06f4e848-242b-4446-baae-33a9dde9139b · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.674964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.674964Z digest=sha256:c7e54a637d110e5254a81aaa33c8a6e31720dd0fee2e28aed9663d2b24074aad

Observation 48e9b2e3-80b4-4576-bacf-75157864c99b · outbound

This paper cites Spa-bench: A comprehensive benchmark for smartphone agent evalua- tion.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Spa-bench: A comprehensive benchmark for smartphone agent evalua- tion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.628063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.679118Z digest=sha256:507f011888b3015e2b04c818feda37d88a8bf4be1f97345f177360f5fb1111c9

Observation 4a044303-7ddb-4633-a91d-549cd3158b58 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Sharegpt4v: Improving large multi-modal models with better captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.683251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.683251Z digest=sha256:3d0d3c61f687abe7c292f6b2e744d0ef0dc64746a17726294ff3973ae15da64b

Observation 189dacbe-1e7a-47f2-b301-7d58e427a451 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.610545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.687230Z digest=sha256:5750bc19c8d0c5a6b42666d2c6cc2548865b7b48e883554b8e6b3f670ff83ece

Observation f4d8f2bb-447e-48e7-b1a4-6d90bfe78e25 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.692030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.692030Z digest=sha256:4b5c949113fdc05a12063f0ed08db75aaea5bdcc847cfc8b9e90560e18ee25c5

Observation 86a2d164-30f3-4be8-878c-2ed042bca62e · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.592111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.696051Z digest=sha256:1456c8b5b615ea905bf77ad07b76b28d2e1bccda5fc3792731a3519d9799b485

Observation 54b7c54b-407f-42a3-a4d7-cae40c478b81 · outbound

This paper cites Vision transformers need registers.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Vision transformers need registers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.580813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.699899Z digest=sha256:29611dfaf14047050e4a0b4db81a77a48f1aff621567ca71e65c231eb77abb56

Observation b05491e6-12f0-4fae-b022-4e138d0b0fde · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers An image is worth 16x16 words: Transformers for image recognition at scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.703723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.703723Z digest=sha256:f765958d0337d20110ab116145c1f4b976aa7f87117c99a9d2045f2d93944f17

Observation 86885e08-9b0e-4839-b105-0bd5e9780b93 · outbound

This paper cites The Llama 3 Herd of Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.707825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.707825Z digest=sha256:87d8472531eda5d94f773c01ddfa132f8f02613531421aa86cd80f5e8e7b053f

Observation f0b27731-56bc-4aa2-b96c-c85fdb6faf2f · outbound

This paper cites Gaussian Error Linear Units (GELUs).

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Gaussian Error Linear Units (GELUs)

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.711903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.711903Z digest=sha256:3fa9b844966ce9f291102b83b549f1a3b7022e836222d60aa1b869539c27056f

Observation ac01ab06-dbdb-4897-a3e3-12b933609d0d · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers CogVLM2: Visual Language Models for Image and Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.716137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.716137Z digest=sha256:61e4980d09480c58768882b97ed030b245ed878a8372d6726aaae31ef064e2cd

Observation 33395e58-9e1a-40ff-842f-eee1d6082bc2 · outbound

This paper cites mplug-docowl 1.5: Unified structure learning for ocr-free document understanding.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mplug-docowl 1.5: Unified structure learning for ocr-free document understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.563433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.720321Z digest=sha256:6439b4435c1af29eb26c6cbc53bfcf804eaf32bf9c5eb07362ab6a39617f2449

Observation 905014d2-2ac6-4584-bced-f09eaa7de993 · outbound

This paper cites Mini-monkey: Alleviating the semantic saw- tooth effect for lightweight MLLMs via complementary im- age pyramid.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mini-monkey: Alleviating the semantic saw- tooth effect for lightweight MLLMs via complementary im- age pyramid

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.551976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.724048Z digest=sha256:ebb4fe6f5fca59a64425e501579f2cc57ead151846d8722e76b8da4cca2877a6

Observation 156e88a4-8f94-4638-83c2-989d28bd2371 · outbound

This paper cites Hires-llava: Restoring fragmen- tation input in high-resolution large vision-language models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Hires-llava: Restoring fragmen- tation input in high-resolution large vision-language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.541358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.727665Z digest=sha256:a2be41b4392521a1f59143c3edb600cb67dd415ce390a62fc19b7501e69640ac

Observation 6493d439-31b1-4e22-815a-5bd29d58c30c · outbound

This paper cites Openvla: An open-source vision-language-action model.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Openvla: An open-source vision-language-action model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.530445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.731361Z digest=sha256:4c8b0218d9fb8bcc40090b5c12aa9c7e09ca42abb60bf2f984cdda204f59c0ae

Observation 057861c3-7ed6-4acb-943c-0fc7d84e2852 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.735094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.735094Z digest=sha256:373225d5740d4975326a01a30f1be1714ead3083cbca3d4b9239fdbcc3867f3e

Observation 01559c1a-9d54-4159-b24d-a22c4e2cb7e2 · outbound

This paper cites LLaV A-onevision: Easy visual task transfer.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers LLaV A-onevision: Easy visual task transfer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.739410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.739410Z digest=sha256:320eaacd0e3c5fd6873e29497dad7f6526e2d44d1d4f3981085a10969a6561d3

Observation dbaafdc9-d0bb-4a5c-8a55-5eb8ee60a24b · outbound

This paper cites STAR: Learning diverse robot skill abstractions through rotation-augmented vector quan- tization.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers STAR: Learning diverse robot skill abstractions through rotation-augmented vector quan- tization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.513557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.743525Z digest=sha256:527acb39ae14c8bd436e2d1e483856463b7c6979f54545ee9d751312692bb2fb

Observation 9773fdbd-b6ca-4420-b586-1ed3a5cdc433 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with 11 frozen image encoders and large language models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Blip-2: Bootstrapping language-image pre-training with 11 frozen image encoders and large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.502224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.747331Z digest=sha256:9fdb6d01e86e991ba945254c785b0d8107c484b8b44e614b429841e1c89194ab

Observation 05ab51b1-b59b-4d17-88f1-e36a767cf6f2 · outbound

This paper cites Flex- attention for efficient high-resolution vision-language mod- els.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Flex- attention for efficient high-resolution vision-language mod- els

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.491062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.751038Z digest=sha256:a7cad8db6f5522145e3850735abdd0dfe037274e740368c2a98ab79f68721fac

Observation 3898cb60-9eb4-4ae8-ba70-2caf4890a0de · outbound

This paper cites Lion-fs: Fast & slow video-language thinker as online video assistant.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Lion-fs: Fast & slow video-language thinker as online video assistant

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.480351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.754766Z digest=sha256:6f2359e2339749f8132e4dddb0a1b625a2f16d9386f2db848612a94254b7e765

Observation 92305da7-4122-4711-a3e1-2bb26cb96def · outbound

This paper cites Evaluating object hallucination in large vision-language models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Evaluating object hallucination in large vision-language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.758353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.758353Z digest=sha256:412875e4a3d08dd3eb83c44373661b30117d0cb972ec25a50e690e5f11f0c797

Observation a1862721-228f-481c-a619-8369a8c60462 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.761893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.761893Z digest=sha256:06d69d24e85bf0b85a9c2f7552b80657d96fcdfff1a295015ca927e0f5bbbda9

Observation aa7da1ae-4093-4815-88dc-9d1598f5cd82 · outbound

This paper cites Optimus-1: Hybrid multimodal memory empowered agents excel in long-horizon tasks.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Optimus-1: Hybrid multimodal memory empowered agents excel in long-horizon tasks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.463546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.766280Z digest=sha256:4c523a80f7e70a9a59610acda43fb87430b97f40ec8f43ca6c6a444016eb305d

Observation 0aa62c3a-5c79-47be-b161-5eb2c705ed67 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.452524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.770199Z digest=sha256:cbc036ea2d84448746e85b59c5060fe4dcab35c597ec5d45686daa33cdf28733

Observation 75e8e79f-8633-4bed-bfe3-4ff0824dd92f · outbound

This paper cites Optimus-2: Multimodal minecraft agent with goal-observation-action conditioned policy.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Optimus-2: Multimodal minecraft agent with goal-observation-action conditioned policy

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.441573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.773810Z digest=sha256:dcd76bd6c277f0d31869d27062b1016e8acb6c86a36e835db6314fba3f946787

Observation aeb7aca1-8e46-4cf1-bbf9-a55e1f82e333 · outbound

This paper cites Improved baselines with visual instruction tuning.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Improved baselines with visual instruction tuning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.430737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.777251Z digest=sha256:82b00788b20ccf391cef6400f9cfc97c15ae35c3412b55a3d3c023a99b32b940

Observation 94754f3a-0b0a-49e2-ab02-bb33b379258c · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.780931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.780931Z digest=sha256:25c0b5457f5546fc3410c41ae90551e1296ed733f85f8db92c90bcdf83b10bdf

Observation 403f895a-a8e2-4982-b704-cf18734199f1 · outbound

This paper cites Visual instruction tuning.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Visual instruction tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.784750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.784750Z digest=sha256:20769ec9be28982cafa9ec8bc29d3028b85acd7fb2340adc16f4c5d190261f4d

Observation c8afdfd2-f1ab-45df-81a9-258884492a40 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.788312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.788312Z digest=sha256:e2d051061b4dd0f9ce5e338070a6b8fc40ddcad00ac048f3741623de1d3e18d6

Observation 99251e67-9df0-41a4-970d-90e9a6a68bf0 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.792560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.792560Z digest=sha256:51b8c76154ae584ffd62b3a7f6c76c70836cff677ad269d95fb64a5a15096fa0

Observation 718f2c8d-4842-4d67-9520-ebb223d1e3b8 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Swin transformer: Hierarchical vision transformer using shifted windows

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.796210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.796210Z digest=sha256:77f5ea4f602d16f87102e5be1b453024f30f313caae38e32582c9fcd7233a452

Observation 80466512-1ad5-4253-a94a-96d2ca4629bb · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.800215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.800215Z digest=sha256:36587f3275d8ec8ccfda3028a708b4d0ce859bc6df76032c5bc4021b4ba66804

Observation b5245f0b-10fa-4f44-b04f-5b197bbd3d48 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.394658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.804301Z digest=sha256:50cba174a6b038210afead6eca11114b43fad3cae1d239620fe4575007b37d75

Observation 5c6f3f2f-88ce-4b28-aa9a-a977052721d7 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.807983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.807983Z digest=sha256:21f77091bd40e2d7fdda18538fc578e56294532ae365d7e72a2e54a36bd85dbd

Observation 4dfc0037-2743-484f-bd91-0de059dd5dd2 · outbound

This paper cites Spatial-temporal graph diffusion policy with kinematic mod- eling for bimanual robotic manipulation.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Spatial-temporal graph diffusion policy with kinematic mod- eling for bimanual robotic manipulation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.383566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.812074Z digest=sha256:8c5731c4dd58ed48e0ccfc617fdd6f990bea6f7e6774ee828b021136e30becb7

Observation b58f3cd1-f15f-4198-b84a-0ab867f706a4 · outbound

This paper cites Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.372698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.816038Z digest=sha256:bed6c5a27aa43811ae1bc161ae24e489fdb8f247a5c2a1fbf39acbc2219cd150

Observation fcf03e1a-d56a-426f-878e-f1710aed48e2 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Docvqa: A dataset for vqa on document images

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.361057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.819840Z digest=sha256:a64d8ad3a75c5f671f83ea3d76bcbf68cd1051231793c550ed389e4a9d6743ab

Observation 2223dd60-7822-4ab6-8de0-6f1d3d81404e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Learning transferable visual models from natural language supervi- sion

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.823527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.823527Z digest=sha256:bfcfec7b6ee156943246de67e64902c27a51e516de549dd5f97cf4ae59a255b5

Observation 68d28a54-be9a-4828-b622-3bef84431d87 · outbound

This paper cites Multi-adversarial discriminative deep domain generalization for face presentation attack detection.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Multi-adversarial discriminative deep domain generalization for face presentation attack detection

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.827229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.827229Z digest=sha256:263159d4a8656bdd106314a9cb530b5003d43b57bc98f216654c9c448942d8d9

Observation 5682eb4d-25ed-4d0c-909a-1a2eb617661c · outbound

This paper cites Detecting and grounding multi-modal media manipulation.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Detecting and grounding multi-modal media manipulation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.337795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.831549Z digest=sha256:d57cbfd73d46e17597a97c8eb53b8b47d9d3c5a832d6df81fe6784538a417ff6

Observation 1ee6b5e4-3a72-463e-91df-12ac496afa7a · outbound

This paper cites Detecting and grounding multi-modal media manip- ulation and beyond.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Detecting and grounding multi-modal media manip- ulation and beyond

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.327055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.835497Z digest=sha256:b40abecdc0319f428b39a051219ca73533f6fcc7db4028a9a1928d3edf555965

Observation 302dd13c-0a73-4068-a2b9-490798c2f71d · outbound

This paper cites Mome: Mixture of multimodal experts for gen- eralist multimodal large language models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mome: Mixture of multimodal experts for gen- eralist multimodal large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.315809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.839306Z digest=sha256:39f123842517280d112b23a7812d21529177dbb3d3c7fa18e500b5ee2d1b520f

Observation 5de446c3-1946-4748-85a2-4b8b58af2f6c · outbound

This paper cites When do we not need larger vision models? In 12 European Conference on Computer Vision, pages 444–462,.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers When do we not need larger vision models? In 12 European Conference on Computer Vision, pages 444–462,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.304740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.843134Z digest=sha256:3900ee9510902768e5d19af62357e737e82c01b082966a30d79dccf3b1cf8bba

Observation f9db3667-dfaa-4807-a952-860ebba09887 · outbound

This paper cites Towards vqa models that can read.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Towards vqa models that can read

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.847335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.847335Z digest=sha256:2a73ab9ab118f90db64a7634b66a1e317fec171fb4414615b19ab6820a3a2065

Observation 36153ffb-5247-473e-a02a-992f25713644 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Gemini: A Family of Highly Capable Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.851125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.851125Z digest=sha256:925848ed0c85cdd5fe45c30c42df139733faef21038b06fb8d74f881f5382cb5

Observation 279cc74f-b201-40d4-95ba-30f69f6556f0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers LLaMA: Open and Efficient Foundation Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.855123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.855123Z digest=sha256:f389b21ea36716889d1b616d01e3ac9c959bac80e372bce247f810439869bf17

Observation 41e6be49-fde2-4170-950c-dea735c419e7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.858993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.858993Z digest=sha256:3515a91f4b02ee48a94db9a03afba84da2b3f69e3be72cccc250d8d43b5fd098

Observation 5b9a5693-bae5-4f98-b923-48684cfba1b9 · outbound

This paper cites V*: Guided visual search as a core mechanism in multimodal llms.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers V*: Guided visual search as a core mechanism in multimodal llms

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.287477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.863115Z digest=sha256:fb872c0686c8b5e5da9162fd603db0e89a7465c5cf4777f4dbcd7df542e3d23a

Observation 8caefd48-b66b-45fd-98dc-526a0336d167 · outbound

This paper cites Gui-explorer: Au- tonomous exploration and mining of transition-aware knowl- edge for gui agent.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Gui-explorer: Au- tonomous exploration and mining of transition-aware knowl- edge for gui agent

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.276332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.867100Z digest=sha256:647b57e5771fde944c1a3e6fe74c606594af70b3ff3cc1a7b52277369fc01489

Observation 369d4a1a-0c92-4c72-b617-b8716adcf56d · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.870758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.870758Z digest=sha256:b76bcc0a55ca9274e41647aae10ce162ab2d2124d55e315c536723856260f9bf

Observation 9bd10117-1474-4938-a48d-c544a1805316 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.875541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.875541Z digest=sha256:01303f643f688ca16dad76ecef3013ed1449626572105b8af88a7fe3a0ec7409

Observation 2e866717-1d12-4ccf-bf1b-c775db334c8d · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.879391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.879391Z digest=sha256:f23f7a7b336ea97262211e3de8b3fcfbe14963f877e54fb7f05ff8466c1266de

Observation 57e8c7ba-2775-43f7-8808-f8bb8be2a97e · outbound

This paper cites Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.264231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.883290Z digest=sha256:71ff4be1a88ba52070f3dfd919d9ebe8b0bc24d4f5cdf48e957dfb83002fc893

Observation 3908d8fc-4886-4a77-916b-13a650794a9b · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.886964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.886964Z digest=sha256:1283f13f3742fb6d1a707af754ceea7a7daed30a80bbafc567a265efcc1796f4

Observation d570105d-11f2-4813-bad3-6a832874fa46 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.239916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.891094Z digest=sha256:cb213e2e2385d904200686caaf670e5dd67085291c23a9a452d4b8b0fb329717

Observation 95e99493-16af-43f9-aacd-7f06afa2fa0d · outbound

This paper cites Modeling context in referring expres- sions.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Modeling context in referring expres- sions

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.228709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.895145Z digest=sha256:991f22df92f6466a40860f473c4df97345e82149f4449cd07035ad1d44a4dcd3

Observation 288adfb4-a5bb-43f1-96d3-a928edd16fbe · outbound

This paper cites Sigmoid loss for language image pre-training.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Sigmoid loss for language image pre-training

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.216519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.899089Z digest=sha256:1d45e12055c85e067d59a1ff76e5d4dbe2f2f061cd6f6be3c9f736cfcf0d044d

Observation 7f24491c-e4a3-4866-b6fe-c2441f230af3 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.903154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.903154Z digest=sha256:e089b350af26573094a54dee946b1264bcdccc386182ef7fb00181d53b1c4d6c

Observation 8fbc5767-86a4-45e6-85df-5737da82d995 · outbound

This paper cites Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.907427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.907427Z digest=sha256:e9b9a7629ac0c3a01b22ccf137a72f0c2dd2a0c6ab79ecbc8f920a7cc65cf63c

Observation ac5d1f9d-a70b-4c35-8dde-30a233ecbc2c · outbound

This paper cites From redundancy to relevance: Enhancing explainability in multimodal large language mod- els.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers From redundancy to relevance: Enhancing explainability in multimodal large language mod- els

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.199802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.911600Z digest=sha256:4f4c146b1529fb89d6129bb85654019404ee97fb7b9e79094a4a77dafa56458e

Observation 5a52d847-6c2a-401f-a6d0-37d6db3f6f4f · outbound

This paper cites an unresolved cited work.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:38:19.186362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T13:38:18.915446Z digest=sha256:de063cfcd78f86db83aae5a8bb54b46c945282734574bf72888554920935eab2

Observation b382da80-0c94-48de-b22d-f2e8fd4ba5eb · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.919283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.919283Z digest=sha256:da887b798e845e440b8be3ba826469b5bf4a4b4bea7f2e0d83722ca2f331a49f

Pith citing papers

Observation 1e720cc7-46b5-45e6-8500-60bd9775742e · inbound

UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries cites this paper.

UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:15:33.790283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T08:14:28.540629Z digest=sha256:90ec37d6fb8d89bdb98a85f2db8026c21dd705664807ccb1f24331138122ce9d

Observation 308599d9-d0d2-4720-a799-43695fdad954 · inbound

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models cites this paper.

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:01.667217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T23:09:32.594194Z digest=sha256:0d6b51aa1c319503b16029ccda76b76869fd3e814dfd2e58728392513ee4fd5b