Pith. sign in

Paper Citation Record · LEDGER

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers

As of 12 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 2 inbound Pith citation observations for arXiv:2501.16297.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16297 v2

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T13:38:18.919283Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T23:09:32.594194Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T23:14:01.665133Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9cecb0cd-669b-499c-b5d1-611b54e4602a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.660902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.660902Z digest=sha256:e99ce2a5287d926b1ce49e79679badc9808f897e4a090b56644fcaa1f198649e

Observation 2930fe71-38c5-4dd1-96c0-cf4eb91f3fa8 · outbound

This paper cites Lion: Empowering multimodal large language model with dual-level visual knowledge.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Lion: Empowering multimodal large language model with dual-level visual knowledge

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.650731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.665980Z digest=sha256:6340519308727938acfba0edf40bc39def59ebb018b3a1e9c30a271468d02d53

Observation 00b8b4a7-9f6e-4a5b-9383-1862af359e04 · outbound

This paper cites Less is more: Empowering gui agent with context- aware simplification.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Less is more: Empowering gui agent with context- aware simplification

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.639220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.670578Z digest=sha256:44ebf2444d0a7235c11a8aceee70b666ba992aab00c845d11a3d672b5840314b

Observation 06f4e848-242b-4446-baae-33a9dde9139b · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.674964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.674964Z digest=sha256:c7e54a637d110e5254a81aaa33c8a6e31720dd0fee2e28aed9663d2b24074aad

Observation 48e9b2e3-80b4-4576-bacf-75157864c99b · outbound

This paper cites Spa-bench: A comprehensive benchmark for smartphone agent evalua- tion.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Spa-bench: A comprehensive benchmark for smartphone agent evalua- tion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.628063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.679118Z digest=sha256:32c9d2d5b48a57a63ba402ded8d5ef83274976f46806838bd8721d33b572fdaa

Observation 4a044303-7ddb-4633-a91d-549cd3158b58 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Sharegpt4v: Improving large multi-modal models with better captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.683251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.683251Z digest=sha256:3d0d3c61f687abe7c292f6b2e744d0ef0dc64746a17726294ff3973ae15da64b

Observation 189dacbe-1e7a-47f2-b301-7d58e427a451 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.610545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.687230Z digest=sha256:8eaa56f3a278e4e2668962dacd35a9de9f7a8a40b4377a4cf8588f6c9b2e2eeb

Observation f4d8f2bb-447e-48e7-b1a4-6d90bfe78e25 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.692030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.692030Z digest=sha256:4b5c949113fdc05a12063f0ed08db75aaea5bdcc847cfc8b9e90560e18ee25c5

Observation 86a2d164-30f3-4be8-878c-2ed042bca62e · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.592111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.696051Z digest=sha256:c3842840d50f8b9c8368d5fed5103f4576a2187c018fa262bdcdef4adbcf6a7b

Observation 54b7c54b-407f-42a3-a4d7-cae40c478b81 · outbound

This paper cites Vision transformers need registers.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Vision transformers need registers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.580813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.699899Z digest=sha256:8a8031af2960cea590816dffe2586356a9774f3f32775e0d46ec81def7d3f042

Observation b05491e6-12f0-4fae-b022-4e138d0b0fde · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers An image is worth 16x16 words: Transformers for image recognition at scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.703723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.703723Z digest=sha256:f765958d0337d20110ab116145c1f4b976aa7f87117c99a9d2045f2d93944f17

Observation 86885e08-9b0e-4839-b105-0bd5e9780b93 · outbound

This paper cites The Llama 3 Herd of Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.707825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.707825Z digest=sha256:87d8472531eda5d94f773c01ddfa132f8f02613531421aa86cd80f5e8e7b053f

Observation f0b27731-56bc-4aa2-b96c-c85fdb6faf2f · outbound

This paper cites Gaussian Error Linear Units (GELUs).

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Gaussian Error Linear Units (GELUs)

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.711903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.711903Z digest=sha256:3fa9b844966ce9f291102b83b549f1a3b7022e836222d60aa1b869539c27056f

Observation ac01ab06-dbdb-4897-a3e3-12b933609d0d · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers CogVLM2: Visual Language Models for Image and Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.716137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.716137Z digest=sha256:61e4980d09480c58768882b97ed030b245ed878a8372d6726aaae31ef064e2cd

Observation 33395e58-9e1a-40ff-842f-eee1d6082bc2 · outbound

This paper cites mplug-docowl 1.5: Unified structure learning for ocr-free document understanding.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mplug-docowl 1.5: Unified structure learning for ocr-free document understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.563433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.720321Z digest=sha256:97744976ce10b1b7c8b5b06bcddae7da912d1063fcc51b6d6523772b4b54536f

Observation 905014d2-2ac6-4584-bced-f09eaa7de993 · outbound

This paper cites Mini-monkey: Alleviating the semantic saw- tooth effect for lightweight MLLMs via complementary im- age pyramid.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mini-monkey: Alleviating the semantic saw- tooth effect for lightweight MLLMs via complementary im- age pyramid

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.551976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.724048Z digest=sha256:7c871050d9ae442beaf9ff2a1f2fb5c959e64800613282b8f8ae8ef445d1691f

Observation 156e88a4-8f94-4638-83c2-989d28bd2371 · outbound

This paper cites Hires-llava: Restoring fragmen- tation input in high-resolution large vision-language models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Hires-llava: Restoring fragmen- tation input in high-resolution large vision-language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.541358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.727665Z digest=sha256:ededda3f80dfafc2b25f6bb2990732df68874997f8d053ac21c997185f7a827d

Observation 6493d439-31b1-4e22-815a-5bd29d58c30c · outbound

This paper cites Openvla: An open-source vision-language-action model.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Openvla: An open-source vision-language-action model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.530445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.731361Z digest=sha256:45bf2483f700387f988b507342bc703c3f776b7185e99d9cee4c17ccb6ebd6c3

Observation 057861c3-7ed6-4acb-943c-0fc7d84e2852 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.735094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.735094Z digest=sha256:373225d5740d4975326a01a30f1be1714ead3083cbca3d4b9239fdbcc3867f3e

Observation 01559c1a-9d54-4159-b24d-a22c4e2cb7e2 · outbound

This paper cites LLaV A-onevision: Easy visual task transfer.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers LLaV A-onevision: Easy visual task transfer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.739410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.739410Z digest=sha256:320eaacd0e3c5fd6873e29497dad7f6526e2d44d1d4f3981085a10969a6561d3

Observation dbaafdc9-d0bb-4a5c-8a55-5eb8ee60a24b · outbound

This paper cites STAR: Learning diverse robot skill abstractions through rotation-augmented vector quan- tization.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers STAR: Learning diverse robot skill abstractions through rotation-augmented vector quan- tization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.513557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.743525Z digest=sha256:2d45a10e629cf38545c7916dc60b28ce43cce3fe77410e01f0a0a9308f0f43ed

Observation 9773fdbd-b6ca-4420-b586-1ed3a5cdc433 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with 11 frozen image encoders and large language models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Blip-2: Bootstrapping language-image pre-training with 11 frozen image encoders and large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.502224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.747331Z digest=sha256:af9402b01a4ea8cf2f980b51f8904dc1177485cf61a83a9e4b42fe147d5ae14c

Observation 05ab51b1-b59b-4d17-88f1-e36a767cf6f2 · outbound

This paper cites Flex- attention for efficient high-resolution vision-language mod- els.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Flex- attention for efficient high-resolution vision-language mod- els

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.491062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.751038Z digest=sha256:455131f082120db99112c7db7e24dc12158dbfb37ce87fe6ab574ab76395e60c

Observation 3898cb60-9eb4-4ae8-ba70-2caf4890a0de · outbound

This paper cites Lion-fs: Fast & slow video-language thinker as online video assistant.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Lion-fs: Fast & slow video-language thinker as online video assistant

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.480351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.754766Z digest=sha256:0d6f59ccb74e0966ce43a0587613892d4640541bc89c1ceedcfa42fd0f33f644

Observation 92305da7-4122-4711-a3e1-2bb26cb96def · outbound

This paper cites Evaluating object hallucination in large vision-language models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Evaluating object hallucination in large vision-language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.758353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.758353Z digest=sha256:412875e4a3d08dd3eb83c44373661b30117d0cb972ec25a50e690e5f11f0c797

Observation a1862721-228f-481c-a619-8369a8c60462 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.761893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.761893Z digest=sha256:06d69d24e85bf0b85a9c2f7552b80657d96fcdfff1a295015ca927e0f5bbbda9

Observation aa7da1ae-4093-4815-88dc-9d1598f5cd82 · outbound

This paper cites Optimus-1: Hybrid multimodal memory empowered agents excel in long-horizon tasks.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Optimus-1: Hybrid multimodal memory empowered agents excel in long-horizon tasks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.463546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.766280Z digest=sha256:3f8859e4a9dc8e118aef392111c10f98ffa2da88ff119a53f4d7eafd809583d8

Observation 0aa62c3a-5c79-47be-b161-5eb2c705ed67 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.452524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.770199Z digest=sha256:29db4b160fca10d1ddefb26d74618b507f6c69bfdf759ab8f56acdb680f53ad6

Observation 75e8e79f-8633-4bed-bfe3-4ff0824dd92f · outbound

This paper cites Optimus-2: Multimodal minecraft agent with goal-observation-action conditioned policy.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Optimus-2: Multimodal minecraft agent with goal-observation-action conditioned policy

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.441573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.773810Z digest=sha256:a6b30f947ea54a535b49984d73a5a446df1b19c59281ee8ddc4d327bfb81b3f4

Observation aeb7aca1-8e46-4cf1-bbf9-a55e1f82e333 · outbound

This paper cites Improved baselines with visual instruction tuning.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Improved baselines with visual instruction tuning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.430737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.777251Z digest=sha256:fee5bb9ad948ca16f2c3efda31178e7e66e0d791719c6b7abf986ddc7e3a9a50

Observation 94754f3a-0b0a-49e2-ab02-bb33b379258c · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.780931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.780931Z digest=sha256:25c0b5457f5546fc3410c41ae90551e1296ed733f85f8db92c90bcdf83b10bdf

Observation 403f895a-a8e2-4982-b704-cf18734199f1 · outbound

This paper cites Visual instruction tuning.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Visual instruction tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.784750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.784750Z digest=sha256:20769ec9be28982cafa9ec8bc29d3028b85acd7fb2340adc16f4c5d190261f4d

Observation c8afdfd2-f1ab-45df-81a9-258884492a40 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.788312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.788312Z digest=sha256:e2d051061b4dd0f9ce5e338070a6b8fc40ddcad00ac048f3741623de1d3e18d6

Observation 99251e67-9df0-41a4-970d-90e9a6a68bf0 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.792560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.792560Z digest=sha256:51b8c76154ae584ffd62b3a7f6c76c70836cff677ad269d95fb64a5a15096fa0

Observation 718f2c8d-4842-4d67-9520-ebb223d1e3b8 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Swin transformer: Hierarchical vision transformer using shifted windows

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.796210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.796210Z digest=sha256:77f5ea4f602d16f87102e5be1b453024f30f313caae38e32582c9fcd7233a452

Observation 80466512-1ad5-4253-a94a-96d2ca4629bb · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.800215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.800215Z digest=sha256:36587f3275d8ec8ccfda3028a708b4d0ce859bc6df76032c5bc4021b4ba66804

Observation b5245f0b-10fa-4f44-b04f-5b197bbd3d48 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.394658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.804301Z digest=sha256:f673a96970d9925be1201f2e1e57caf77c9ea3bfcb7e34a1e21b90af21f91350

Observation 5c6f3f2f-88ce-4b28-aa9a-a977052721d7 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.807983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.807983Z digest=sha256:21f77091bd40e2d7fdda18538fc578e56294532ae365d7e72a2e54a36bd85dbd

Observation 4dfc0037-2743-484f-bd91-0de059dd5dd2 · outbound

This paper cites Spatial-temporal graph diffusion policy with kinematic mod- eling for bimanual robotic manipulation.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Spatial-temporal graph diffusion policy with kinematic mod- eling for bimanual robotic manipulation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.383566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.812074Z digest=sha256:4e6a99b28614038025ade0cb7aa361cd63a618a04070d0b0e122fd08f124c4c6

Observation b58f3cd1-f15f-4198-b84a-0ab867f706a4 · outbound

This paper cites Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.372698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.816038Z digest=sha256:91491085556137ceaea9b7536d1bc0f3a34ee04e145bae5d768fd27e4bc01af6

Observation fcf03e1a-d56a-426f-878e-f1710aed48e2 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Docvqa: A dataset for vqa on document images

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.361057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.819840Z digest=sha256:97a2899ae702b41087052f35e1b53c3ad492a4c95a1801e0e17111cecefd8eaf

Observation 2223dd60-7822-4ab6-8de0-6f1d3d81404e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Learning transferable visual models from natural language supervi- sion

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.823527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.823527Z digest=sha256:bfcfec7b6ee156943246de67e64902c27a51e516de549dd5f97cf4ae59a255b5

Observation 68d28a54-be9a-4828-b622-3bef84431d87 · outbound

This paper cites Multi-adversarial discriminative deep domain generalization for face presentation attack detection.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Multi-adversarial discriminative deep domain generalization for face presentation attack detection

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.827229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.827229Z digest=sha256:263159d4a8656bdd106314a9cb530b5003d43b57bc98f216654c9c448942d8d9

Observation 5682eb4d-25ed-4d0c-909a-1a2eb617661c · outbound

This paper cites Detecting and grounding multi-modal media manipulation.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Detecting and grounding multi-modal media manipulation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.337795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.831549Z digest=sha256:5486ecc64f488ca39407bffea3972fbd8bf402ac5b75e993d281680eff5773d4

Observation 1ee6b5e4-3a72-463e-91df-12ac496afa7a · outbound

This paper cites Detecting and grounding multi-modal media manip- ulation and beyond.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Detecting and grounding multi-modal media manip- ulation and beyond

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.327055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.835497Z digest=sha256:5a7a5b0d4853a3c9472b3c937e7c07a8083b0cf2b9db637c6510bff29f6e7d38

Observation 302dd13c-0a73-4068-a2b9-490798c2f71d · outbound

This paper cites Mome: Mixture of multimodal experts for gen- eralist multimodal large language models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mome: Mixture of multimodal experts for gen- eralist multimodal large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.315809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.839306Z digest=sha256:7bb2c79db698848bbc2d2dc5a825a51b169323289af7521668343956d2b2db4a

Observation 5de446c3-1946-4748-85a2-4b8b58af2f6c · outbound

This paper cites When do we not need larger vision models? In 12 European Conference on Computer Vision, pages 444–462,.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers When do we not need larger vision models? In 12 European Conference on Computer Vision, pages 444–462,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.304740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.843134Z digest=sha256:049ab905fc5a7c98d1efc5e35381917ed8ed097ec716cc22e4403aee82adeb8a

Observation f9db3667-dfaa-4807-a952-860ebba09887 · outbound

This paper cites Towards vqa models that can read.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Towards vqa models that can read

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.847335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.847335Z digest=sha256:2a73ab9ab118f90db64a7634b66a1e317fec171fb4414615b19ab6820a3a2065

Observation 36153ffb-5247-473e-a02a-992f25713644 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Gemini: A Family of Highly Capable Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.851125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.851125Z digest=sha256:925848ed0c85cdd5fe45c30c42df139733faef21038b06fb8d74f881f5382cb5

Observation 279cc74f-b201-40d4-95ba-30f69f6556f0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers LLaMA: Open and Efficient Foundation Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.855123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.855123Z digest=sha256:f389b21ea36716889d1b616d01e3ac9c959bac80e372bce247f810439869bf17

Observation 41e6be49-fde2-4170-950c-dea735c419e7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.858993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.858993Z digest=sha256:3515a91f4b02ee48a94db9a03afba84da2b3f69e3be72cccc250d8d43b5fd098

Observation 5b9a5693-bae5-4f98-b923-48684cfba1b9 · outbound

This paper cites V*: Guided visual search as a core mechanism in multimodal llms.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers V*: Guided visual search as a core mechanism in multimodal llms

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.287477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.863115Z digest=sha256:aa9f7cb5b7b4b533c1d80c3620d564e924d50af97f79ddb8bebc661daac98d01

Observation 8caefd48-b66b-45fd-98dc-526a0336d167 · outbound

This paper cites Gui-explorer: Au- tonomous exploration and mining of transition-aware knowl- edge for gui agent.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Gui-explorer: Au- tonomous exploration and mining of transition-aware knowl- edge for gui agent

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.276332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.867100Z digest=sha256:ae2469358a616ac6f4018c110ecffa0216f1c1f32690d76a3788fd4ec93b3011

Observation 369d4a1a-0c92-4c72-b617-b8716adcf56d · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.870758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.870758Z digest=sha256:b76bcc0a55ca9274e41647aae10ce162ab2d2124d55e315c536723856260f9bf

Observation 9bd10117-1474-4938-a48d-c544a1805316 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.875541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.875541Z digest=sha256:01303f643f688ca16dad76ecef3013ed1449626572105b8af88a7fe3a0ec7409

Observation 2e866717-1d12-4ccf-bf1b-c775db334c8d · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.879391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.879391Z digest=sha256:f23f7a7b336ea97262211e3de8b3fcfbe14963f877e54fb7f05ff8466c1266de

Observation 57e8c7ba-2775-43f7-8808-f8bb8be2a97e · outbound

This paper cites Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.264231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.883290Z digest=sha256:b7eaaed29e98e4dbcad301dd126f08a5fc926d689822de97789222f9e5c4acdc

Observation 3908d8fc-4886-4a77-916b-13a650794a9b · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.886964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.886964Z digest=sha256:1283f13f3742fb6d1a707af754ceea7a7daed30a80bbafc567a265efcc1796f4

Observation d570105d-11f2-4813-bad3-6a832874fa46 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.239916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.891094Z digest=sha256:1abca11abc2c07a3f09a425ab7f63adbb33e8f3af86568e6c2154c1ef2b9ade8

Observation 95e99493-16af-43f9-aacd-7f06afa2fa0d · outbound

This paper cites Modeling context in referring expres- sions.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Modeling context in referring expres- sions

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.228709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.895145Z digest=sha256:a2616e6827a193dc331ccc2d0112ce43374ddd78b7ec2712a8fbd363a5e48e45

Observation 288adfb4-a5bb-43f1-96d3-a928edd16fbe · outbound

This paper cites Sigmoid loss for language image pre-training.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Sigmoid loss for language image pre-training

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.216519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.899089Z digest=sha256:a111b34827b2d62188b4b90730f7ca4b100ecec2a61ddd08d532fd4717163eda

Observation 7f24491c-e4a3-4866-b6fe-c2441f230af3 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.903154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.903154Z digest=sha256:baf69dc3471e063d5a7ad991e82f90a8c83e60003a7d067ed31485fb2b8452f9

Observation 8fbc5767-86a4-45e6-85df-5737da82d995 · outbound

This paper cites Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.907427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.907427Z digest=sha256:e9b9a7629ac0c3a01b22ccf137a72f0c2dd2a0c6ab79ecbc8f920a7cc65cf63c

Observation ac5d1f9d-a70b-4c35-8dde-30a233ecbc2c · outbound

This paper cites From redundancy to relevance: Enhancing explainability in multimodal large language mod- els.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers From redundancy to relevance: Enhancing explainability in multimodal large language mod- els

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.199802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.911600Z digest=sha256:f35be5311bab40abaa526ccb543cfdf6f74c3f8858da793aaafe098b6a67db8c

Observation 5a52d847-6c2a-401f-a6d0-37d6db3f6f4f · outbound

This paper cites an unresolved cited work.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:38:19.186362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T13:38:18.915446Z digest=sha256:478cc341f8c3f7ee1006e9ebee3549367532cf22b25aeb3775b6c5bdf1afaa0a

Observation b382da80-0c94-48de-b22d-f2e8fd4ba5eb · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.919283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.919283Z digest=sha256:da887b798e845e440b8be3ba826469b5bf4a4b4bea7f2e0d83722ca2f331a49f

Pith citing papers

Observation 1e720cc7-46b5-45e6-8500-60bd9775742e · inbound

UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries cites this paper.

UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:15:33.790283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:14:28.540629Z digest=sha256:f267d10a3534eb8a37be710b4b38da8c0af4ae3ce89fb5da09c5fd00ab76480c

Observation 308599d9-d0d2-4720-a799-43695fdad954 · inbound

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models cites this paper.

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:01.667217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T23:09:32.594194Z digest=sha256:6bf43c5e4b2ba70c11f7526ca4310c6478fa79143cfaf26489602e8c082772b0