Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T13:38:18.919283Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 2 inbound Pith citation observations for arXiv:2501.16297.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T13:38:18.919283Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T23:09:32.594194Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T23:14:01.665133Z
66 of 66 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9cecb0cd-669b-499c-b5d1-611b54e4602a · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2930fe71-38c5-4dd1-96c0-cf4eb91f3fa8 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Lion: Empowering multimodal large language model with dual-level visual knowledge
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 00b8b4a7-9f6e-4a5b-9383-1862af359e04 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Less is more: Empowering gui agent with context- aware simplification
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 06f4e848-242b-4446-baae-33a9dde9139b · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48e9b2e3-80b4-4576-bacf-75157864c99b · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Spa-bench: A comprehensive benchmark for smartphone agent evalua- tion
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4a044303-7ddb-4633-a91d-549cd3158b58 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Sharegpt4v: Improving large multi-modal models with better captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 189dacbe-1e7a-47f2-b301-7d58e427a451 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f4d8f2bb-447e-48e7-b1a4-6d90bfe78e25 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86a2d164-30f3-4be8-878c-2ed042bca62e · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers InstructBLIP: Towards general-purpose vision-language models with instruction tuning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 54b7c54b-407f-42a3-a4d7-cae40c478b81 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Vision transformers need registers
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b05491e6-12f0-4fae-b022-4e138d0b0fde · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers An image is worth 16x16 words: Transformers for image recognition at scale
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86885e08-9b0e-4839-b105-0bd5e9780b93 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0b27731-56bc-4aa2-b96c-c85fdb6faf2f · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Gaussian Error Linear Units (GELUs)
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac01ab06-dbdb-4897-a3e3-12b933609d0d · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers CogVLM2: Visual Language Models for Image and Video Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33395e58-9e1a-40ff-842f-eee1d6082bc2 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mplug-docowl 1.5: Unified structure learning for ocr-free document understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 905014d2-2ac6-4584-bced-f09eaa7de993 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mini-monkey: Alleviating the semantic saw- tooth effect for lightweight MLLMs via complementary im- age pyramid
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 156e88a4-8f94-4638-83c2-989d28bd2371 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Hires-llava: Restoring fragmen- tation input in high-resolution large vision-language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6493d439-31b1-4e22-815a-5bd29d58c30c · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Openvla: An open-source vision-language-action model
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 057861c3-7ed6-4acb-943c-0fc7d84e2852 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01559c1a-9d54-4159-b24d-a22c4e2cb7e2 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers LLaV A-onevision: Easy visual task transfer
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbaafdc9-d0bb-4a5c-8a55-5eb8ee60a24b · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers STAR: Learning diverse robot skill abstractions through rotation-augmented vector quan- tization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9773fdbd-b6ca-4420-b586-1ed3a5cdc433 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Blip-2: Bootstrapping language-image pre-training with 11 frozen image encoders and large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 05ab51b1-b59b-4d17-88f1-e36a767cf6f2 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Flex- attention for efficient high-resolution vision-language mod- els
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3898cb60-9eb4-4ae8-ba70-2caf4890a0de · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Lion-fs: Fast & slow video-language thinker as online video assistant
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 92305da7-4122-4711-a3e1-2bb26cb96def · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Evaluating object hallucination in large vision-language models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1862721-228f-481c-a619-8369a8c60462 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa7da1ae-4093-4815-88dc-9d1598f5cd82 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Optimus-1: Hybrid multimodal memory empowered agents excel in long-horizon tasks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0aa62c3a-5c79-47be-b161-5eb2c705ed67 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mon- key: Image resolution and text label are important things for large multi-modal models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 75e8e79f-8633-4bed-bfe3-4ff0824dd92f · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Optimus-2: Multimodal minecraft agent with goal-observation-action conditioned policy
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aeb7aca1-8e46-4cf1-bbf9-a55e1f82e333 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Improved baselines with visual instruction tuning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 94754f3a-0b0a-49e2-ab02-bb33b379258c · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 403f895a-a8e2-4982-b704-cf18734199f1 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Visual instruction tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8afdfd2-f1ab-45df-81a9-258884492a40 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99251e67-9df0-41a4-970d-90e9a6a68bf0 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 718f2c8d-4842-4d67-9520-ebb223d1e3b8 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Swin transformer: Hierarchical vision transformer using shifted windows
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80466512-1ad5-4253-a94a-96d2ca4629bb · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5245f0b-10fa-4f44-b04f-5b197bbd3d48 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5c6f3f2f-88ce-4b28-aa9a-a977052721d7 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dfc0037-2743-484f-bd91-0de059dd5dd2 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Spatial-temporal graph diffusion policy with kinematic mod- eling for bimanual robotic manipulation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b58f3cd1-f15f-4198-b84a-0ab867f706a4 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fcf03e1a-d56a-426f-878e-f1710aed48e2 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Docvqa: A dataset for vqa on document images
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2223dd60-7822-4ab6-8de0-6f1d3d81404e · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Learning transferable visual models from natural language supervi- sion
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68d28a54-be9a-4828-b622-3bef84431d87 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Multi-adversarial discriminative deep domain generalization for face presentation attack detection
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5682eb4d-25ed-4d0c-909a-1a2eb617661c · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Detecting and grounding multi-modal media manipulation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1ee6b5e4-3a72-463e-91df-12ac496afa7a · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Detecting and grounding multi-modal media manip- ulation and beyond
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 302dd13c-0a73-4068-a2b9-490798c2f71d · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mome: Mixture of multimodal experts for gen- eralist multimodal large language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5de446c3-1946-4748-85a2-4b8b58af2f6c · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers When do we not need larger vision models? In 12 European Conference on Computer Vision, pages 444–462,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f9db3667-dfaa-4807-a952-860ebba09887 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Towards vqa models that can read
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36153ffb-5247-473e-a02a-992f25713644 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Gemini: A Family of Highly Capable Multimodal Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 279cc74f-b201-40d4-95ba-30f69f6556f0 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers LLaMA: Open and Efficient Foundation Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41e6be49-fde2-4170-950c-dea735c419e7 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b9a5693-bae5-4f98-b923-48684cfba1b9 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers V*: Guided visual search as a core mechanism in multimodal llms
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8caefd48-b66b-45fd-98dc-526a0336d167 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Gui-explorer: Au- tonomous exploration and mining of transition-aware knowl- edge for gui agent
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 369d4a1a-0c92-4c72-b617-b8716adcf56d · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bd10117-1474-4938-a48d-c544a1805316 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e866717-1d12-4ccf-bf1b-c775db334c8d · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57e8c7ba-2775-43f7-8808-f8bb8be2a97e · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3908d8fc-4886-4a77-916b-13a650794a9b · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d570105d-11f2-4813-bad3-6a832874fa46 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 95e99493-16af-43f9-aacd-7f06afa2fa0d · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Modeling context in referring expres- sions
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 288adfb4-a5bb-43f1-96d3-a928edd16fbe · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Sigmoid loss for language image pre-training
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7f24491c-e4a3-4866-b6fe-c2441f230af3 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fbc5767-86a4-45e6-85df-5737da82d995 · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac5d1f9d-a70b-4c35-8dde-30a233ecbc2c · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers From redundancy to relevance: Enhancing explainability in multimodal large language mod- els
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5a52d847-6c2a-401f-a6d0-37d6db3f6f4f · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b382da80-0c94-48de-b22d-f2e8fd4ba5eb · outbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e720cc7-46b5-45e6-8500-60bd9775742e · inbound
UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 308599d9-d0d2-4720-a799-43695fdad954 · inbound
Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.