Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:53:21.337583Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2501.10967.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:53:21.337583Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:33:53.715127Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T13:33:59.190592Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bd728695-fe16-4c4e-93bc-57d78dbf47a0 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding URL: " 'urlintro :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e1ed59a-42e7-4036-a0a8-c8cae395c280 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5786a460-1c18-437c-ae02-9830bba80b54 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding GPT-4 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f1e3c33-a0b7-4528-93c8-83bc459fe542 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65ca5de0-af9b-4ccd-bfa6-d8ecacc75b25 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51e49b39-3cea-45e4-9068-3b340d657146 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51464a4a-ffcf-4210-9e2f-cfb33c83271f · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b14393b-8007-46ad-8a5f-cb8277c9670f · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3880363-f873-48c8-98a8-aff26caa6c42 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daf33d11-56da-4485-8da8-ded14d604195 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad34788c-ce41-4ede-95a2-2043812b615f · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a3d22f0-2778-4a2d-b7fd-3ff154b867d7 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Conditional Positional Encodings for Vision Transformers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62e118f4-ef7a-4979-9a29-dcbec9eedb22 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe7a0637-b79c-455d-8e71-86b8aa567b47 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e2de7f6-551f-4427-b9d0-d7f619aad12f · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding DeBERTa: Decoding-enhanced BERT with Disentangled Attention
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bb4cd20-1813-4b85-9ac8-ff9322f3d676 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9533b3e9-3db5-4cba-92e3-c9ed8fe00f30 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72d86c38-aa0b-43ab-840e-59b631c2b669 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba211248-19a1-49d9-b6d6-02649e595073 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0f1687e1-d1c9-4d6f-bf6e-eafa10da89d8 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9a6236e-7d34-4ff6-944a-789e22ec86c2 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding How Much Position Information Do Convolutional Neural Networks Encode?
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e230440-2b43-4c7f-921f-1b2945662584 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 98ca37df-6dc3-4f06-81a4-12e2a5c9b12a · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c79ad036-6786-49a8-964c-c40a3f80f969 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb81aa44-c08f-4884-9f02-50748bdcd999 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67edd8ee-aac5-4079-86b4-86f418161e38 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Evaluating Object Hallucination in Large Vision-Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b37fb5a-0c7f-4e07-809f-f8b5c6e41022 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Textbooks Are All You Need II: phi-1.5 technical report
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 653918df-6e49-4619-86e0-f6f66b08ef8e · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b86dd734-b85d-4b76-9323-3f3313008965 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 921330f4-c344-4f16-9472-e13babacc8df · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61e2ef02-53ea-4aab-a0eb-e557248a08ac · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c6983f93-e11b-42cd-8369-71c4f0683dd0 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1deeb4d5-9968-4346-96a7-3a7dfebe1448 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding FiT: Flexible Vision Transformer for Diffusion Model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d8a54e7-659d-4750-bce3-9efeac8d4b21 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f5c5029-2e62-4966-8d94-02d2658ead84 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 150582c0-c9c2-4ed6-be1e-53a5d8668d19 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bffab2de-dbd3-446a-8bb9-b4ad4b8a78bb · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2d8c1cc-8a4b-46b3-a8c8-3cfb23593de9 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2958f6a8-b860-469a-aa49-36b3470ff17b · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Self-Attention with Relative Position Representations
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03336ce4-4f74-4aba-9e35-a8ac106aa1e2 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9148d516-cbaa-498f-aadd-62d0af40512c · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97612d31-27b7-4869-82a5-10360bfc0a6e · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 976fda51-b607-4cab-bcc2-1bd10567e1a2 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ead18f3-da29-4786-ad75-6e6de0d29d8c · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a87b37c-8e59-498d-b94a-983660750c0a · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d89f973f-e032-4b6e-ace7-5b0087e73553 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c7921cb-2645-4a42-9211-5a3e3634d3bf · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 788b66ae-eae0-43de-85b5-8aa41fbca580 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1f5ddc7b-37f5-4309-86b9-676567c03777 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f50923cd-52be-477a-a241-05610718db17 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Mitigating Object Hallucination via Concentric Causal Attention
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3822327e-6fb4-4b4f-96c6-9600459ea71a · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90c41d2a-8e4a-4132-bf31-981baf459a15 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d719f9fd-9191-4d10-9ea0-df0d727b03ec · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3a9e902-f5cd-4546-8b40-8a0e844289c3 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a51fdb93-0565-4608-a31a-323286fc6904 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86ea2dec-3ff0-4369-8491-984c2e3211bb · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d05c34b-f13f-4580-8ad8-275f0d5552ff · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 892db1c5-979e-4703-bb70-41129f00a538 · outbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding TinyLLaVA: A Framework of Small-scale Large Multimodal Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22035ef3-ee65-4c95-a16f-ec251a2eaf41 · inbound
ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.