Pith. sign in

Paper Citation Record · LEDGER

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding

As of 14 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2501.10967.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10967 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:53:21.337583Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:33:53.715127Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T13:33:59.190592Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bd728695-fe16-4c4e-93bc-57d78dbf47a0 · outbound

This paper cites URL: " 'urlintro :=.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.113126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.113126Z digest=sha256:2fcfe52412c3d7d7448cd0ed53946a25208db3f51079c0c28f18a00406891955

Observation 1e1ed59a-42e7-4036-a0a8-c8cae395c280 · outbound

This paper cites write newline.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.117733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.117733Z digest=sha256:4231797cb19949ad4f2f57fe902403cf32747d378913828fe4a3ad7dcd917171

Observation 5786a460-1c18-437c-ae02-9830bba80b54 · outbound

This paper cites GPT-4 Technical Report.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.122055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.122055Z digest=sha256:7d03ac5573b2d8345576405ad71830343e7b7f7f49b9fc988fff8301cc02055b

Observation 3f1e3c33-a0b7-4528-93c8-83bc459fe542 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.126218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.126218Z digest=sha256:2f6cb1a20a385d4b3b9c66266b1e341c9c31f461f2fbff337846179f6200378b

Observation 65ca5de0-af9b-4ccd-bfa6-d8ecacc75b25 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.130751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.130751Z digest=sha256:9202fc02bdf94c3ef823a006e9515566db334e715e66d6a2fc964cb348e85013

Observation 51e49b39-3cea-45e4-9068-3b340d657146 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.134552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.134552Z digest=sha256:5644a57024373fff6bd4ad568936db4bc4272d9d3e3c457759159b409b17ba34

Observation 51464a4a-ffcf-4210-9e2f-cfb33c83271f · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.138626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.138626Z digest=sha256:c8f1380557d18c2813b42ea46302283ebfd6743b0353fff11811765b55f4342d

Observation 1b14393b-8007-46ad-8a5f-cb8277c9670f · outbound

This paper cites LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.142546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.142546Z digest=sha256:daabcc8011bf43caff90471c35bb55939fb8e4046b8f1576e57c15ab4fc47fba

Observation b3880363-f873-48c8-98a8-aff26caa6c42 · outbound

This paper cites MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.146912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.146912Z digest=sha256:85eb760392fdb6767e72c49fcd6d487147b8926a6466c59cdbc1e1b6b3952d6e

Observation daf33d11-56da-4485-8da8-ded14d604195 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.151267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.151267Z digest=sha256:13ba3d870a995738f39697bb6df2e3ecfb6144d48f483f3a57fe5ebcd9adc1e5

Observation ad34788c-ce41-4ede-95a2-2043812b615f · outbound

This paper cites VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.155051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.155051Z digest=sha256:3f98ef24f69fa012fc758f69de135fd3b279bfe7dc95da6926d6f6dc993eb466

Observation 0a3d22f0-2778-4a2d-b7fd-3ff154b867d7 · outbound

This paper cites Conditional Positional Encodings for Vision Transformers.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Conditional Positional Encodings for Vision Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.160206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.160206Z digest=sha256:2664cffa74221ac89fb60a810a4220deb9fea73e52eea0bef4b47a27c2ebe129

Observation 62e118f4-ef7a-4979-9a29-dcbec9eedb22 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.163914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.163914Z digest=sha256:f91f186c705c1984004750bc4d59b7463d6d38a9b5045b73c49a062b81c593ac

Observation fe7a0637-b79c-455d-8e71-86b8aa567b47 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.168110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.168110Z digest=sha256:e86121f8ab94bb34607dcfe90df6b863269b0ca80756c12823ac84ce653c745b

Observation 6e2de7f6-551f-4427-b9d0-d7f619aad12f · outbound

This paper cites DeBERTa: Decoding-enhanced BERT with Disentangled Attention.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding DeBERTa: Decoding-enhanced BERT with Disentangled Attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.171635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.171635Z digest=sha256:a86a74d154e1c07bd5fccded0aac08e92c0450ddf3548c79f435ab487f188809

Observation 4bb4cd20-1813-4b85-9ac8-ff9322f3d676 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.175649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.175649Z digest=sha256:6e357b2ad619f5f61838d8251302f6226e2a60cf1e78e0fd03abdf3dec9e5e78

Observation 9533b3e9-3db5-4cba-92e3-c9ed8fe00f30 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.179430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.179430Z digest=sha256:5e6cd21c6a529ddd18697d5e51058ac1eec6272e32c2b844962564bb751e3344

Observation 72d86c38-aa0b-43ab-840e-59b631c2b669 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.183417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.183417Z digest=sha256:c6200b0474a162acfb0f783a1101cc9b597a7715e4aa2b38e5b94c2e23647b85

Observation ba211248-19a1-49d9-b6d6-02649e595073 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:53:21.898479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T18:53:21.187231Z digest=sha256:7f5e34dea6c89418ba3b0473b4a98391255f995919d7aa09bac68876a8fd920d

Observation 0f1687e1-d1c9-4d6f-bf6e-eafa10da89d8 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.191079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.191079Z digest=sha256:3023e4047420e392b1b1b777ef2284a162f625069dc20207b449b744cd69a9cb

Observation c9a6236e-7d34-4ff6-944a-789e22ec86c2 · outbound

This paper cites How Much Position Information Do Convolutional Neural Networks Encode?.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding How Much Position Information Do Convolutional Neural Networks Encode?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.194797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.194797Z digest=sha256:9e6d00ff288a1965a12f55429522d72b4daa8bf42d18d6baed566127af2ba98c

Observation 2e230440-2b43-4c7f-921f-1b2945662584 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:53:21.878556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T18:53:21.198710Z digest=sha256:dd01bbb8c9c89d50c3ddafa9dff2549278ede179bde0159f38f367febedd919e

Observation 98ca37df-6dc3-4f06-81a4-12e2a5c9b12a · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.202262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.202262Z digest=sha256:dff6cd27e9759ca21390fbb4b329cfb599c625ad68feb1099734de9f26b5e422

Observation c79ad036-6786-49a8-964c-c40a3f80f969 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.205868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.205868Z digest=sha256:c2cff7619b7b5b8a0ba769a16b48c870afbfd8ef316e24364bc303f9993589c8

Observation eb81aa44-c08f-4884-9f02-50748bdcd999 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.209443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.209443Z digest=sha256:caf5ab3a02297af266c52da4820b4a0415b7503899934e8e5678220ede2e0010

Observation 67edd8ee-aac5-4079-86b4-86f418161e38 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Evaluating Object Hallucination in Large Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.213156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.213156Z digest=sha256:38644e77a7afac80872c51219a9d1a94b32b88fd6aaa26691e28f7d6afa25069

Observation 4b37fb5a-0c7f-4e07-809f-f8b5c6e41022 · outbound

This paper cites Textbooks Are All You Need II: phi-1.5 technical report.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Textbooks Are All You Need II: phi-1.5 technical report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.217368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.217368Z digest=sha256:5d1f0f5439b5211373373d1d3e52ce57e9f0597bb5c41e21828b66f677fce773

Observation 653918df-6e49-4619-86e0-f6f66b08ef8e · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.221344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.221344Z digest=sha256:da55735a733cde9a103c1a21fa130afd8b2a7d1beda8f4c313699cb7243fd2a5

Observation b86dd734-b85d-4b76-9323-3f3313008965 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.225192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.225192Z digest=sha256:7db92616e42018e5ade03ef28000bfae1a52b6938db1cb37399092ea6d3ce392

Observation 921330f4-c344-4f16-9472-e13babacc8df · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.228841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.228841Z digest=sha256:1dab5ebfcdf49c27b87331710d5dc8501253b2d74133c2f4accea46e3893712d

Observation 61e2ef02-53ea-4aab-a0eb-e557248a08ac · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:53:21.830621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T18:53:21.232248Z digest=sha256:da2f3ded7d8251e6848bd4808dfba17f550c52a3e2af58463aa5187250a5eae0

Observation c6983f93-e11b-42cd-8369-71c4f0683dd0 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.235324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.235324Z digest=sha256:deabf09521067c6085c4317ba80ebaa6215b8835eaca15791aa30efeecf3897c

Observation 1deeb4d5-9968-4346-96a7-3a7dfebe1448 · outbound

This paper cites FiT: Flexible Vision Transformer for Diffusion Model.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding FiT: Flexible Vision Transformer for Diffusion Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.238911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.238911Z digest=sha256:0efbbc7f2405bcc8ead0e6d0a897bd68a83d4b480dfa7d04d1a71c981b6b10db

Observation 4d8a54e7-659d-4750-bce3-9efeac8d4b21 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.242563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.242563Z digest=sha256:17054c575282f5a423ea1edfca4191d18cae90af8adf5baad693f2323ab26742

Observation 4f5c5029-2e62-4966-8d94-02d2658ead84 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.245981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.245981Z digest=sha256:62282b469174dc8cefd0d5a63e1fbdd7b6655e83e3147da19f577192ec808ea1

Observation 150582c0-c9c2-4ed6-be1e-53a5d8668d19 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:53:21.798081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T18:53:21.249253Z digest=sha256:fcbfe1d4cf64f90d2971219ede3dd11d7f043e3b8083b42f91000566109ab711

Observation bffab2de-dbd3-446a-8bb9-b4ad4b8a78bb · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.252516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.252516Z digest=sha256:d897bdd2a303f4f00f760d483d6d7e627ebb3807f1d2fcbb08e083a4880e8f17

Observation b2d8c1cc-8a4b-46b3-a8c8-3cfb23593de9 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.255966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.255966Z digest=sha256:c3e581062a0ae562b65400fa643ad9c7d655d8dd4056f7627bcf543efbc84249

Observation 2958f6a8-b860-469a-aa49-36b3470ff17b · outbound

This paper cites Self-Attention with Relative Position Representations.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Self-Attention with Relative Position Representations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.259932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.259932Z digest=sha256:0029441c89efd3434e334b509e35acd3ccfae734146c923a3fe30ea58fd02c6a

Observation 03336ce4-4f74-4aba-9e35-a8ac106aa1e2 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.263723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.263723Z digest=sha256:c0b68b829429f579cabb375b176216e656d953cad81295226e5f86a4673503ac

Observation 9148d516-cbaa-498f-aadd-62d0af40512c · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.267210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.267210Z digest=sha256:2c455ace68d244c6d40c6ee7f7309c85c8b59e7d4056acea8ad955ce717125c5

Observation 97612d31-27b7-4869-82a5-10360bfc0a6e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.270877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.270877Z digest=sha256:cb2836715eb653ebccaf81e9fbeff5e6aa93084ef9eb7227410cf9cda175cfe6

Observation 976fda51-b607-4cab-bcc2-1bd10567e1a2 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.274697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.274697Z digest=sha256:c0ea51b169428b158728d6c9a5c0fff0852d247c4a07037f50ec506988614dae

Observation 8ead18f3-da29-4786-ad75-6e6de0d29d8c · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.278186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.278186Z digest=sha256:1cf1ce4b3ae2a86434e6efdf3eac789f14ca253f6e86def51ea03cb52b019e6a

Observation 2a87b37c-8e59-498d-b94a-983660750c0a · outbound

This paper cites Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.282078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.282078Z digest=sha256:0bce6f66ada3954c02dfe476a58e258c6e0ae66be341a4660929e36588e8e25c

Observation d89f973f-e032-4b6e-ace7-5b0087e73553 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.286291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.286291Z digest=sha256:32fae66a274aaec3109d1f18e4882d028bab75fdcd988a4eebbd060bebae0081

Observation 6c7921cb-2645-4a42-9211-5a3e3634d3bf · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:53:21.746478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T18:53:21.290614Z digest=sha256:98441cc44a3cff98015a0dc352417befb12db3a7a6ccdbef7aea477a8e1a20b7

Observation 788b66ae-eae0-43de-85b5-8aa41fbca580 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:53:21.732702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T18:53:21.294788Z digest=sha256:4e02a671afb559779b9f67dcdbc391bacdd15ecc26cd71f5562042846ceb2d26

Observation 1f5ddc7b-37f5-4309-86b9-676567c03777 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 49

Resolution
verified exact
doi, observed 2026-08-10T18:53:21.376297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T18:53:21.299539Z digest=sha256:88677ac5906b45965fcb20009082d54c9c258960e45b8c7b9bab5bf1c0530787

Observation f50923cd-52be-477a-a241-05610718db17 · outbound

This paper cites Mitigating Object Hallucination via Concentric Causal Attention.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Mitigating Object Hallucination via Concentric Causal Attention

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.304184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.304184Z digest=sha256:48f0758a314142a02b60b5f04571d0b3ff826ad5045b963f1bae3f9de005a0dc

Observation 3822327e-6fb4-4b4f-96c6-9600459ea71a · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.309192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.309192Z digest=sha256:d0227ece90ea4f58c6ae02958715c16a022d120ae1cfcae7c69a4755bf5df007

Observation 90c41d2a-8e4a-4132-bf31-981baf459a15 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.313329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.313329Z digest=sha256:51f6cdee2d8189891441ded4e4f859d4479e7c63629e7b5c950c36304da501e9

Observation d719f9fd-9191-4d10-9ea0-df0d727b03ec · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.317589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.317589Z digest=sha256:9d9679e88135d00c3b1d146f966a50e11061d68754dae22f00b7c8023b365537

Observation e3a9e902-f5cd-4546-8b40-8a0e844289c3 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.321944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.321944Z digest=sha256:da1f78177cc64be718782f52a19206c7f8a908238d244ab660d005cc9b1a5a76

Observation a51fdb93-0565-4608-a31a-323286fc6904 · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.325758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.325758Z digest=sha256:dd8e800d125b3689d889449c0128ed71cadf4aac7f22c92c8e3d2d1af441fff0

Observation 86ea2dec-3ff0-4369-8491-984c2e3211bb · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.329427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.329427Z digest=sha256:9a9ab15718c65dd2e3acb0eccea74db3f3df05c9eac207063a8950220a3a7bc5

Observation 8d05c34b-f13f-4580-8ad8-275f0d5552ff · outbound

This paper cites an unresolved cited work.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.333128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.333128Z digest=sha256:59d58cf849023f0bd80177a96c2525484dc47e66d963dbec425eb3cb7e3456d4

Observation 892db1c5-979e-4703-bb70-41129f00a538 · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.337583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.337583Z digest=sha256:69b902feb58489927bad30b6ba622afd2542ca73ee21b1a866968f291ceffac2

Pith citing papers

Observation 22035ef3-ee65-4c95-a16f-ec251a2eaf41 · inbound

ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models cites this paper.

ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:59.270505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:33:53.715127Z digest=sha256:d953fc1d3eb38a13e363939f43d432f3138897925fe29ef16fa8d1fb7a264bac