Pith. sign in

Paper Citation Record · LEDGER

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2412.09616.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09616 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T15:50:49.083652Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T09:22:16.152985Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7567a43a-81e0-4548-bf55-f8e0092efe2b · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.309183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:661298e39490a1ca052104c704089f4f47750af762d0fe23870e11ce306bc1a3

Observation 88919fe1-327b-4224-af79-8ca2fb029b26 · inbound

LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops cites this paper.

LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:22:16.157643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T09:18:26.804728Z digest=sha256:61b025b919b46e39442ba10d9cc8cc8aad50e9ac0fc1aaa706edd59a1858fb61

Observation d8b617ae-c426-43db-9522-598b88ade219 · inbound

Mitigating Coordinate Prediction Bias from Positional Encoding Failures cites this paper.

Mitigating Coordinate Prediction Bias from Positional Encoding Failures V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:12:23.580165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:11:30.734633Z digest=sha256:9763ccd9443de43c6ac40d55aafcb64f7b4ff28ef016b2fa4feb080691ee8a4b

Observation 19e45ffc-f358-4f8f-8222-cda983d2581b · inbound

Internalized Reasoning for Long-Context Visual Document Understanding cites this paper.

Internalized Reasoning for Long-Context Visual Document Understanding V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:53:27.762573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:53:19.148407Z digest=sha256:697d5c501a2ab481087f0973901a79b038063e97d211f9d7cf25a31967ae4c0b

Observation 17745e31-98ca-40a7-a894-555f4669ce45 · inbound

Internalized Reasoning for Long-Context Visual Document Understanding cites this paper.

Internalized Reasoning for Long-Context Visual Document Understanding V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T15:50:49.083652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:50:49.083652Z digest=sha256:fa427330834ae3cb316957a34a3084ca4c1dac7198da4af33dba721a9744a610