Pith. sign in

Paper Citation Record · LEDGER

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models

As of 11 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2412.12735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12735 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:52:57.225192Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2dcfc16d-07e7-4520-b819-f58cb25181c5 · outbound

This paper cites URL: " 'urlintro :=.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.008460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.008460Z digest=sha256:e179284459b5cd16f2b249433cc1e8c8f9d814a7d144434a65facc6501bad10d

Observation ccf1e047-37db-4ede-a275-7a60d030dab3 · outbound

This paper cites write newline.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.014345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.014345Z digest=sha256:fe081f8f579a34c88b13d9209d7c7b676fc7a9b62edb80ce0cb62a61fea25082

Observation 81c81897-6316-46f1-b41c-93a5ea47eff0 · outbound

This paper cites Training-Free Long-Context Scaling of Large Language Models.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Training-Free Long-Context Scaling of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.018380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.018380Z digest=sha256:a8ade104f4db7b9c1f7c5d34d77c992195a5c7b58c9791f478e68bc4937ad5af

Observation a9dd32fa-fb61-4bec-99dd-d3059e8d4dc5 · outbound

This paper cites Why Does the Effective Context Length of LLMs Fall Short?.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Why Does the Effective Context Length of LLMs Fall Short?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.022710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.022710Z digest=sha256:3e05ef48e4f3f26d141b1128cf6fce60bc540fa2d526907d32da90c692c7f8b4

Observation 4351508a-781c-4c02-a763-0042b6bd6997 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.026749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.026749Z digest=sha256:91493473c019054491955a6d1f88b60cb3e900fdfedc9b44448d72f77d0c89f3

Observation a2a2a2c2-20de-4064-bcca-d5c16c17ae04 · outbound

This paper cites LongAlign: A Recipe for Long Context Alignment of Large Language Models.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models LongAlign: A Recipe for Long Context Alignment of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.030461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.030461Z digest=sha256:99318be4679fbffdc0da9eb25db8990e6aad63a2a20778438379d3f410d07d5a

Observation 29d34610-de14-4dab-b07d-a1ecdabf9cb4 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.034144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.034144Z digest=sha256:d492b33e49d7c631b071bb455760538b90425984c1e98a1f9e7fe2de01c819fe

Observation d547b8b2-6939-400e-8813-961cde1eec81 · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Extending Context Window of Large Language Models via Positional Interpolation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.037972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.037972Z digest=sha256:d0ca2b26eef32ea624ec122a66d20ae839b096fb852881683c61a91e9aef6207

Observation a08b6390-119e-410e-a0e1-19ed29d913f7 · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:52:58.103697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:52:57.041765Z digest=sha256:a42211a2079eaa03da00eccc331bf2b2d5dbc34d461df524ce5622fa1dc4d59e

Observation af0ddcb4-6675-4e21-a28b-365687bb4553 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.045374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.045374Z digest=sha256:a8e3ee9a79286717b7d39e5546949b3c3d2a1c4a1657a9d5486062100d5179c5

Observation b558fd6f-c168-489c-9493-836982d6a573 · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.048852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.048852Z digest=sha256:2774eb7d374d855e68c41307bfb6b6b49af615db1f53a5b438535753eae9a483

Observation 5d061051-915d-4015-b771-20bb819c5115 · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.051960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.051960Z digest=sha256:05858e4cb4dbed9c8028a7187435ecdbc811a4f2d94403768ca7e57381d2512f

Observation dd8a30ca-cf04-43e0-92ab-4b9919b549bf · outbound

This paper cites SlowFast Networks for Video Recognition.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models SlowFast Networks for Video Recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.055375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.055375Z digest=sha256:5f1ba74a98d19add12683a1af016a8b4953c2a77bdae6a738ea8fc84d6d0589a

Observation a1834679-23ab-4180-9f05-16c2e846a4e7 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.058793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.058793Z digest=sha256:132b88dc4c920f30fa20b2be14f5ed31997c74b968a1757131f130e774ed7332

Observation 89300531-3d8e-4416-9ad8-794e839ee30d · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.062343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.062343Z digest=sha256:30a8afc0a82d21ec5c866caf9167cc96ad9f2a380a94dbde62e364bbfd6dac18

Observation f743acbe-c706-4107-a2f2-b81bf92ef702 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.065731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.065731Z digest=sha256:f1b9e8ef77c24282cbd71bd2b3fa018338cd37670bdae1b0fd31ed7bdadf78ba

Observation 9c075e71-43ea-4df9-bb9f-93a9c2ea915a · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:52:58.081212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:52:57.069160Z digest=sha256:4485a396afed9c68a764536bac3716c9b0d554d4ddc7bd124cb2e188be535e3c

Observation f07333e1-9073-4a84-8116-313acf5ca227 · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.076132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.076132Z digest=sha256:a04928f18e82366ea1ccff1db04eea1acd9072a070e81eae190f1d53ec5a8248

Observation 5f492940-9c1e-45b6-b636-a26dfe13d8dd · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.079187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.079187Z digest=sha256:a159943b6da8a7a79ee8c8ae5b7cbb6725eb16065fb585e26c51e60797a421c6

Observation bb9c01c4-71d7-40b8-b144-77fa99e4f328 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen - Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen - Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.082728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.082728Z digest=sha256:5d27821181276a2ad2d10b6c02a93909374f1c981427ee6092289d936cd4e7bc

Observation 675a55bd-2e84-47bb-a331-f6457955a0e4 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.085731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.085731Z digest=sha256:0038d84979ef5e303209dd016778deec09f1ae3a7e634e3d4b3efa4f68b2de23

Observation 961d1d6e-3efb-4fc7-9d2e-12dca1d937e0 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.089252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.089252Z digest=sha256:e7d5f152eef97405942e0eb65f72c14e562ab07374e5fcd1ce84fb00216acd50

Observation 9d7ee437-7529-442d-9bd9-58f0ba233421 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.092848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.092848Z digest=sha256:4c0ad354f61ee813c7e0d781eaec8b91b8ada1444f5f5a382e27f8d8ef79bf63

Observation 2da2df34-2ee6-4cf2-b86e-b5042ec76380 · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:52:58.063682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:52:57.096191Z digest=sha256:3a9b99e90387c59354464d20b9b74f3396677d1a01df5ed0fc44f61919f537e9

Observation 978d9c79-25e9-482a-a527-83062b2b5560 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.099585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.099585Z digest=sha256:99f6644e0d5dde41eec5fcc7bd1db04ccde21bbeee5f295eaabb1f9c7183302e

Observation b5b17f06-a8c1-409d-bc31-b403998fca68 · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.102968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.102968Z digest=sha256:3efc734f903e037dc7b7ccc285ec38c209bcf8a393e3ba7ae37bd6d975a05684

Observation 12952629-216e-4600-8d99-dee8bdbf7ebd · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.106265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.106265Z digest=sha256:b6aac074a667be6bd6c0e7daa7e81f2bc4124eb3e09eed07b9c0bf9a16accdf7

Observation 0522cd55-09d1-497e-a29d-8c1ef02e290a · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.109673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.109673Z digest=sha256:44cb5ff3958df7ca7e70f13f6929274c0bbb05748623d2e9cff456ae096e079e

Observation 9da6e79b-8d81-45b3-90e9-6d425b8fec82 · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:52:58.052571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:52:57.113288Z digest=sha256:30ece48631518a325f530132dd0fb19109d4ec901a129f65a0c1597018493473

Observation 10fc13bc-9154-495c-9644-f2a0722aa6bf · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.116326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.116326Z digest=sha256:98214f5d0cd388af87710264bd59383f538f9aa538daa9e2fa02c4d1d10233bd

Observation bac83baa-919a-4e5c-b093-0183f15e2861 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Improved Baselines with Visual Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.119801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.119801Z digest=sha256:61880905201552d675f5aad427161f98ac303174ab20bf1631cba5e66116e630

Observation e815a367-7e79-413a-bb5a-7aadad4a93d9 · outbound

This paper cites Visual Instruction Tuning.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Visual Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.123240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.123240Z digest=sha256:6384c1f6c6f9355e7e5d95c57481eea691166cbea5bd4f66f4daf7097716c688

Observation ddbe5737-ef40-4d1a-912a-537731729393 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models MMBench: Is Your Multi-modal Model an All-around Player?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.126790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.126790Z digest=sha256:5aa3ade950f6fbde8b46ea3fd092662c09b96c1c0668566f75aabed2ae501b7f

Observation 98dfef0f-a47d-4f1e-9919-aafa5b46d6cd · outbound

This paper cites MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.130337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.130337Z digest=sha256:292e0d968440e8dde321192ae16e10d64e4af38339e4f316eef70fd7a81f4277

Observation 3bea4e0f-9112-4b05-bd55-2e3dcc715367 · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:52:58.041224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:52:57.133759Z digest=sha256:bdc86b13b0f1a834bc899eee1b874d117c3d9323279391c53a93e501ea8ead28

Observation 0dfaab4b-8d98-403c-8cba-c45de2b8efb7 · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:52:58.029301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:52:57.136930Z digest=sha256:28d352039ea92b795fa177bf990caf6a4320e2b06bac1ee75046c7a298f066e7

Observation c4d95314-2c89-4514-98be-049927005d14 · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:52:58.018345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:52:57.140054Z digest=sha256:c15a3fd6b3ed1d0a25c223ab2eb089a914796279deb03dc47ab377c5c5a5cc2b

Observation 709b09af-8f3b-4fb4-85e5-822d1e66504a · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models YaRN: Efficient Context Window Extension of Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.143311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.143311Z digest=sha256:19be995f0f210c1cd72c643603854448716da9d160fc47f06678be28f52a8aeb

Observation 0f3ed3e4-2436-4fa5-b4fd-3e1941f08c1e · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.146701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.146701Z digest=sha256:f4a199159e2dfb100bf0e6866388d8a3a2a14be5efa84b6b3745bcbe3cb1d9b2

Observation ef15c730-7ab6-4724-817f-3f416cd663c5 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Code Llama: Open Foundation Models for Code

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.149783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.149783Z digest=sha256:fd39d0571cca9852e3255f136dc74eb84793c1d46ef60790fc42890d5ea047f4

Observation ca2486ac-745f-496d-a356-8c207822161f · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:52:58.000973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:52:57.153312Z digest=sha256:c1ae297386b9490d4d17b7b3a0d7c01651c8fdf175afdb5dca735ee322c8b531

Observation cb65775c-9940-45b6-89db-55280d35a964 · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.156602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.156602Z digest=sha256:238b0be34016efc3208979e0c1e4b4c0da3e631641163475712ded1252a7f611

Observation dc934de4-1b24-41b2-b737-87bd7325a224 · outbound

This paper cites A Length-Extrapolatable Transformer.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models A Length-Extrapolatable Transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.159727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.159727Z digest=sha256:dd0bf8e5a7c884962555919893c4c410a9c026d079be5cc0e468f20ac2954cc8

Observation a756a6fe-e069-46d5-ab28-00674315cbd6 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.163241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.163241Z digest=sha256:e8b57ebfcea98fcb07195f86c5792acd9fcbb6a15e6b41df6b57a0a1acb77e3d

Observation e22036f3-f643-4e47-bafb-7be7896b70a4 · outbound

This paper cites The Llama 3 Herd of Models.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models The Llama 3 Herd of Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.166879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.166879Z digest=sha256:16c9e7c0bb5ef235392d81dca12c2f59a7f2f906c22d9c92fcb2ddeea4844cd0

Observation 80ae5947-48ac-4754-b805-abdefc72b2b1 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.170062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.170062Z digest=sha256:79eb1dcfcbd40814e41928dab8fe7bb32e6418799820e9e5fc8c25eda0da0f51

Observation 9c6370aa-cf45-430f-85b5-ead9134862ca · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.173528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.173528Z digest=sha256:4e67eb0a381a276c7286d87c2e5c1a7736d17b7a51483febd87b7e7f4b9f5c47

Observation 1d1af7f1-7ea6-4bf5-b7af-944bef1cc46d · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.176808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.176808Z digest=sha256:64fee2e0145da090726fd954c072e1bf32363be126ca8a77f21c38ba306bf5ac

Observation 86f9e18f-2976-4668-a5b2-9efddf3b0a3c · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:52:57.989794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:52:57.180300Z digest=sha256:9979ff3919922e7ae9c5f62c38c8d7bd8249db8eb5bec647a46008afe3cd050f

Observation 5c9abfd0-f032-491f-8137-ab3180bbf295 · outbound

This paper cites Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.183419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.183419Z digest=sha256:26f49135599f8a24428000f78dbdd7baa20dee49d4d2c1330579968825c0592b

Observation 355e61a9-5474-44fd-8a7c-854e220cc851 · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:52:57.978405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:52:57.187039Z digest=sha256:99f0486a624148dcbd0b410cb7ea8127c4658ff719af2baa3affe293d8c01878

Observation f4d7ae47-6970-412c-bc50-f5048af92f5d · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.190148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.190148Z digest=sha256:eaa86a811eb9527b4e0d73438d8c1dd7b825509f0c4e0b87c8bace378164220d

Observation 242a68af-21f3-40db-b0b2-456ba3e78a2a · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.193369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.193369Z digest=sha256:fbe2eca0d4aa12b12f23068f431d6de1b9a46765f0a6558c2bb0d14d920e4b5f

Observation 0aa71577-1f22-478c-9ae9-b8689e182619 · outbound

This paper cites Long Context Transfer from Language to Vision.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Long Context Transfer from Language to Vision

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.196807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.196807Z digest=sha256:b8682532896c6e9682aab7329f3003f480a3358e9b85b5b1b23b3c4e840d99c7

Observation 9f7ee3c2-44ea-46d2-83c0-05b0f1986b22 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.200325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.200325Z digest=sha256:12c3ed75d3fc32b7c895ea53b0388239f71409ca94c5e075e2ebcdc72aa86092

Observation 556af101-79d9-4455-997b-c2bc0867fcf4 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.203832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.203832Z digest=sha256:c957dd375f56b38bc6b1f0db7c82dd5964ad542ff1e78f1d6cc9e403888ae8bc

Observation 139c5f36-f880-4f55-8ca9-d5c75a11ef7a · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.207839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.207839Z digest=sha256:5f7329a7bf39ba84eb31ba619f052f8433f36d7fe810eb089ebf5eeaf762bfde

Observation 12074ea3-c5e6-478b-a1fe-74048382e5d1 · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:52:57.961123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:52:57.211293Z digest=sha256:5aa72133f3d9c25f161db918f8c27bb5223d639f5f2603ae598420cdf865cff3

Observation 10647a01-123a-413e-90a5-e15ca231085b · outbound

This paper cites write newline.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models write newline

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.214469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.214469Z digest=sha256:1ac85f4fdd8511da16bc00f694a74fbe464a10f96197438ac7f11bdd0300b4d3

Observation fce838dc-b3fb-460e-96bb-29b3c9796e81 · outbound

This paper cites @esa (Ref.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models @esa (Ref

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.218185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.218185Z digest=sha256:80e2e2f743c1203a9aadb6ba9820c6baf01b00c9c1c4d3a83735729285d52cc2

Observation 8be063b8-29c6-44cf-8ba5-167b8a7b9650 · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.221740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.221740Z digest=sha256:fb15d00e138c1f1ff94fa724e16ba91bd6791b0301782ce4abaf8d98da7a4459

Observation e28feb92-f2e7-4e4f-98c7-76051339bc59 · outbound

This paper cites an unresolved cited work.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.225192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.225192Z digest=sha256:281d99f0638edae206a2d43e68410cd67f1f17f5811c95fa7da1221dfd2a93cd

Pith citing papers

No inbound Pith citation observations are available.