Pith. sign in

Paper Citation Record · LEDGER

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models

As of 7 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 2 inbound Pith citation observations for arXiv:2508.19650.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19650 v3

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:40:40.564012Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T06:16:07.090870Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T08:16:47.773102Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved15
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65e74ff3-ca57-4317-b64b-72922a10f5e6 · outbound

This paper cites Pixtral 12B.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Pixtral 12B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.375797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.375797Z digest=sha256:4b9f8e2d0c52af85af7b7e3aad6de8c9732d0de733a84db88f6d5298433fe953

Observation b145b217-8f77-4661-b58a-012bfb679beb · outbound

This paper cites an unresolved cited work.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:40:41.324211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.379826Z digest=sha256:2e1f1fb55dea3cb50739cefe1fbc51e4046ce1c7193bd326cec0f956562843ef

Observation 37c1ab81-66d2-4275-8218-f814e6f8b074 · outbound

This paper cites Claude-sonnet-4.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Claude-sonnet-4

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.315764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.382974Z digest=sha256:0fa42d68bff037b3481a388a24a4e8c5a8d6c04d226a58d0f83703f2d818ab57

Observation 5689d440-2051-4c0b-bc89-e1e4b1c65b74 · outbound

This paper cites Minigpt4-video: Advancing multimodal llms for video understanding with interleaved visual-textual tokens,.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Minigpt4-video: Advancing multimodal llms for video understanding with interleaved visual-textual tokens,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.306761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.386685Z digest=sha256:3222070fb5bd40dee41f8d18f234e9d5dc0cf57da56acf51f91aba145f02c590

Observation 2bece409-1543-4315-8abb-85aba8fc7f43 · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.297697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.389989Z digest=sha256:b7c54257b90ad927794f4717b39ff2615b88ac97d0e4337d33f78ec4a1cf54c7

Observation 059fc12a-ced7-4d08-af87-903701ead891 · outbound

This paper cites Qwen2.5-VL Technical Report.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.393101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.393101Z digest=sha256:7826882385518b180368b714c2cfb4323f83208d318edf7f529113fea580369f

Observation c353df1a-85d1-4771-aa1b-1739b7ecfa43 · outbound

This paper cites Doubao-seed-1.6.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Doubao-seed-1.6

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.288645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.396513Z digest=sha256:ff7e5768c714c7444052f8bb77b4ac48e031190f81c41e22d73595442f37c488

Observation 8b13b3b2-8aa9-48d4-a1ed-33d456936220 · outbound

This paper cites an unresolved cited work.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:40:41.279774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.399914Z digest=sha256:2462b896ffc0283ec34e2b14be0a93a23824e1ba7b0cbce9960f3b3f357cbd27

Observation ce52ab39-d687-408f-801d-df85a4ecd06c · outbound

This paper cites LongVILA: Scaling long-context visual language models for long videos.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models LongVILA: Scaling long-context visual language models for long videos

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.270684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.402904Z digest=sha256:3639098ab161b2d0e34228952180922caaeae3f93260c921a2786e6ff935ca77

Observation 1b2fed83-7c48-4c90-bd44-ae8aa70838ac · outbound

This paper cites Videollama 2: Advancing spatial-temporal modeling and audio understanding in video- llms, 2024.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Videollama 2: Advancing spatial-temporal modeling and audio understanding in video- llms, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.260969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.405886Z digest=sha256:a9d4116638140e29258c27313719fa185a47c5bfa152d579727084f0dde1fc64

Observation 9d43fce6-8ccf-46cc-b5e3-cc9bff5c269d · outbound

This paper cites Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic ca- pabilities, 2025.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic ca- pabilities, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.251670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.408752Z digest=sha256:a5994f001b4d14fecbdcd5155cc83ca78a60c8a09b1a7dd6a37678ad3f7c5e04

Observation f5ab3313-f21c-47f7-864e-7f20a81a650e · outbound

This paper cites Video- mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Video- mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.242533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.411871Z digest=sha256:bd764a5b450d92b7ebc26ee238959a9626917e8e4269967f331c2b8d261ee489

Observation ea51068d-30b1-43a7-97ad-ad4c47deb42d · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Ego4d: Around the world in 3,000 hours of egocentric video

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.233387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.414832Z digest=sha256:9d8fcc7c70d24f582cd89ba3896a9042689627e7aec29873351c87be0d05a3f9

Observation 00566207-44b0-4dc9-9499-bec7e5bbe829 · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.224035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.417531Z digest=sha256:77f25c876261c95e1a8dd301b83595eeea57b0a43828260b7a658bcb64a1a4cd

Observation 2307a43e-3c88-4356-b064-ede6b14e023f · outbound

This paper cites Scaling Laws for Neural Language Models.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Scaling Laws for Neural Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.420423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.420423Z digest=sha256:4452d3755c07ccc47f85822636b57409c7ee4f1e635ba7255feaf6295c89d212

Observation b9dc044b-cff9-4370-b2ce-6b458c9ccfb6 · outbound

This paper cites Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.423571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.423571Z digest=sha256:2103c35428a486945b64236f1600d908c13b18c398a7496129eb9e692b8afbd6

Observation 1f10baab-a03f-448b-bc7f-5ebb11321d01 · outbound

This paper cites LLaV A-onevision: Easy visual task transfer.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models LLaV A-onevision: Easy visual task transfer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.209634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.426686Z digest=sha256:63f03e5c605c94c5c1d169c1f20de9852c34ef0d177c39541bb3b9708530658f

Observation dbb12601-14cf-4791-b2e6-2e166ac5db85 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.201130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.429618Z digest=sha256:2c5e4f8260b9a560f92c1dabefc4fe6d3f7b209ec05c07856ac237203a1c5e1f

Observation 1fc36da3-17db-4a3c-826c-a94a1c94c0bd · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Llama-vid: An image is worth 2 tokens in large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.192619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.432539Z digest=sha256:77e0aa60ff70fadc2941cd726c2053b9adb01e340495168284f07871f6648d7a

Observation a871293f-60e7-4828-a97a-939032e264de · outbound

This paper cites Video-LLaV A: Learning united visual rep- 9 resentation by alignment before projection.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Video-LLaV A: Learning united visual rep- 9 resentation by alignment before projection

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.183525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.435609Z digest=sha256:85860d29c3974ce43776f74be90189b5c55bf4d3d3e4f963406d9caeff3d93d3

Observation adce1126-a8a0-4236-b4ca-617e4f826596 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.438516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.438516Z digest=sha256:e42adbe1fc4088c46b99b6ebb21523f8e016666de2f9c202d0f0ec2e6d69f189

Observation 38c45cc1-d510-49ce-b333-a65a322c07fd · outbound

This paper cites Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.174770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.441934Z digest=sha256:5471357cb949f2c22975259818f511e8db7f6d5db6ea9af3ffefeac9b843d54c

Observation af6c663a-86dd-4ae8-b784-9963b6deba10 · outbound

This paper cites Scaling laws of rope-based extrapola- tion, 2024.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Scaling laws of rope-based extrapola- tion, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.165563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.445184Z digest=sha256:41469f761969ef7f813e551f8a51b3caabb9852801bf882c562e13af79cbc31f

Observation 6753fc71-50bc-4a1c-9b66-c935817c839c · outbound

This paper cites TempCom- pass: Do video LLMs really understand videos? In Findings of the Association for Computational Linguistics: ACL 2024, pages 8731–8772, Bangkok, Thailand, 2024.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models TempCom- pass: Do video LLMs really understand videos? In Findings of the Association for Computational Linguistics: ACL 2024, pages 8731–8772, Bangkok, Thailand, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.156567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.448237Z digest=sha256:ef9bdd41f0549f5afe410348b12981fe73140884a2a803480f51c39a7690c73f

Observation d7681b96-e8d3-4c04-88e3-5b215a9226eb · outbound

This paper cites Nvila: Efficient frontier visual lan- guage models.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Nvila: Efficient frontier visual lan- guage models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.147633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.451552Z digest=sha256:75b77c0dadfbcecab7a9f64a21721928452b5a51d9df5f54aaf3d79d6d8243ca

Observation a0edd6fb-bef5-477d-9122-9dd90f9cbd0e · outbound

This paper cites Nee- dle in a video haystack: A scalable synthetic evaluator for video mllms.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Nee- dle in a video haystack: A scalable synthetic evaluator for video mllms

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.138613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.454556Z digest=sha256:99f2fc520fb1b674effed5134aa1a8aa580d56016bb2314b43f0295f38923752

Observation dbd00143-8289-4c02-8ecb-57b2ec42e662 · outbound

This paper cites Video-chatgpt: Towards detailed video un- derstanding via large vision and language models.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Video-chatgpt: Towards detailed video un- derstanding via large vision and language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.128830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.457639Z digest=sha256:fbb2698a2726303de6123b5533788b238ef30aa670dc7689f03f3308a6beba9a

Observation b27ab64c-bbdf-4ef0-90ff-b9379f5b0262 · outbound

This paper cites The serial position effect of free recall.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models The serial position effect of free recall

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.118795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.460901Z digest=sha256:f250015ce585148a1d402737466d5bb78be53336337a72b6492b80f20646c1da

Observation ab257ea7-ec9f-4d31-9354-ee63024834ad · outbound

This paper cites Needle in the Haystack for Memory Based Large Language Models.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Needle in the Haystack for Memory Based Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.464235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.464235Z digest=sha256:de120540ff9023c7f564956ab0015a35b884addcdeda951d073353a637928937

Observation d8901391-803d-4ff6-b76f-753967c912c4 · outbound

This paper cites an unresolved cited work.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:40:41.109398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.467639Z digest=sha256:8941ff76c12b2a1398693c6a5256b7d8801514ed8f54d27aca9ec31f88a62307

Observation 9c773163-96be-4efa-9531-554b984c2f0e · outbound

This paper cites Gpt-4 technical report, 2024.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Gpt-4 technical report, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.100555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.470740Z digest=sha256:a1b204f67b260eda390c06801b95607ee0f42c15fd2b19c02198aa3f8b93d36c

Observation 4605b57d-522c-4e9d-b7b3-845e14079c8b · outbound

This paper cites Too many frames, not all useful: Efficient strategies for long- form video QA.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Too many frames, not all useful: Efficient strategies for long- form video QA

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.091034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.473979Z digest=sha256:400541c8dfd57edcd8d41ebde02a429c933e009a825bbfe6858fbbda19ec08da

Observation bb56a904-e9af-4aa6-bde5-e33822a78dcd · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.477147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.477147Z digest=sha256:3a00f68c1d5a65ed64927631095fc69855034c43a82f1032c87b8cec761a0d9d

Observation 09913029-dae3-4a30-9500-59136aa2bbb5 · outbound

This paper cites Video-xl: Extra-long vision language model for hour-scale video understanding, 2024.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Video-xl: Extra-long vision language model for hour-scale video understanding, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.081785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.480593Z digest=sha256:e7fd5702d42ed61497fe0b304402ec8346f34829cab581dad60c4b123e8c716c

Observation e6cce75b-3f48-4e2e-aa07-a121e0e2311c · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Moviechat: From dense token to sparse memory for long video understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.483719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.483719Z digest=sha256:8985a9d9021a7c8efdc680b310fd5d3deff2f4ff87e7b8628672f408ee7b097e

Observation 99f41bf0-272c-438f-aee6-58fabe8059a7 · outbound

This paper cites Real-world anomaly detection in surveillance videos.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Real-world anomaly detection in surveillance videos

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.065754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.487027Z digest=sha256:3e53a918f8ed4788edcbc839fb65a6fbf3e2d4619a3f2c31d786d63563a88fa0

Observation d376f8ca-6071-4fce-ae16-dbcbd90f2658 · outbound

This paper cites Mimo-vl technical report, 2025.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Mimo-vl technical report, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.055225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.490189Z digest=sha256:077ede1197fd486180cb395e596afa68ec909475c962cc994f32c8e3d22f9223

Observation e1c435f9-c960-4abd-ac46-0f9a93d2e30d · outbound

This paper cites Glm-4.5v and glm-4.1v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning, 2025.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Glm-4.5v and glm-4.1v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.045168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.493475Z digest=sha256:66d3e327201c79d06cd75161e4c777c3ff47ca16e35180a4e7a1a4977656a17d

Observation 5170d3ad-81ff-486e-9603-2d99133f9c50 · outbound

This paper cites Multimodal needle in a haystack: Benchmarking long-context capability of multimodal large language models.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Multimodal needle in a haystack: Benchmarking long-context capability of multimodal large language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.034671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.496721Z digest=sha256:55bb42b4d32a26df31a16add41e295ae55b7b9d6750862ea9b9dd54d812ba965

Observation 9ce82354-d169-42d4-b055-f69ac1ffb12e · outbound

This paper cites Lvbench: An extreme long video understanding benchmark, 2024.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Lvbench: An extreme long video understanding benchmark, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.024585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.499800Z digest=sha256:b600e1f51ac4f0debeac4cb9b490c2485e78089e97470bb640ae480319cdd02f

Observation 32b649c9-7fce-441a-b8e3-b0e6c0690088 · outbound

This paper cites Needle in a multimodal haystack.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Needle in a multimodal haystack

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.013993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.503121Z digest=sha256:fe3128351f6077a011b21b97d6e8fdd01f9f5e1b621deba422d53b3a24ad9121

Observation a6e007d2-b029-4615-9b19-1fba9a0470be · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.004212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.506176Z digest=sha256:695f50473d3a08af7ffa58516fc7fedcf986021809948ac33bc1968dad420f31

Observation 92174007-f620-44f5-b9a0-82522511e19b · outbound

This paper cites Mit- igating object hallucination via concentric causal atten- tion.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Mit- igating object hallucination via concentric causal atten- tion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.993509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.509168Z digest=sha256:c31d94d77f1dc5d8c24041ee0b92ccde7c8e7a462ecdf789e67ad860e30e4158

Observation 84336aff-ced5-476d-8666-c87a70868ed4 · outbound

This paper cites Qwen3 technical report, 2025.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Qwen3 technical report, 2025

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.982894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.512623Z digest=sha256:23182ee463258b8a94b2d68e40fbc25ea092ca145e5043c0ffa6fb315c811cf2

Observation f5b55030-f5c0-4c8c-9071-6ce23f26e5fd · outbound

This paper cites Re-thinking temporal search for long-form video understanding.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Re-thinking temporal search for long-form video understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.972907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.516157Z digest=sha256:fc3f5dc0658ab8411789db319b62eca3ada471992d548ce31096879e11336d0e

Observation 1e268707-996f-4d21-abc5-7fba0a759ac6 · outbound

This paper cites Lv-eval: A balanced long-context bench- mark with 5 length levels up to 256k.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Lv-eval: A balanced long-context bench- mark with 5 length levels up to 256k

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.519503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.519503Z digest=sha256:87d7cc7d2a74d3c0a3b018ad1642a938ffa432414b57a0783fd0e06925c1027f

Observation 9da279b6-3d77-42df-979f-b468a4b7351b · outbound

This paper cites Videorefer suite: Advancing spatial- temporal object understanding with video llm.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Videorefer suite: Advancing spatial- temporal object understanding with video llm

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.962573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.522801Z digest=sha256:708de3714e7ec1dd3cfdb2ce92a127ef5b0b7e7972225652370d81b30b58ce06

Observation a6aeea15-5916-4ff5-9296-8ba06b9838ef · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.525881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.525881Z digest=sha256:74969cad8a92d989e708bffe98f924d920ff0251083accbc74ff95e7b4fcad34

Observation f74b40df-1c1c-4347-b75b-3322c6cf4af1 · outbound

This paper cites A simple LLM framework for long-range video question-answering.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models A simple LLM framework for long-range video question-answering

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.952156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.529685Z digest=sha256:8b967fc067198d2e0b6225a0a1c576d31b0b4bc20fea31a100b35d3b6a3b61b6

Observation 5ef51ecc-6e95-4b19-a3be-6fbefba61492 · outbound

This paper cites Long context transfer from language to vision.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Long context transfer from language to vision

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.941578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.533027Z digest=sha256:9e289e7157f54126a7d0d3f3ab5c5eea2c686874b84d133f98ad8323904164bc

Observation 495fcdb7-43d6-4704-b7bb-7734549ef176 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.536185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.536185Z digest=sha256:bf9217cd354e86555e768b3e56a76b65f5bc295129a052d7402b0d867b9c978a

Observation c4ad9bb9-4dd1-4262-9ca1-62d36a96c2d2 · outbound

This paper cites Mmvu: Measuring expert-level multi- discipline video understanding.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Mmvu: Measuring expert-level multi- discipline video understanding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.931073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.539944Z digest=sha256:f015a6c1a68839639b453f23665e7d436b7de5e06880486096db9becc6f8c69c

Observation d365e751-59a2-4134-a938-4fcb7e40192b · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Mlvu: Benchmarking multi-task long video understanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.919974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.543150Z digest=sha256:3bd11c3f2b74eab3789665cc4a157afec9a8332b816af8cb6364daed87ce5918

Observation 5b09bca6-5e5e-4d1d-8e5a-92b3ea07ebcc · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.546410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.546410Z digest=sha256:0979ef1fcde2aa4f14046fc59086107c0c0342a1d62446504272873b7e07b997

Observation e7f879f7-f741-436c-a306-c62beb11157c · outbound

This paper cites Detection and tracking meet drones challenge.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Detection and tracking meet drones challenge

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.908840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.549951Z digest=sha256:5d371ad2cbe9e109b56cac0255622dd5a23af022397bfa1b094c50019e7e2356

Observation 7b3cf1e6-74a7-4c33-a225-b56aaea33ff1 · outbound

This paper cites Hlv-1k: A large-scale hour-long video benchmark for time- specific long video understanding, 2025.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Hlv-1k: A large-scale hour-long video benchmark for time- specific long video understanding, 2025

Reference 56

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T15:40:40.897972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.553142Z digest=sha256:dc50ee490326358a8ca95c14012e6d7842642df6ae1b263b65c3b8e176fa78fb

Observation 2f25c3d7-ceb1-4d3f-b2cb-85be3c223669 · outbound

This paper cites three chairs,.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models three chairs,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.886292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.557088Z digest=sha256:66615986c6249c92a65d8083ffee780cde3b4955faed645d65e09ea0367e389a

Observation 2747719c-ecce-4797-a12a-6d8f5f63f1d0 · outbound

This paper cites the lamp is on top of the desk,.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models the lamp is on top of the desk,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.768054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.560393Z digest=sha256:e50bbce7c92bf9e3b7886fabf8af1bb9283e6f8aa5385dedc20fa93840592451

Observation b2ae1e6c-6403-459c-8c95-e8b2abe56a88 · outbound

This paper cites nearby" can be replaced with.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models nearby" can be replaced with

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.757346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:40:40.564012Z digest=sha256:e269912d0e6e6befa4006fdedf6890ef84d83c966018b045ad300c627ac20d74

Pith citing papers

Observation 1bfb9d1e-a881-429b-9703-92f7ff47c0af · inbound

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs cites this paper.

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:58:05.983666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T06:55:01.619441Z digest=sha256:10d4b16ae36f391f34754d5bdb87dbedb134dd341b68f106c8388dd380e5ca65

Observation d29541e9-3524-4b6a-a74a-0cb96578d36d · inbound

M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks cites this paper.

M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:16:47.774503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T06:16:07.090870Z digest=sha256:074191bf3d4816e83560c87384ee42821b2e96624b8eda7c13e0f7d5ded38800