Pith. sign in

Paper Citation Record · LEDGER

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

As of 10 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 2 inbound Pith citation observations for arXiv:2505.24158.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24158 v1

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:37:24.588694Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:27:56.359929Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T03:06:18.765685Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact1
  • verified fuzzy52
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b0f98ab-5ecc-4387-99bc-7583abcaa79a · outbound

This paper cites GPT-4 Technical Report.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.067831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.067831Z digest=sha256:bc7a212400f3656bedc91e10e64de42b5b7a3c6c5e951fbac5ffdaf53a784874

Observation 2a25c717-9def-4308-8ff1-c6049a818e78 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.151840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.151840Z digest=sha256:72200ca29e674d4d5925ddd1b13030816f1310db856779871493874ebfb497e4

Observation 68155e93-d9bf-49c6-b5da-6cad2496cd05 · outbound

This paper cites Vqa: Visual question answering.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Vqa: Visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.236696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.236696Z digest=sha256:db589bbb243f5d337e99e18bd4e94b8bbfce6775ff5589ca0f0161100c09d7dd

Observation 5bc3e1f1-f3a4-4f92-a82e-04cd9341d2bc · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.323580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.323580Z digest=sha256:fdc00bcb7779af88bab029dfccfeca1f9a018f1f2ec20aa831c7e98ac9485c82

Observation e7c4e58b-8675-472d-a7ed-ac3f59df9c45 · outbound

This paper cites Solving mixed-integer quadratic programming problems with ibm-cplex: a progress report.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Solving mixed-integer quadratic programming problems with ibm-cplex: a progress report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.421314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.421314Z digest=sha256:8e70b757149ac4f103694d14049af0b04335b58afdb15de13fe1fac7bb0fb431

Observation 02f8b0ea-6d17-4e2d-9aba-f79e9be26d52 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders On the Opportunities and Risks of Foundation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.506733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.506733Z digest=sha256:d452f67523f613e891d509865f388554de4f43deea3e4d9667c59b4a631f5e76

Observation 8b1037ce-82e4-47c4-a89d-2e18f8978df0 · outbound

This paper cites Language models are few-shot learners.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Language models are few-shot learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.620624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.620624Z digest=sha256:db0ff7aaa0e4cc1ec8f268558b8b583f2c04647013bd0d6d08de5f9e5df58cdd

Observation 896ac7d1-7c57-4717-a50c-28a592b161b9 · outbound

This paper cites Hourvideo: 1-hour video-language understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Hourvideo: 1-hour video-language understanding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:34.498288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:16.681958Z digest=sha256:57de3ab700d3f92bf5a3dffaf8ec8734fea860da59ca56beb0e79d6d4f34b6fc

Observation c3568a3a-b413-41e3-88e1-e0f954ffb8ae · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Sharegpt4video: Improving video understanding and generation with better captions

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:34.329297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:16.776741Z digest=sha256:0c11d5b76113540fc6b8ae387e2fa519b6bfbb535aefa6281f79df63c7b7e297

Observation 27e4ed29-6118-42b5-b76a-559978453cd2 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.941581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.941581Z digest=sha256:c990471c6345d24255f286cd6231007d7a77bdc2c3535602183e0f57556a21c9

Observation e3b44931-b510-4d07-8757-d9f2446ea122 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:34.139239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:17.185352Z digest=sha256:0d337dd9ebb21904a7888070d213ccfeafa51b7639df4756bfa676511960ded3

Observation 762e141e-19cd-44e0-bc51-3a40d50f8291 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:17.344219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:17.344219Z digest=sha256:f2804f8ff0e3e1697a6571ffe5aa29b4b5786a876ef48236e972466bcecd4863

Observation d484031f-54be-415b-940f-bb6fb78340da · outbound

This paper cites Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:33.935943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:17.455893Z digest=sha256:b32fd17f832637f6bbea868c7fdf34b6e87eb8210421b55b1ae4c4adb7932d71

Observation 376adc86-d204-4089-8f9e-cecd44d61290 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders An image is worth 16x16 words: Transformers for image recognition at scale

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:33.752601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:17.620244Z digest=sha256:9bac117c912b04839e64b6ca06929bf7e1385beffdb4484660ce43152407b0e0

Observation 2676de86-cf63-4741-8c40-e4de5cae84b2 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi- modality models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Vlmevalkit: An open-source toolkit for evaluating large multi- modality models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:33.587582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:17.746820Z digest=sha256:6fa56e20effacdb1fef0c233f13a4f72e5b3564ec7936d03568fb881910f7d38

Observation 7f21be05-08c7-49b4-9be9-780019e38bd2 · outbound

This paper cites Slowfast networks for video recognition.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Slowfast networks for video recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:33.431518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:17.902941Z digest=sha256:0db3c02250ecd9c7a4ea3081e85bbf3eaa8e7493ad9c03a15ca5c3ab02253b6e

Observation 809b72ba-d3a4-41c5-8700-937a966fc819 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:33.155237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:18.023177Z digest=sha256:8070cc4e1d3401e2fc1b310ab16eb62e511f83fa079445fc63cdec5a3a56934e

Observation 47dd4bc0-84f5-4db0-b975-1d6a06d65372 · outbound

This paper cites The Llama 3 Herd of Models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:18.121807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:18.121807Z digest=sha256:3a4441070ebbc09d63346a584a9b85b6140cc86a669f5dc27301480db41662c8

Observation 2381eef4-f583-4f76-94e7-c73d273512bf · outbound

This paper cites M-llm based video frame selection for efficient video understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders M-llm based video frame selection for efficient video understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:32.959706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:18.286872Z digest=sha256:c7d0dcb8558b6e6b340f1a48d920412b6c8dabac6ce5183549644d577b6efccb

Observation 3750f7a5-fdfc-4ed9-aafa-7789b8c5546b · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Chat-univi: Unified visual representation empowers large language models with image and video understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:32.786989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:18.409318Z digest=sha256:7944b279653feebc65ad438ecd0e96b1a923ade69459610f2553a76df4dcbd64

Observation fa668bc0-cee4-43e7-b805-46e964487e50 · outbound

This paper cites Language repository for long video understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Language repository for long video understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:32.625784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:18.555490Z digest=sha256:2d7a815da04b64a920531d30c1223d6df57a83415c54058d597d213cdd517477

Observation a8d39fa3-b1cf-4f27-b2a2-2587ffad92a7 · outbound

This paper cites An image grid can be worth a video: Zero-shot video question answering using a vlm.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders An image grid can be worth a video: Zero-shot video question answering using a vlm

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:32.441482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:18.682198Z digest=sha256:6cba7385d816f8adbe71292a710d7cd7aac4875007a26300265cfd397365f3a7

Observation 4241b8f0-2af6-46c9-9a51-1bf2b12aaab9 · outbound

This paper cites Lmms-eval: Accelerating the development of large multimoal models, March 2024.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Lmms-eval: Accelerating the development of large multimoal models, March 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:32.178385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:18.798283Z digest=sha256:9c7a35a523d15b2755f9acc9f07691e3437722854b74a319a35c8a1bfffd1636

Observation 291baed7-2303-4cf2-977a-841f9a3c3bf3 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:18.912546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:18.912546Z digest=sha256:29c1821e5c237e3f19d8c85bac36ad03f502cd44292128b8e351dec79b3a114f

Observation 5d203498-19e8-4011-bb17-7995405bec65 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:31.953262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:18.979088Z digest=sha256:78c66ef951bc840f321da3e964fe26e9cdf1810319ecd01de0d1cc252e84bc20

Observation 60b66c9c-95ff-4c99-aa4f-806c7132fb93 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:19.138328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:19.138328Z digest=sha256:b8728ecc471c52cc0193950b6d56fa7760a292308901431c3aa429af7a75e78b

Observation 01bfaaa7-1cec-48f4-8706-3dc35bce83d9 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Llama-vid: An image is worth 2 tokens in large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:31.759563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:19.222339Z digest=sha256:8e7be9a3978b2be7de34216dae4f9a9ab147e9f7a3310860ccd1bdd1fdfa39be

Observation 5d5a836a-f36a-4515-98b8-f0983000e696 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Video-llava: Learning united visual representation by alignment before projection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:31.614389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:19.291562Z digest=sha256:021b94c5e9fc13c69547caa6732f18372354cdcfa395bdd5bbc18805e11819d3

Observation 638b84e3-ccde-4770-8b05-ab1b71fd4652 · outbound

This paper cites Vila: On pre- training for visual language models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Vila: On pre- training for visual language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:31.426866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:19.353735Z digest=sha256:a67ef7e13ae422f16f018ea498a78d3b17774f09980cf5d548a8c5c34ae501a0

Observation c00adb3d-44ed-426f-a8e7-1acd59c1535a · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:19.424939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:19.424939Z digest=sha256:0207e2e539933456a53003438717c00985dd270f6d7119a59952bd14243c6c95

Observation e16f8ea2-1456-4418-80ed-0254c24804d7 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems (NeurIPS), 36:34892–34916, 2023.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Visual instruction tuning.Advances in neural information processing systems (NeurIPS), 36:34892–34916, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:31.189946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:19.493444Z digest=sha256:550cef8aea20cec8d9148d7928bf61a19aa985f38ca33680c01b99f920b0bece

Observation 6954443f-d472-471a-ba74-4e194b2f62cf · outbound

This paper cites Timecraft: Navigate weakly-supervised temporal grounded video question answering via bi-directional reasoning.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Timecraft: Navigate weakly-supervised temporal grounded video question answering via bi-directional reasoning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:30.918499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:19.566948Z digest=sha256:0550c365a0fb34fd5b63c8a8857f58f33f1595bfd16f1ef45b9c767ae61e4371

Observation b426fb8e-9c4a-44d3-9d07-e7078607e366 · outbound

This paper cites Lost in the middle: How language models use long contexts.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Lost in the middle: How language models use long contexts

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:30.659000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:19.644740Z digest=sha256:88654964406d5053a92be5d7fac78672b690d817810dc65f9f7e2354119a119b

Observation 2f4cc9f1-282b-41ed-9b30-625553421422 · outbound

This paper cites St-llm: Large language models are effective temporal learners.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders St-llm: Large language models are effective temporal learners

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:30.370543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:19.721077Z digest=sha256:ed588af93adc96ef7e68135ef250e63f4e486a39cc1cafebbdfcee42a0d952cf

Observation 81f55e11-6d72-4da8-a94c-73e27f7c50ec · outbound

This paper cites Bolt: Boost large vision-language model without training for long-form video understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Bolt: Boost large vision-language model without training for long-form video understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:30.122315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:19.791097Z digest=sha256:b9a3c9a6a49dfff43acd293c8154be06807e1506ebb110111219cdc95bf89ac8

Observation 5342c0cd-7585-46c2-a601-c2740e0a62c8 · outbound

This paper cites Drvideo: Document retrieval based long video understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Drvideo: Document retrieval based long video understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:29.987173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:19.859288Z digest=sha256:f02a33765fbb61895eeeb31a4cd1368c1701b516f2db5daffe07ca0976e4056f

Observation d4d72409-db44-44cc-b222-457dff76e831 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:19.932480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:19.932480Z digest=sha256:a10d14bdb2a659db31516f738963408bbf78644c6be8340078ad28e00c7c605f

Observation ffbf24f6-cd17-462d-8bb0-3c82bb673924 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:29.773619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:20.004759Z digest=sha256:806bdbe5b9029b8dffbed3b77afeef8ee5b70b5fe10b0335355fb00f9ad34b0c

Observation ae471a22-25b4-492d-a176-bca78cc195e3 · outbound

This paper cites Morevqa: Exploring modular reasoning models for video question answering.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Morevqa: Exploring modular reasoning models for video question answering

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:29.610912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:20.072028Z digest=sha256:0454be3cefd88575e509b9e616bdac66a090526ee9d95d220a1ff312cbac07f0

Observation 0abc113c-9916-4985-825b-dc805945947b · outbound

This paper cites Branch-and-bound algorithms: A survey of recent advances in searching, branching, and pruning.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Branch-and-bound algorithms: A survey of recent advances in searching, branching, and pruning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:29.526632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:20.145930Z digest=sha256:78162a95fd67f75b1d946ec7605ee2b472cf571dd590647e0807c5a74fff8fce

Observation d1c22e62-5148-4f3c-bd01-6c185a24dec6 · outbound

This paper cites Chatgpt: Optimizing language models for dialogue, 2023.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Chatgpt: Optimizing language models for dialogue, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:29.389856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:20.236870Z digest=sha256:7c04380de3e75691f5cb871a6574910c6366616ac3152ff859d434100a864922

Observation 02da2d3a-496d-4709-bcc1-4a52e6737496 · outbound

This paper cites Too many frames, not all useful: Efficient strategies for long-form video qa.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Too many frames, not all useful: Efficient strategies for long-form video qa

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:29.235070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:20.321049Z digest=sha256:2679b4432216459e64fa0b8493b986127269d84bd45b3e4a3545cbc133963303

Observation deb2fcb6-9589-4fa4-b275-b8777c51689d · outbound

This paper cites Momentor: Advancing video large language model with fine-grained temporal reasoning.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Momentor: Advancing video large language model with fine-grained temporal reasoning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:29.090406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:20.447778Z digest=sha256:83bec18614b75f6c98a564d089594c4b3d6654adfa350f9e61236d245ae5a686

Observation 9347cb9c-82d3-4a9e-bf0e-53f2dcf2f7f3 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Learning transferable visual models from natural language supervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:28.954666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:20.542794Z digest=sha256:1ac6b396d04bec1c00a36a56cdfbe7542d4e33fc4a713c9910c60c618a514723

Observation 3bc64dc1-212e-4d68-9cf6-0e9ba1c2a852 · outbound

This paper cites The knapsack problem: a survey.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders The knapsack problem: a survey

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.635358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.635358Z digest=sha256:b4e495156775ece995bc5b6b1acb0f376acb78d78e31ab20a720352a9c14193f

Observation 312e0497-9293-483e-9372-348a673c1466 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.727692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.727692Z digest=sha256:3b55c37e03358b64d5a8e0946a1d4a46357b4c059905172d56637f74c14e3a16

Observation f19b5cbc-5429-425a-8010-6d84e9d4ed61 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.806587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.806587Z digest=sha256:d59bdc39a51abbb004d392286157b2f022271150a1b9335d0811df2f6081b712

Observation fbd4598c-ac5d-4e0c-aeba-4bd4cc9806cf · outbound

This paper cites Two-stream convolutional networks for action recognition in videos.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Two-stream convolutional networks for action recognition in videos

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:28.843270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:20.873682Z digest=sha256:475dda42ed540146bd8e7fd0f65e4a18c57e3b6cc357b13b16de4540c307e09c

Observation d462dcf3-d607-4bff-9bed-dd3901c0c289 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Moviechat: From dense token to sparse memory for long video understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:28.689205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:20.984827Z digest=sha256:833a04e6af167ba8ddea5255d257fc270af4f89c6d215f7dd491a4e80958e327

Observation 7b755651-2206-48b2-9d89-86b81dba6344 · outbound

This paper cites Mdp3: A training-free approach for list-wise frame selection in video-llms.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Mdp3: A training-free approach for list-wise frame selection in video-llms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.085807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.085807Z digest=sha256:122452002492cbd27f416461865e7b53b8ab512851e46a4835cdc6e499cd60e8

Observation 542d090a-bde5-476e-93c8-631aaba47c51 · outbound

This paper cites Adaptive keyframe sampling for long video understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Adaptive keyframe sampling for long video understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:28.544301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:21.159282Z digest=sha256:6fc6ecb641d652b1d35dcc50da807050993fd819c86b45607c871668187dd62f

Observation fb5be959-6296-4046-9e8b-69de40a997a2 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Gemma: Open Models Based on Gemini Research and Technology

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.246351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.246351Z digest=sha256:855619d6cf08c7abcf44944aee2b6501fc5650a981bd988b7436000225b2ee37

Observation 0bb7053e-8653-4f74-a8ed-f522d94005a6 · outbound

This paper cites Cambrian-1: A fully open, vision- centric exploration of multimodal llms.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Cambrian-1: A fully open, vision- centric exploration of multimodal llms

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:28.389747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:21.360353Z digest=sha256:b3e32b89873dd2ce2123037d1d81d8709df8c213b9c82ce884eddc9435d15450

Observation a09a5b92-a3c3-46d7-9c28-29122d677e3c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.439956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.439956Z digest=sha256:a5ae14c4a3ab5cbc6f1202ee6ec3eebc512f9580c41ed64f27ed13655306c06e

Observation 46bcfb17-ae86-4fea-9e96-e026cddb3bfc · outbound

This paper cites Attention is all you need.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Attention is all you need

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:28.221238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:21.559745Z digest=sha256:075ac7e2fb1ee224593af671036a0a7c7f333c17730baa3d9f6a412f53c4b945

Observation a9736535-42eb-4e5f-8817-6603449fa2e5 · outbound

This paper cites Show and tell: A neural image caption generator.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Show and tell: A neural image caption generator

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:28.070943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:21.692309Z digest=sha256:cdff538efc1ea850dba6bc0dfd046c3a1423c4f0c8ac3130e17534d12839d378

Observation cf88f0d5-6501-4dae-b17b-c9deec7298aa · outbound

This paper cites Efficient large language models: A survey.Transactions on Machine Learning Research (TMLR), 2024.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Efficient large language models: A survey.Transactions on Machine Learning Research (TMLR), 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:27.896903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:21.768128Z digest=sha256:d4a97faba701ca3e373d5dadb1283a0f510b66fb6d9429fabc898b841ff1aaae

Observation 8e98f985-ca8e-4fc2-a759-67f3c756a9be · outbound

This paper cites Weakly supervised gaussian contrastive grounding with large multimodal models for video question answering.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Weakly supervised gaussian contrastive grounding with large multimodal models for video question answering

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:27.715659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:21.846512Z digest=sha256:05a02b163d5ae95b7b203d1e4b52d26517afc370a18f2009f4934c15022938d0

Observation 88731b3a-d85e-4308-ab9d-076a50bdcc2a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.942911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.942911Z digest=sha256:30d0065b6685eb5e74c491e3645d828cbb08e2889491b40b88aa16b48b9e3a19

Observation 26926a99-6fe9-4763-ae56-44f168a64c7d · outbound

This paper cites ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:22.035335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:22.035335Z digest=sha256:f53c3702300b8359f21fdbac523811ef25e743afeb4506c71dc8f36385904e8f

Observation ae110e84-5baf-4ba8-94c9-2236d33e803d · outbound

This paper cites Videoagent: Long-form video under- standing with large language model as agent.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Videoagent: Long-form video under- standing with large language model as agent

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:27.552815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:22.139175Z digest=sha256:674646176400a7257562054d20a4885cce12fb1fc94f21d0aae11184a467288e

Observation eb9742ac-5b10-4d64-a3ed-357f9bcc1d79 · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Videotree: Adaptive tree-based video representation for llm reasoning on long videos

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:27.342292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:22.236365Z digest=sha256:3a09ebd82cbf0c4e00418cb84fe707e09b3c5215c26ca03a26b47981fd62be2d

Observation 7cd90885-7735-4b77-8764-6ba6759e772e · outbound

This paper cites Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:27.171688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:22.338769Z digest=sha256:da819b3d656e290e3c2151108779534f358943b28b0eb024b9341fb9a4160cba

Observation 061e4745-6e27-4634-a3ca-076def2997d1 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:26.974968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:22.450389Z digest=sha256:2b87a0c0035c89a59070445acca15e14f427e00f5f92a95f645c2115f2d4ed92

Observation 7998ee8f-53e2-406a-944f-7d9d6e446f8b · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Next-qa: Next phase of question-answering to explaining temporal actions

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:26.807231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:22.581586Z digest=sha256:9331d1dded3c294414c7036f9aecf89e28fe032b2e95ff7001c10787bcc31b04

Observation 519ac6c8-ad9d-415d-bb49-44f01783acce · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Can i trust your answer? visually grounded video question answering

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:26.664010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:22.691030Z digest=sha256:5dcec90a2207059854aea8bbd02a3d78c100724a1605cc4441f524d30150b30c

Observation c3b083f2-ee51-4cb1-a99d-b3b5c2ef56c0 · outbound

This paper cites Effective long-context scaling of foundation models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Effective long-context scaling of foundation models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:26.500839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:22.786503Z digest=sha256:79ef0d6404fed63642b5c90bd1e89bfddb72520b9aea8fca0cd40b19e27b9445

Observation a990e067-485f-4cf9-852f-ee77af800c35 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:22.900514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:22.900514Z digest=sha256:f1af24b461e68b7ef5bc95e0a4cdefd16cf368b96b5894968edcad6b815aabad

Observation cba55f33-db7f-408d-99fa-78317f5baaca · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:23.032653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:23.032653Z digest=sha256:261955b5c93bdd90692ddf6d7d56b5161adcc72953295ff2e2b09eba3007a3a6

Observation cb3d9dd3-06aa-45e8-bdb6-fdb75c4eda2e · outbound

This paper cites Zero-shot video question answering via frozen bidirectional language models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Zero-shot video question answering via frozen bidirectional language models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:26.332155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:23.130555Z digest=sha256:ed99c588b5bf80e7249f8e3bcc979de957c5ea20115a63659e9d6ed7bbf9e1c0

Observation 51652350-edb5-4a59-8448-5a2ab454b129 · outbound

This paper cites Vid2seq: Large-scale pretraining of a visual language model for dense video captioning.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Vid2seq: Large-scale pretraining of a visual language model for dense video captioning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:26.186839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:23.230304Z digest=sha256:7fbb9c67d785a96614db189d7f5e1a8fb77516aef6ee8a0a6bfb340233a59d2c

Observation 9298836e-13e0-461f-ab6c-1abfb7ea459e · outbound

This paper cites Dense connector for mllms.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Dense connector for mllms

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:26.019452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:23.342782Z digest=sha256:445f75eca23af334e38ac81a54ddba019e9b42ef7e1ba7fdef2241f897c9a202

Observation 0ef5db10-d2dc-4306-898f-c0a74ce8d362 · outbound

This paper cites Generative Frame Sampler for Long Video Understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Generative Frame Sampler for Long Video Understanding

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:37:24.889986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:23.440876Z digest=sha256:9a72a760c7ff416e3ce6b10ce691ea9a3a855de2e9ae2245fd9a28d9f948a5b4

Observation e9d53f51-3628-4ccd-8161-bafbb69f4d9a · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:23.584263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:23.584263Z digest=sha256:1b0ddf5a5326b9692e26e3096cdf18663cdf432941b2aef8fede822fcb9d21e5

Observation 1388fada-bbd9-4c81-83ce-e5b474ed503f · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Self-chained image-language model for video localization and question answering

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:25.913672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:23.714914Z digest=sha256:d18788f45aae869ffcf07eac2ebf70931437e064a4b3c76c24b3720390ac7645

Observation a2d295cf-2c3e-4d3b-864f-226d08d397da · outbound

This paper cites Frame-voyager: Learning to query frames for video large language models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Frame-voyager: Learning to query frames for video large language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:25.779824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:23.826563Z digest=sha256:17af7c0c4a85014529f7cf325ee83cde630031aeb77eae9e40b5d17e0d5d40c6

Observation 49e98b79-f970-452f-9e47-f1468db5e4da · outbound

This paper cites Sigmoid loss for language image pre-training.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Sigmoid loss for language image pre-training

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:25.554077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:23.882899Z digest=sha256:a15b1d36e22e76c5e549d6b6aafe1fd14bb79a827d622e1f69c24312129cdf48

Observation 842a74e2-9b22-453c-b188-65da3f3cc9e1 · outbound

This paper cites A simple llm framework for long-range video question-answering.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders A simple llm framework for long-range video question-answering

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:25.380735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:37:23.986712Z digest=sha256:06bd8e0cb2cfcc27cbec0bf1248f895a1a2f0687a504ce072c8517492d72703e

Observation d0f16f33-24b5-4af9-ad4e-4b8e07355c6d · outbound

This paper cites Long Context Transfer from Language to Vision.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Long Context Transfer from Language to Vision

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:24.093163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.093163Z digest=sha256:ff89157454276407d60c772b6693961c101a22e512f4180a86d6b41154d75c79

Observation 79b9db04-5fc5-48b4-aadf-697adeb8f07e · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, April 2024.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Llava-next: A strong zero-shot video understanding model, April 2024

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:24.190869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.190869Z digest=sha256:ce9d77f9bf75e3f7757183e5322afddb289772dc677a66597fff237f9a9f1e7d

Observation b30ebedc-b77a-4463-b3a6-ef8c7a773f6d · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:24.258273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.258273Z digest=sha256:ab38168341b90a98a9fc072aeff418578827bcb17cf3621ffee92eec755bdeb8

Observation b028f069-8365-4ff5-903f-0d9928d4f313 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders MLVU: Benchmarking Multi-task Long Video Understanding

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:24.363162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.363162Z digest=sha256:6556874c1c6064c99f1dea41b1ccf4c611635403934159ad05a28d41c3f498c5

Observation 4793def1-2e8a-4082-a776-2a2c5df55d59 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:24.461768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.461768Z digest=sha256:416db74fdb76d7cbce8a4660ea9a1e88a1e5cfc10e6567331f0954e1e4ae9536

Observation 447bbede-b0be-4fbc-9feb-22827ba5a7bd · outbound

This paper cites Apollo: An Exploration of Video Understanding in Large Multimodal Models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 84

Resolution
malformed identifier
no resolver link, observed 2026-08-07T12:37:24.588694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.588694Z digest=sha256:302bcf5656cb390f4185d1c2c155353f20c4ae329acb6eac99c4782178840947

Pith citing papers

Observation 76dbb3a0-84e8-434c-9c6f-48b3029920af · inbound

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs cites this paper.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:56.359929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:56.359929Z digest=sha256:849bc7202b2b6102be2e7e63972b33a0c369ea999329e046088ce6f603d83e5b

Observation 767c7c7e-b77e-4aeb-8c50-e2dfe77869ac · inbound

CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding cites this paper.

CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:06:18.772638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:06:09.753634Z digest=sha256:de0cd8077c50d951c25ed555b6db7e8965c2b9ca7c3f68a90857fb71a4d8b884