Pith. sign in

Paper Citation Record · LEDGER

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering

As of 20 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 2 inbound Pith citation observations for arXiv:2412.12833.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12833 v2

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:44:50.672180Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.436903Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 7deb4e54-a544-43cf-9ec6-4ae0d1de3e4e · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:49.965596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:49.965596Z digest=sha256:97da70b77d57272c554431f600955c12e3768a917c106354a80b021ae8b53098

Observation 3be2706f-a41f-4581-80a9-75c517c53179 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:49.972988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:49.972988Z digest=sha256:b412d0fae4124dd8f1f6df17eb0182dea6779f9d3938ae352440ce97fe089fbb

Observation c21441fc-495e-413e-ade8-ce4e89f3dea1 · outbound

This paper cites Memory Consolidation Enables Long-Context Video Understanding.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Memory Consolidation Enables Long-Context Video Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:49.986395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:49.986395Z digest=sha256:4d489b42e4b2bd22ef2698e89edf52a99a288cc312b851e78c678c925478151a

Observation c36730f4-6ed6-4fee-be05-11bdb36ab249 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:49.996073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:49.996073Z digest=sha256:b9e572139787ef8c0d9618b28144e445fe4d72616d7ceb5ca83ea17bdd5f8a3b

Observation e07e8417-78d9-4864-b5c2-a8077e335d3b · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.004298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.004298Z digest=sha256:58a6afc97fee8cac69a566c9dac3f7370df7fc5d045460d2866c15eee7a9c0fe

Observation baf08990-0c88-4e02-a7a9-08728f4048a6 · outbound

This paper cites Xmem: Long- term video object segmentation with an atkinson-shiffrin memory model.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Xmem: Long- term video object segmentation with an atkinson-shiffrin memory model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.017130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.017130Z digest=sha256:8b25a75c10e3bd13e0f67749878691938547bbaeb8480aa03abb1cbee2cbc2ef

Observation ebaf58fb-25de-4edf-af88-ca03f5b24131 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.026792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.026792Z digest=sha256:bb4018aec32608e8486f6a9bd0417f8b64c8f994838d3332a2fb123dd7af7207

Observation e3c85fba-9546-4263-8206-8e8ead4454e2 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.041938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.041938Z digest=sha256:64569ab1b2ebb09a7f15aaa859eec16f2f23032e593ea4829caa9dba8c173f24

Observation f3bf6c2e-3ffa-4700-881f-4528e9510437 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering PaLM: Scaling Language Modeling with Pathways

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.052198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.052198Z digest=sha256:838eb4313623580c07252d837a961ca90428ef4a421fc04994f88be48fde21da

Observation 459de546-0dde-49ba-8586-8f548a6d0328 · outbound

This paper cites Scaling instruction- finetuned language models.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Scaling instruction- finetuned language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.702135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.058668Z digest=sha256:a7b8dcc0559c9a487c5596e976ceb8587930be0179540dbe08209d001a58b39a

Observation 42793921-42a3-4e57-ac57-e495bf7b8f8b · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Instructblip: Towards general- purpose vision-language models with instruction tuning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.672190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.064581Z digest=sha256:80002b6e30564e5ed680439041382e33f64dd48fbf2400d906de83770abc319b

Observation 39776e8a-ce3f-4fb6-b1a0-b2a823373b02 · outbound

This paper cites Egovqa-an egocentric video question answer- ing benchmark dataset.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Egovqa-an egocentric video question answer- ing benchmark dataset

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.649998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.069806Z digest=sha256:d7ab14712180aa4f1f75b225686c4bc879e4d1f9b2d5fa7dd0ad5198cff0a2d6

Observation ecacce4a-9f99-4463-b9bf-3f7744a82d22 · outbound

This paper cites Eva: Exploring the limits of masked visual representa- tion learning at scale.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Eva: Exploring the limits of masked visual representa- tion learning at scale

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.624327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.082700Z digest=sha256:d27eaee6393b9cd64342d5181cf62b0046b8d4e639aad95f7ec140b05311b048

Observation c3d9022f-9226-42e2-9a86-01c96d77e172 · outbound

This paper cites An empirical study of end-to-end video-language transformers with masked vi- sual modeling.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering An empirical study of end-to-end video-language transformers with masked vi- sual modeling

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.603947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.088868Z digest=sha256:d892b12bcc541fd44076aed6fb19b411ff0868a2acedd959c9e333a1bd1fcd11

Observation 5505f809-a4e5-421c-9ce2-9e63472a52d2 · outbound

This paper cites The faculty of language: what is it, who has it, and how did it evolve? science, 298(5598):1569–1579, 2002.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering The faculty of language: what is it, who has it, and how did it evolve? science, 298(5598):1569–1579, 2002

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.583196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.098603Z digest=sha256:aa6fbe915db4e47a17db883979607a2d6f34b8397f6d0cfcd6bb1945092aa0cd

Observation 7cf23d61-6dcd-4c4e-b4bb-ede4d7c5d4af · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.559863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.104176Z digest=sha256:8a58e2dec5b82fc7633cbe216f7fae2e9c2ae7713ab59514836f58cefce5d224

Observation 58ad06cc-9093-4e80-ba80-1d5356e57f6d · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.116939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.116939Z digest=sha256:52f7390ba16cde227b74d820eb079ad9ba5019b31aeb1310ddb303c9fa385d0b

Observation 5e83701a-aab5-4b91-ae2d-0fce57c80f23 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.540885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.129443Z digest=sha256:dc9fa5e52b1101fc3c2519b5883774eb95adf25c464c2acc129922e120b10f14

Observation e49bfb55-3144-4a2f-b33b-59b8d134b026 · outbound

This paper cites Language Repository for Long Video Understanding.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Language Repository for Long Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.143816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.143816Z digest=sha256:fffb055d4f8b708a3084199fc859b265301b9ef1cc20b09de383bcc70ec62f5d

Observation 730ded9d-3710-4b50-a4b5-2f055009d5b3 · outbound

This paper cites The Kinetics Human Action Video Dataset.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering The Kinetics Human Action Video Dataset

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.152682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.152682Z digest=sha256:cf408ded73c8b3116f442eaa0af33635acbd385febe65b05cc00951993397fb8

Observation fbc6bf7b-8517-413f-9621-03f0cf75958d · outbound

This paper cites Segment any- thing.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Segment any- thing

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.162208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.162208Z digest=sha256:024bcd08f32c4662765faf3c08630912d0d1e07fe5a2b1463ad337020389ff8e

Observation b9ccc714-7929-4e32-bee2-724ade01dcd7 · outbound

This paper cites Revealing Single Frame Bias for Video-and-Language Learning.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Revealing Single Frame Bias for Video-and-Language Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.169552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.169552Z digest=sha256:441f555d446af6f70c386b8ae4c8203c17b9ff62dd6a1e308a61e7736088dbfe

Observation 5b289956-d3ba-4c97-b720-a1299016f50c · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.456643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.180145Z digest=sha256:c7a50fc78ec307bd0e2be3d6ace072a7d68a79add5bc1785aa7cc48e8f682b15

Observation f605c85a-e8f9-4e88-8dd3-3541048d8410 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering VideoChat: Chat-Centric Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.188674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.188674Z digest=sha256:db43d3ff96f588a331b013fae75ab3e499345d08cbca447631cbb7538be80fae

Observation d2712700-82b0-4fc0-8d28-6e25ba04c1da · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.197811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.197811Z digest=sha256:b8fbae4010f5e46ebacf836a94628897ec8010e81912386b090d1d6274136c37

Observation bca4a85f-1150-4670-9679-34d68d7426aa · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Unmasked teacher: Towards training-efficient video foundation models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.371539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.207634Z digest=sha256:0422b59f4044d733108bc973d067f9b1ffbf3f348e1be8dc895034176d4f53df

Observation 96801e0e-8f50-450f-b028-6908cc30c484 · outbound

This paper cites DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.214223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.214223Z digest=sha256:45feb4909b6cf04af9593fee75ee484e0d7fa28e7c02283a8021b4f67034bbb7

Observation 8446f119-07ab-4432-b5aa-d7509bc90b28 · outbound

This paper cites Tgif: A new dataset and benchmark on animated gif description.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Tgif: A new dataset and benchmark on animated gif description

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.219545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.219545Z digest=sha256:83234a8f40b0767e0ee67774c71a72e99264e7eb82706288f0b61f8b2c1c3fb0

Observation b7821931-0fb9-40ea-bf10-beb77dd26734 · outbound

This paper cites Invariant grounding for video question answering.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Invariant grounding for video question answering

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.324152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.224643Z digest=sha256:a9951bb2aee8b34acb1865121162d57ed965fde0b51762e002a707e5cef3cdec

Observation 9e6bce0e-dd50-408a-aed7-b0a6e70f14f2 · outbound

This paper cites Visual instruction tuning.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Visual instruction tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.230391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.230391Z digest=sha256:7bc78e724b1bb405a465566b00889cea0c54833645e45a2640d5b90aee885342

Observation 70e1b3fc-56d0-42da-9e53-68b7fae9adba · outbound

This paper cites Image Captioning in news report scenario.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Image Captioning in news report scenario

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.236627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.236627Z digest=sha256:a9ab68f92d8d05b5a49fb6922b9ba9d8bf4070ecc1cdddc486a4a546a078c0b8

Observation f1a3515d-bb66-45f2-ae1a-6a6050f1e808 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.253627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.253627Z digest=sha256:3e205a5310d63d4a109c578419d540d48d2a2f72b39d3c05b144f9d166eda82e

Observation b85211e8-21e5-4408-9977-0f9a05d48c3c · outbound

This paper cites Language Models are Few-Shot Learners.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Language Models are Few-Shot Learners

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.260820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.260820Z digest=sha256:2fc4033eba25af6aaccf34d18bc965adc81698e75e3404b6cef70d8b17d3f046

Observation 6f8ce990-2b20-43c6-b479-97139a1ae00f · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering ClipCap: CLIP Prefix for Image Captioning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.268438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.268438Z digest=sha256:c930cc3a557892bcab4f579b7a6fcd0827f3330018db449155f5b42260b30765

Observation 1e77a4b0-b44f-41a9-a51e-b03d7320effb · outbound

This paper cites Introducing chatgpt.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Introducing chatgpt

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.274398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.275019Z digest=sha256:d4fa76a971ee8bb01b9593443445cf23b66a9a7db83aee85b1d9a2f4cb467a98

Observation 42b513cb-85b4-4f4f-9b4a-7939a07d9b88 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering DINOv2: Learning Robust Visual Features without Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.287494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.287494Z digest=sha256:a948a7857d4423f0f5114ffa13cf9d61b136166a19b3a4e2dfe22b96c0468ca7

Observation d5f55249-37c0-45a3-9733-bd3b24ec8439 · outbound

This paper cites The language instinct: How the mind creates language.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering The language instinct: How the mind creates language

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.228315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.295917Z digest=sha256:582d4e7eacd18c4563c858955082ecabb662eb06a2447a16ef940341a50aa5cf

Observation 3d4344dc-e917-48a8-9db2-96f9fb52e066 · outbound

This paper cites Streaming Long Video Understanding with Large Language Models.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Streaming Long Video Understanding with Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.302367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.302367Z digest=sha256:cfb54247e3c42ff956239b6c667a6cd4341dc4505f4002aa84cd701493f48f9b

Observation f0f7bcc5-b8db-4543-96b7-a2f3ec1cc135 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Learning transferable visual models from natural language supervi- sion

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.311340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.311340Z digest=sha256:ec8c85d63dc343fb5f4d941f3cfc0a30e3e0e02575aa5eb1ea6f3260d7d1888c

Observation b4cf9ec3-f3d7-4196-b36c-00a50127b88e · outbound

This paper cites Understanding Long Videos with Multimodal Language Models.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Understanding Long Videos with Multimodal Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.317635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.317635Z digest=sha256:22f7090043ac8a46379ebb772ec0ca0f9c4d42db1b39ea834442ace9e5b46d22

Observation c47738b6-95eb-42db-ba9f-0139b38ed4b8 · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.170911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.329655Z digest=sha256:f0200e30c43518c2548c2d8a827e8229e1a4d3307080b1f2e13d83c6680dd608

Observation 129f7344-ec23-4378-b47d-8a072dd77879 · outbound

This paper cites Fusecap: Leveraging large language mod- els for enriched fused image captions.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Fusecap: Leveraging large language mod- els for enriched fused image captions

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.125471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.338125Z digest=sha256:38241c35dab034b37eca64433ab4481cbe73bb214f7723bfc8f955344c72b13a

Observation 8380ebfa-8481-47a7-ab0d-9282dc5caa99 · outbound

This paper cites Talking about large language models.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Talking about large language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.095985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.352522Z digest=sha256:305f0cfbda1f4a2d4a4692b4e1833531ffd61ca1c9a8d76509cb3755a6a4e825

Observation 77f71f1d-843c-4dfd-b58e-b38000e6fa67 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Moviechat: From dense token to sparse memory for long video understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.058154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.368163Z digest=sha256:4325e379bb6f004c0c06fb43ddb3485ccb72cc351587844d22eb1f9424da9a83

Observation bd70c989-4d89-496d-8f9a-fa1479aa1f85 · outbound

This paper cites DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.378668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.378668Z digest=sha256:5ba8e03a2fb283f46063f7adf0530fde4b4d365b15294e9a92462c0779973f19

Observation 24f70b62-e83e-48ee-a573-1d3d39749a8c · outbound

This paper cites Koala: Key frame-conditioned long video-llm.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Koala: Key frame-conditioned long video-llm

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:52.028251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.399514Z digest=sha256:cde86763e9afd3cc42f2899a51c7f460a6b842c847b585c088fa0d750d083e82

Observation b6df8fd2-1875-498d-968d-5377f094bf1d · outbound

This paper cites Galactica: A Large Language Model for Science.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Galactica: A Large Language Model for Science

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.415820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.415820Z digest=sha256:2545f70d88d80d1f91b45a593106f3bcd879bc3e605fbc2f855e589663ad7255

Observation 8f02d7a7-a679-4bfb-ab42-cabd6dd4f453 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering LLaMA: Open and Efficient Foundation Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.427239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.427239Z digest=sha256:cf50f9b7be2a26ac2c8ba6fd92f1f178ece49bc69d7fb6a4aca72ac5ec2fea43

Observation 85b72e60-54db-4ba4-a526-492fda7f5264 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.435738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.435738Z digest=sha256:bd570115bdda848683d874d84143eb5f583fef2ab3dc5a3eb590683c70d6568d

Observation 1bc027d8-1995-4db6-b25b-3a116ff609c0 · outbound

This paper cites Computing machinery and intelligence.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Computing machinery and intelligence

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:51.993008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.446592Z digest=sha256:d50425f7a9977406a180a1e04ea16a116830849e82948f8cd5219f3bd1da07d4

Observation f2c664a1-9aac-49a4-9801-7c29e34152d5 · outbound

This paper cites Attention is all you need.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Attention is all you need

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.453453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.453453Z digest=sha256:884dfc26f0367a04de76ed9b3c5c6f65edf9435e158fcc9fd3bea47cb81475f9

Observation d70177a6-aa19-42c0-9f6a-30dd8af9135f · outbound

This paper cites ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.459739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.459739Z digest=sha256:83a2902850520c1152fc61f721ab50cfeee3d9db80f8d52dc39e0cf751b7ecad

Observation 27daf447-93fd-489f-96da-22433a18bbbc · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.468860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.468860Z digest=sha256:4e9f8f84319d774b00b71a64085807b7b4edb2b80f50206d7a22690e74e10d29

Observation 36d93877-eaca-47ca-95b2-7fefe806919c · outbound

This paper cites Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.476422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.476422Z digest=sha256:73a67538ed8b719bc73d2020d70be79136b5d3acd5a2d32331bde57d9089c7d3

Observation 168359ec-92d4-4ff7-be2d-8a08834f87b4 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.497730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.497730Z digest=sha256:3e308bfd8d17a06c3e5c76ac73b024ea9b38932f99dbe728b091493e8c51fbe4

Observation 59c53a00-e26c-4fc3-ad06-d19b9a5aba0a · outbound

This paper cites A large cross- modal video retrieval dataset with reading comprehension.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering A large cross- modal video retrieval dataset with reading comprehension

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:51.923186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.517706Z digest=sha256:8dfd6a93247da3da1829e304ba4fc853f92f3c6a1cdf08dcf9fb05d5ab74f2fb

Observation 1bce6f61-334e-4adc-b45a-892d1cfe4f3d · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Next-qa: Next phase of question-answering to explaining temporal actions

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:51.894177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.527453Z digest=sha256:1d1e13ae91da9f628cc2951f4fe4203d25a0388a04bd104616a2badb5f5e716f

Observation 16e2c91b-832e-470d-8db2-dcfbb173dadd · outbound

This paper cites Video as conditional graph hierarchy for multi-granular question answering.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Video as conditional graph hierarchy for multi-granular question answering

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:51.869532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.551545Z digest=sha256:5c1bba6608c1006e70bc876ff48e72ead987d8eff8a65d105788783368de9df7

Observation e4ffe5f4-c9d2-42ec-a1d7-bcd935e56afb · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.561368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.561368Z digest=sha256:a9f1b3cd6ca2c42792fe9ecbb23ed784a0b81fa13616e89fd818ff7bd23a5ac1

Observation 7ac0a1a6-b20a-4c90-a3ac-a3664442ae4e · outbound

This paper cites mplug-2: A modularized multi-modal foundation model across text, image and video.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering mplug-2: A modularized multi-modal foundation model across text, image and video

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.568823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.568823Z digest=sha256:8f9750500c5184706a466e527a600845bfb07436277183809c60da343ed0d39b

Observation 2645119b-22bc-472e-8c14-06d5518652da · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.577987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.577987Z digest=sha256:9058716b1ae16525679253366b86b047c6b2d50f82ecdda557151754e91dc8d0

Observation eb13a480-ead0-4234-ab84-56823a11744e · outbound

This paper cites VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.583896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.583896Z digest=sha256:462ab8d1b24c33dd640801185ad9f7e3124f6d82376c017855044c8229f5a7e9

Observation df907a64-66ec-4748-a479-99be86623005 · outbound

This paper cites Just ask: Learning to answer ques- tions from millions of narrated videos.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Just ask: Learning to answer ques- tions from millions of narrated videos

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:51.804454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.593654Z digest=sha256:743daccdcc3cc189b728d70e51c108f6c6ab9027c4a904ca68b42e9423e12cdf

Observation e89c310c-625b-48ae-af7f-6b72aac62960 · outbound

This paper cites Zero-shot video question answering via frozen bidirectional language models.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Zero-shot video question answering via frozen bidirectional language models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:51.781688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.602419Z digest=sha256:4f5098fbf25ba087c3f9ad58aa78ce432bf701a449594ec789b53dfc3442e2ef

Observation c8cbdc7d-43b5-4171-a8d2-750e90176b8d · outbound

This paper cites DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent).

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.610768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.610768Z digest=sha256:736c75ec119a4cfeb94daa100cfde3e2c9f74bb8d6d1e42a7b02a2379c820c67

Observation a7a8556c-9426-4a4b-a846-eb87ab881c96 · outbound

This paper cites Hitea: Hierarchical temporal- aware video-language pre-training.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Hitea: Hierarchical temporal- aware video-language pre-training

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:51.759999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.618288Z digest=sha256:6e5f7f2bd582822b8bbc8fa7c392b45f6f60d4aa9d62adb39e29cf46f7bfcc6c

Observation 5613c14b-5b87-482b-88e4-6f3b69eb3f1d · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.626507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.626507Z digest=sha256:2100c4c4380807898f652a512594224345e2a50087ffe8089539e1690f35e8af

Observation add73827-01e9-40e2-a263-df818fe155c7 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Self-chained image-language model for video localization and question answering

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:51.738864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.639033Z digest=sha256:079dd121c13b0a7cb1293f1bf125fbae0362a4e5461143275e459e797ba23cfa

Observation 1e15c84f-a3f7-4480-a4c2-9413a1cb7b0d · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:51.702530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.645539Z digest=sha256:315861dbafb7d6aa85b0e0c04f40ad71b7fb3c73e5bdb29104b5431c580df467

Observation 1dde4ae7-c105-4d57-9d7d-fbfb707794a0 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.656933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.656933Z digest=sha256:660ac6f5775c62625e76be4a23b84e54cb8e1ae21d8c23f7f9e4ea559cb39908

Observation a2b53762-26be-46d9-bf18-8476c458ca29 · outbound

This paper cites Streaming dense video captioning.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Streaming dense video captioning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:44:51.660449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:44:50.664971Z digest=sha256:423ee65e109483dd460d8deab08b95807cd13a5f40f6edcb613ccde614a4373f

Observation 8732139d-07c6-4e1d-8a75-29f261e19278 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.672180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.672180Z digest=sha256:7907a0278c96fdde529262ce4153f86cf36f9585ead7116ddff092add8b78b42

Pith citing papers

Observation 4cb02e7a-cd0c-4d8a-aa94-689115e96931 · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.436903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.436903Z digest=sha256:8d0da47f76c5ba3ed0061f1ab57403af80dfab2731ccd4f3a0e3a54368dd9042

Observation 0e828d07-d481-4c27-b553-2541daef71b7 · inbound

Prompt-Adapter Context Routing for Parameter-Efficient Multi-Shot Long Video Extrapolation cites this paper.

Prompt-Adapter Context Routing for Parameter-Efficient Multi-Shot Long Video Extrapolation FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-07-08T04:34:31.024844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-08T04:31:31.924913Z digest=sha256:d7450a21db2034c85aee9bd2f524668c3fba7e9acebb709af66921ac73564ad1