Pith. sign in

Paper Citation Record · LEDGER

VideoChat: Chat-Centric Video Understanding

As of 11 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 100 inbound Pith citation observations for arXiv:2305.06355.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.06355 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T23:30:00.457974Z

measured 159 of 159 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 100 of 209 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:41:47.548723Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact29
  • verified fuzzy28
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

90
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 4f1d7275-bcfb-42f1-8244-508f8918f414 · outbound

This paper cites Openflamingo, March 2023.

VideoChat: Chat-Centric Video Understanding Openflamingo, March 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.809695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:d7ad8c1fb1e77029e756c6b07a637fbfc8c41d09875a5a2242e7b34703ce71ed

Observation 243c31df-d978-462a-b328-eee78fd535be · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

VideoChat: Chat-Centric Video Understanding Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.749928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:be4c606c90a70801b2b0004077e90b1e0d716bb525f21505fd24354948e86400

Observation 0b04bd30-56ac-4b1e-97a0-6c38aacacf6b · outbound

This paper cites Language models are few-shot learners.

VideoChat: Chat-Centric Video Understanding Language models are few-shot learners

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.754276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:c92900d53e41a713f28c3551900293178e390d0019bbe511a5f36f8c6af3f944

Observation 38cbb19d-6359-4c6a-8cbd-94ac5dabec98 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

VideoChat: Chat-Centric Video Understanding Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.758413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:4a5b0bc4a5e68f32096bd0977978f09285d404cc166f8d34d86422599d0403e0

Observation d6438f78-fafb-4231-869f-24ece27e1510 · outbound

This paper cites InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges.

VideoChat: Chat-Centric Video Understanding InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.504755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:42105826f144842850d371b09adacd0c220ff4d629ec30c3d1ade68ac0066cdc

Observation e3d6e35d-5704-4e6a-80c4-1fd9563f1deb · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

VideoChat: Chat-Centric Video Understanding Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:30:00.510458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:c1a624448e85aaa022f845947b140389f499204ad1bdccb01f943cae84fe6c3d

Observation f8fcdbca-4181-4598-b423-7a3f47e98c71 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

VideoChat: Chat-Centric Video Understanding Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.769975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:0d42505ac79bc400fbd8e19f89098d734139f7dbdf7feda5a8c5f0e3ed144171

Observation 5a610f2f-ed84-41c7-ad5b-5ff2f14d57fa · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

VideoChat: Chat-Centric Video Understanding Scaling Instruction-Finetuned Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:30:00.519803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:e39d83d33e19a7911c222c1b02c647c00d32c364a66cacda58427c5a5a368a52

Observation 06c4edb0-be53-4a0a-a1e8-1d2d72fec63e · outbound

This paper cites an unresolved cited work.

VideoChat: Chat-Centric Video Understanding Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-13T23:30:00.778125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:a1fb6579854cded05af8c2f98568298d6c46598570e36a4a1ad0bc6d8d8fa38d

Observation 448d3caa-667f-4136-bec8-b84ba513e11e · outbound

This paper cites Stablelm: Stability ai language models.

VideoChat: Chat-Centric Video Understanding Stablelm: Stability ai language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.781707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:8277cb9f5468561dc40364e3bf4d5137b268f6f341982ffc1c582616414687c3

Observation df5dde4e-4fda-4135-b59e-247aacefc6ec · outbound

This paper cites An empirical study of training end-to-end vision-and- language transformers.

VideoChat: Chat-Centric Video Understanding An empirical study of training end-to-end vision-and- language transformers

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.785457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:43ba274e3105e45715fb57227fdf36609f2d20434b04993dc3dedcd9b8b694b5

Observation ab787156-95d0-45d3-be18-1de421898834 · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

VideoChat: Chat-Centric Video Understanding VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.538637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:cf9744e55eccc3fe38985fa0cbf37b1356fcba4855801c33b3cc1c22ce5c3848

Observation 16a193ea-9225-410a-a0c7-982b383dc647 · outbound

This paper cites Scaling up vision-language pre-training for image captioning.

VideoChat: Chat-Centric Video Understanding Scaling up vision-language pre-training for image captioning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.794460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:45993eb436f8bae74622aafa2c57007aa9e86bf024090b950c3c74aa1686ca74

Observation 4a99b478-b6f0-497d-913d-cb6931eee182 · outbound

This paper cites Language Is Not All You Need: Aligning Perception with Language Models.

VideoChat: Chat-Centric Video Understanding Language Is Not All You Need: Aligning Perception with Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:32:23.026112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:4be5d36abf5ca9cb5fda4176d46e70b86274c5972a546631f28488bf29c3275c

Observation 5d6d4010-0a8e-4b89-a0fd-e76c14b4cc0b · outbound

This paper cites Tag2Text: Guiding Vision-Language Model via Image Tagging.

VideoChat: Chat-Centric Video Understanding Tag2Text: Guiding Vision-Language Model via Image Tagging

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.498208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:3f1511a94c1d125b1ca2ed6b101779f964a8f6eb1ae2164004585e2589a10da8

Observation 90bcb349-2d2f-40f7-8277-c3e4861f24d1 · outbound

This paper cites Dolphin: General video interaction platform based on llms.

VideoChat: Chat-Centric Video Understanding Dolphin: General video interaction platform based on llms

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.814032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:9a95e6397385fb9cc622ad1b4bca5c9ea8c22778ced8ee295eb5af587244eabb

Observation 6fc4f452-f14b-4c0a-8de7-91bd62e3ea64 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

VideoChat: Chat-Centric Video Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.821325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:511d1e6804c4b56cd43178cc2966c8092005d3cae8d83a34202a6e024f34f0db

Observation 072ee880-f327-46cd-82a0-9d9f90a051b9 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

VideoChat: Chat-Centric Video Understanding BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:30:00.555232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:08728551e0e8cec10bb24937316147827063ba03e6f3e7774e6542d9e31a7d7f

Observation dc230461-5565-47ef-92c1-357270b0a820 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

VideoChat: Chat-Centric Video Understanding Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.830565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:3cfab135c27fde6fb680f5812c9356a06d8f5362909d7e22d01271611784773d

Observation a5b874da-8fef-43d5-8329-ae9d9f30a5a9 · outbound

This paper cites UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer.

VideoChat: Chat-Centric Video Understanding UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.561864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:4e3c97ce74a5e6db563533f786e1d7301b4fdd86644a37d14b932aa0b976ef2c

Observation ab5c4fc4-b242-4bda-93f8-b7df99450714 · outbound

This paper cites Unmasked Teacher: Towards Training-Efficient Video Foundation Models.

VideoChat: Chat-Centric Video Understanding Unmasked Teacher: Towards Training-Efficient Video Foundation Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T23:30:00.572703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:f84e9077ea38e6acfb4e38e20374fdd35e419285b5d8158deb52b8a1c5305740

Observation 95867805-c515-4620-b73c-0fb3e2bf0c0a · outbound

This paper cites LAVENDER: Unifying Video-Language Understanding as Masked Language Modeling.

VideoChat: Chat-Centric Video Understanding LAVENDER: Unifying Video-Language Understanding as Masked Language Modeling

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.578691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:b167fe0b3e007fe478f2e1b415d32bd0b0a1eb99996de9ccdfbc3d221d33a44e

Observation 78b1790f-89ba-42ad-9701-6e8853b2cb3b · outbound

This paper cites Learning Spatiotemporal Features via Video and Text Pair Discrimination.

VideoChat: Chat-Centric Video Understanding Learning Spatiotemporal Features via Video and Text Pair Discrimination

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.585193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:06071fd0e49fb5dedbf25a0d7cd8c72def98cda26a780e9d31e4b01faf635544

Observation 93260ef7-5c24-485d-93a9-050fee0c7104 · outbound

This paper cites TaskMatrix.AI: Completing Tasks by Connecting Foundation Models with Millions of APIs.

VideoChat: Chat-Centric Video Understanding TaskMatrix.AI: Completing Tasks by Connecting Foundation Models with Millions of APIs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.591327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:55fec9d12a723522c8f5df0effa3ab20088fd710fdcce4570f258b838d75369e

Observation d5d1a375-3258-4c0a-8cce-5d2ae4dcc3a5 · outbound

This paper cites Visual instruction tuning.

VideoChat: Chat-Centric Video Understanding Visual instruction tuning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.857147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:02d0bd3d38f9cba0d87f92cf538ad96cc90dd6f05a34761c36cde1b786f22b20

Observation dd337dbf-d56d-4135-a28d-da0da5248840 · outbound

This paper cites InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language.

VideoChat: Chat-Centric Video Understanding InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.598826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:fa897ae129e5c67948871805c6b00c0c74cdd8f49c9467ffd572edce58838af1

Observation 8fa12758-6454-4ca9-93b2-12ee0c0b55f7 · outbound

This paper cites End-to-end learning of visual representations from uncurated instructional videos.

VideoChat: Chat-Centric Video Understanding End-to-end learning of visual representations from uncurated instructional videos

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.866241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:f76839b9ac3807e123db2dce2d1a35510d156bd926058e91bd28e705bc5a8ec2

Observation b3b60760-69fe-4464-a775-e10d6d0dfea3 · outbound

This paper cites Cross-Task Generalization via Natural Language Crowdsourcing Instructions.

VideoChat: Chat-Centric Video Understanding Cross-Task Generalization via Natural Language Crowdsourcing Instructions

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:57:29.846007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:2a102071f1395be94dc05cb17c31f0fe52c332fbca1bffed97fab7f20687a08d

Observation fb55e015-b5ab-44a4-8637-d8ecb7375671 · outbound

This paper cites Gpt-4 technical report.

VideoChat: Chat-Centric Video Understanding Gpt-4 technical report

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.873300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:d35b4cbdaf15ea82b0dab4b758fca669cae1730f6f5ada2537883364656d28bd

Observation 3fdb47b0-ef44-4665-a138-0980a99c4e32 · outbound

This paper cites Chatgpt: Optimizing language models for dialogue.

VideoChat: Chat-Centric Video Understanding Chatgpt: Optimizing language models for dialogue

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.878360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:8f007c326f4dddec9202fbeff3316f738eabb478e921e3b3f62f956a76fb1dda

Observation c8762e9d-e69c-4a4d-a13b-fc8b84b1dbec · outbound

This paper cites Im2text: Describing images using 1 million captioned photographs.

VideoChat: Chat-Centric Video Understanding Im2text: Describing images using 1 million captioned photographs

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.762228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:323e1df6b079822f1a8268b0cee97873a5f5b3dda9dca73dc4f1433a52f65822

Observation 37268e7b-c1b3-4473-ac6f-e599e20b9bed · outbound

This paper cites Training language models to follow instructions with human feedback.

VideoChat: Chat-Centric Video Understanding Training language models to follow instructions with human feedback

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.766064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:7b516376389f1a6fbbd08f27a823a678bd56d671a23ddadcfc97ffcd80b7068e

Observation 1f5592a9-5b6c-4a81-bcce-e25bde195f02 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

VideoChat: Chat-Centric Video Understanding Robust Speech Recognition via Large-Scale Weak Supervision

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:30:00.609860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:7679de49429c55f5b36bc6c3d6fb6059119573b0b7936bb5b18477e7441f1e53

Observation 837ee617-26ac-4b68-91ef-73d73f32acd8 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

VideoChat: Chat-Centric Video Understanding Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.789810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:41029293dcfe43ebccc87f76e0900a7bee17f0a169a9cfe7794e5af4b625acfb

Observation b18076b8-4fd7-4dcd-90da-eb9297708857 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

VideoChat: Chat-Centric Video Understanding Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.801744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:a0c4c06f2cc361ae1143b6962ba2afbc43283043bdf22005449d01e2f2bb71e4

Observation 7825d19f-eb53-4f0d-bfd3-6f58d4a5bcf9 · outbound

This paper cites How Much Can CLIP Benefit Vision-and-Language Tasks?.

VideoChat: Chat-Centric Video Understanding How Much Can CLIP Benefit Vision-and-Language Tasks?

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.616454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:cb300881f8be70605e37568607ff6704dae8986a13156a257bccaec35c030d49

Observation 74912950-f8ee-4edc-b5f1-af1b1f977abf · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

VideoChat: Chat-Centric Video Understanding HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:06:45.816188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:d01cde2fee6f3fc9c6b1b4e991b1116f6015bd22ac96de786b6e287b210d0a9e

Observation 1a772ef6-6cb4-483e-925b-dac2e4448560 · outbound

This paper cites Murphy, and Cordelia Schmid.

VideoChat: Chat-Centric Video Understanding Murphy, and Cordelia Schmid

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.839773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:ff46ef9e802003622dae868b367d19e288ba5ebaff03f73d2d0b4df6abe03277

Observation e12117b1-6c89-4e1c-a853-cc27b8436508 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

VideoChat: Chat-Centric Video Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:30:00.627028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:c89a20c800abe8b1dc41625c10c04a9bd74901161a98f99785e7118d3645561e

Observation 34774cd3-7749-47dd-b53b-e80404f37492 · outbound

This paper cites Hashimoto.

VideoChat: Chat-Centric Video Understanding Hashimoto

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.850477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:479ca9a9f6591740a2b3b1e2ce8a096c73b72138792b4a44fce12ac5eb609110

Observation 8837ab19-0e7f-4e4c-9102-0676edff31b7 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

VideoChat: Chat-Centric Video Understanding Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.854067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:bc105b50de5f4cbac1ce20d4bfde77b3da0dc97685de78bc780ef6af10b5b4b4

Observation c1539f79-b1e3-49b0-a27e-acefe3128c87 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VideoChat: Chat-Centric Video Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:30:00.639179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:05f56f3bd8872a944dce93993f443c329b84886556f2994dce35af90e9eb303c

Observation 3a02762d-64c5-4419-932f-1de2d9e32d0a · outbound

This paper cites All in One: Exploring Unified Video-Language Pre-training.

VideoChat: Chat-Centric Video Understanding All in One: Exploring Unified Video-Language Pre-training

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.646028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:3ac660c5c4f160ed4806d4c6bc1f70b5b9b00564c2b210926853fb8a845f0027

Observation 8289f058-a571-428a-8d0a-0d3825f37c8d · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

VideoChat: Chat-Centric Video Understanding Videomae v2: Scaling video masked autoencoders with dual masking

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.773726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:e7a8653f88b59b32e16b8204fb8fb7a961d8c048781c7fb7c94ec236cd1a93c1

Observation 5187eb9a-34b1-494a-9c95-6e71d691d44c · outbound

This paper cites Internimage: Exploring large-scale vision foundation models with deformable convolutions.

VideoChat: Chat-Centric Video Understanding Internimage: Exploring large-scale vision foundation models with deformable convolutions

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.826143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:41ec9d968b1b4af07bcb6f396aa637612a5554a93b762a92d98804a875e9d8a7

Observation 01890349-770e-4c2c-9b11-0c15604e7c9d · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

VideoChat: Chat-Centric Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:fea6eeca6d2de47d6ba605c10b48e38a5e6b84c2fde11fdcd395222921aad2ce

Observation 71cca06b-0916-4d52-a063-8629de8e0e6c · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

VideoChat: Chat-Centric Video Understanding Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:30:00.663677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:a9ba761591f4665306302140d3fbb447b5ed2ef3fc29d5fb238b9bd35392d23e

Observation 52ccaa0b-309f-4256-bf89-57f0f0adba60 · outbound

This paper cites GRiT: A Generative Region-to-text Transformer for Object Understanding.

VideoChat: Chat-Centric Video Understanding GRiT: A Generative Region-to-text Transformer for Object Understanding

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.669366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:ed333b1f614a6fa46ed52f35fd18584d7eb4d97535fc1ea9bc6d57f8c032f0c6

Observation b9b1d4cd-aa91-48eb-b583-b569e1e5952e · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

VideoChat: Chat-Centric Video Understanding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.675840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:f89ea8d8d9914332ca39398c0c227bde16e67f009cdc1ddc0a43cde2acf8ac90

Observation 013852b1-7a80-4f7e-b826-b5e1872d8a59 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

VideoChat: Chat-Centric Video Understanding MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:17:58.977376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:a72636b48fcaa5cfaadf19406c3c517b6484fbe8c4bd342e85930e8e68be1209

Observation 25df6251-3b0c-4dbc-8297-b07522936587 · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

VideoChat: Chat-Centric Video Understanding FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.687974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:14ee22ea345c76772648e74ac9a715de1aa82ba5004229f81e5ec313d9631a16

Observation 90037464-168e-4e64-8291-c62f5b392ac3 · outbound

This paper cites mplug-owl: Modularization empowers large language models with multimodality.

VideoChat: Chat-Centric Video Understanding mplug-owl: Modularization empowers large language models with multimodality

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.860438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:c4bee805c3d59dc232c828ca36d49d761dff41c64e547a65ded08ecd466f381a

Observation 33860558-ca6c-4577-a7c3-470807f3a221 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

VideoChat: Chat-Centric Video Understanding Florence: A New Foundation Model for Computer Vision

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:38:09.598269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:2813fbbd3ed295b0565fe0712147d93632989daf63efcabdbe116a0e92a55bae

Observation 5345611c-17fc-423d-9288-17a8db0c6650 · outbound

This paper cites Merlot reserve: Neural script knowledge through vision and language and sound.

VideoChat: Chat-Centric Video Understanding Merlot reserve: Neural script knowledge through vision and language and sound

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.836023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:99ec04625c8ef85f1dfcc08fcb494250fcfad44109236b4a0d3dba949e021c26

Observation e4788a76-ddfa-4a4e-b0ca-ca9660cf8c48 · outbound

This paper cites Merlot: Multimodal neural script knowledge models.

VideoChat: Chat-Centric Video Understanding Merlot: Multimodal neural script knowledge models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.847015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:1f3e5ebd491a14331e1b2f81cd2c45dbba43429eba076d6d7f7fff071b19ff60

Observation 4ad12f5e-f804-442c-856d-7776cf171457 · outbound

This paper cites GLM-130B: An Open Bilingual Pre-trained Model.

VideoChat: Chat-Centric Video Understanding GLM-130B: An Open Bilingual Pre-trained Model

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:41:34.452877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:c3bf9c7fa00c26e20b8ac65d630cfc165f9755a365ca02f127810f5ef22602f7

Observation 00743eca-d43d-4f66-8b69-91eb324e821b · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

VideoChat: Chat-Centric Video Understanding OPT: Open Pre-trained Transformer Language Models

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:30:00.739742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:2d41fb64d9b37918256c6153d6fd93b0907508d54abd302e185a783b6955eee0

Observation 8a858f91-89d1-4bed-a889-c4f5b6260a7f · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

VideoChat: Chat-Centric Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:30:00.745883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:b096fa5dff6c1245ae28810a7c19663d90f00763231e93303da9a0fdca9f48a4

Observation acf393d8-c28a-4406-8a0d-58d01bf290c3 · outbound

This paper cites Describe the following image concisely.

VideoChat: Chat-Centric Video Understanding Describe the following image concisely

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:30:00.870039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:cd6a218663b42e34b8b2df8e9aef30e4b3bc809c548c78ebfbf25e49baed1430

Pith citing papers

Observation ae4c2164-6c10-4cf3-84a5-8080ac845dcc · inbound

Otter: A Multi-Modal Model with In-Context Instruction Tuning cites this paper.

Otter: A Multi-Modal Model with In-Context Instruction Tuning VideoChat: Chat-Centric Video Understanding

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:43:47.862703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T02:43:47.775691Z digest=sha256:c315bc864a256022decc3f3c97bf160a9d8de0d2573c3ddcc388ea08924282c9

Observation 15850b87-1bb4-4c0d-8e85-e048d6ba2f49 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:56:42.342701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:3322cac590ae0bef68e08bc2971f09239c8ed6874272d3b9712ce910a45298c2

Observation a365cf16-1b32-4160-b406-b9cbe99f8809 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models VideoChat: Chat-Centric Video Understanding

Reference 272

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:28:39.220206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:3d55c1979824fb41a7a1d5f8b16d01116a4659dc95526a2bffbbe43c2caba528

Observation 14ee9171-1cfa-45b2-95a4-f627cf2fd942 · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VideoChat: Chat-Centric Video Understanding

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:30:22.509068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:121f03635b774379e7d77ab25b010e96e5c36025c8426dde0e83935c869260e2

Observation 6abe26b8-437b-4edb-af3c-8c6a28a9c8bc · inbound

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension cites this paper.

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension VideoChat: Chat-Centric Video Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.879549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T16:59:50.495335Z digest=sha256:f57d7d276e1777c177c14fc065164c93f8f998a0ff9669df60632c707d85f9c0

Observation a811ed80-2b38-4a2e-9e15-e364ded8ccdb · inbound

A Survey of Hallucination in Large Foundation Models cites this paper.

A Survey of Hallucination in Large Foundation Models VideoChat: Chat-Centric Video Understanding

Reference 130

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:21:00.932415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-16T15:21:00.778049Z digest=sha256:b32b65b94f442a6f7112a498249bc7fbc6e9b49418bf7083c9c9c74e1aceef40

Observation 48b1f484-3639-4ead-a4b8-01f3bb01bddc · inbound

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration cites this paper.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration VideoChat: Chat-Centric Video Understanding

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.688069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:646d126a9c4f958801afcddad370c40bedbf445b7b337ac64d5a051db15249f0

Observation 59472c27-18c4-4b95-b968-f6b1b8535212 · inbound

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection cites this paper.

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection VideoChat: Chat-Centric Video Understanding

Reference 63

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T18:08:01.350653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-14T18:08:01.166072Z digest=sha256:14f158e4caddcf33bce7ee521e82e5feaf7e1608b9fd2500ab06fa39c46d58f7

Observation 884e9606-2235-40c7-b760-7076483c38d2 · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark VideoChat: Chat-Centric Video Understanding

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.099555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:3166837eefd6a7b730dd75d3619950b28683abb5cc4e38c53c8d3598540d22ea

Observation 42f7cdf9-0992-4f98-b24e-edcbcea51670 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks VideoChat: Chat-Centric Video Understanding

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.879549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:a98411188cad415e822e962fd5c7b893f17496859088eb102038a808ad2cf1e9

Observation 239ac708-0def-4865-a4f0-ecb9cda9ba29 · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction VideoChat: Chat-Centric Video Understanding

Reference 84

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T14:25:59.287768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:673d2f7cd03239ae87be43f38e9eb9624da6e75ff6355797fc463e9af7b6a9c3

Observation 1b876105-a280-431c-ad71-0b69f8f9fd21 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? VideoChat: Chat-Centric Video Understanding

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:46:16.711637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:78b345a5d48aa456d629e69c41afec3d5ebd269b229e2b662f9ff94aa2478a20

Observation a1e21a52-fc73-4886-9710-2f2d0bce3ea3 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites VideoChat: Chat-Centric Video Understanding

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.879549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:d7429b93569d4fcb5496706cb9008bb995d4ff3f97cbf55082e49ca4d8888ce2

Observation fac87720-fb17-40cd-bd06-1a749f8a1017 · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning VideoChat: Chat-Centric Video Understanding

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:21:57.962597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:6d69bd8dd4cb6de6b17bbda41e3af210bde9f123d017696618d912bfe4e4370e

Observation a73ef674-3145-4c45-a6a8-0b1191e21e94 · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:55:26.590413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:e445be22501cef767b167ee645187ff4fab3e9d2c558ec552ea06a1a71a079c0

Observation bd69b2c4-4d4e-44ed-91b8-551ffb3f7747 · inbound

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives cites this paper.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives VideoChat: Chat-Centric Video Understanding

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T00:23:39.617174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:e391edbbd071dedb058bf90f3a2fbbf8f10c6629954aa7d94ecd300c8790929d

Observation 4015d4dd-a620-425c-920c-7210c1415dc2 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output VideoChat: Chat-Centric Video Understanding

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-17T10:46:28.625942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:24a5b0ec73dd22d6c3a0a988dcdc88d5a3cafa066bf59c112499b662ffdbb1f4

Observation 00f85e25-659b-4493-a6d5-69c64df88929 · inbound

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models cites this paper.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models VideoChat: Chat-Centric Video Understanding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.879549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:0da01f0973114548c99578a568202f21356a0f4eca30e64f88cd8e1043f0d561

Observation a6cc0f65-2612-433c-aac3-aaeb76eeb852 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models VideoChat: Chat-Centric Video Understanding

Reference 222

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:20:36.423910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:4dba849d1dbc076f4fb9f9162d0f6eed5a3bd88be81d88e1eeebfff19227e059

Observation 0652cd27-6b42-445f-8066-cb5879896a28 · inbound

CogVLM2: Visual Language Models for Image and Video Understanding cites this paper.

CogVLM2: Visual Language Models for Image and Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.722354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:edc496a1797c0cf8f797aa1170e61c3aa2bd89e8c3b6f858aaf7c7d638711212

Observation 945090e3-98eb-48ff-a09f-1ec94a9f1394 · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data VideoChat: Chat-Centric Video Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.879549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:f6a8de3bea839f69c24cd95b80a40f4d880848051f1db4657e86c7f49887040e

Observation 79b5e6c1-457f-4a0e-a0ab-f3722dbca13d · inbound

VidHal: Benchmarking Temporal Hallucinations in Vision LLMs cites this paper.

VidHal: Benchmarking Temporal Hallucinations in Vision LLMs VideoChat: Chat-Centric Video Understanding

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-23T16:58:12.013422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T16:57:12.821916Z digest=sha256:295dd55aee35edcd32ad0303fd7594daa3a88c95af9b152b680a769cd8d8e1c0

Observation ac1221c8-b61b-4081-8dd9-81e78e4a7685 · inbound

TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos cites this paper.

TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos VideoChat: Chat-Centric Video Understanding

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:12:43.943872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T08:08:01.889675Z digest=sha256:644ffc8356a97741def47be0ff86b7b898ba3f925b061dc361b9dc6096888044

Observation 56d1e175-a9ee-43b8-95c6-8e79068f73e4 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling VideoChat: Chat-Centric Video Understanding

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.879549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:9b391d8b09cee97d635cf2e3d6e06dc7517b289f48bc21ced9642812df734e4c

Observation 419570b1-1595-48dd-9558-1b75b0c809b7 · inbound

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks cites this paper.

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks VideoChat: Chat-Centric Video Understanding

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:51:36.324238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T19:51:36.137985Z digest=sha256:0e89f00e080a1455cfa28fa87563db67e937566deed56fe257c7d520f0b3072a

Observation 729ca95e-ca18-4b1a-b029-2b6b82e69251 · inbound

Friends-MMC: A Dataset for Multi-modal Multi-party Conversation Understanding cites this paper.

Friends-MMC: A Dataset for Multi-modal Multi-party Conversation Understanding VideoChat: Chat-Centric Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:47.548723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:41:47.548723Z digest=sha256:189968f37fa7bdb5d95fe3626eec0dc29fbfb1030e15a82d482b58f4face8c60

Observation 8d0bae63-f0bf-4c94-845d-8e3eb0568304 · inbound

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks cites this paper.

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks VideoChat: Chat-Centric Video Understanding

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:05:29.173667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T07:05:08.716223Z digest=sha256:3cbed0296bce89c78be07f3361dc7cd2089057fe9bed925e756e2e2a72a568ef

Observation bc44888b-8097-4936-ade0-db1a18d98b24 · inbound

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment cites this paper.

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment VideoChat: Chat-Centric Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T00:47:12.242711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:47:12.242711Z digest=sha256:8491d49cc71039df40f4bbf5a66f28eac03d00c226e047de28e1387bfcbd006b

Observation ab47944e-5706-4a95-8f8d-40b98bedb0f2 · inbound

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model cites this paper.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model VideoChat: Chat-Centric Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.533205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.533205Z digest=sha256:2586c2cf4432761a0edf42e23f1642da7d1ebdad2ef4706ea223b6957aabdcb3

Observation 349855bf-91db-4b2c-b76a-3e2ed6b978dd · inbound

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval cites this paper.

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval VideoChat: Chat-Centric Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:00.849222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:00.849222Z digest=sha256:0bcca3abade478e0c85b2d65ce32a3c46b71cce6052352813662b4940f1d0b60

Observation 7c3bf99c-f3e6-405f-a78d-f220407e7788 · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling VideoChat: Chat-Centric Video Understanding

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.438898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:4dc4b92202502bf65a568f583fd6776b96de2f97216ee67df21a632a36d612a0

Observation 225d188f-b04d-49e7-b30b-e3790e154e73 · inbound

Online Video Understanding: OVBench and VideoChat-Online cites this paper.

Online Video Understanding: OVBench and VideoChat-Online VideoChat: Chat-Centric Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:40.045508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:40.045508Z digest=sha256:b5cbdec5cf3b52bf1f192229d0a8c7ee37624d0ca80c860fb273203ce6c26038

Observation 5ce8e1bb-2c42-4e4a-9e91-a9918787bddf · inbound

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs cites this paper.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs VideoChat: Chat-Centric Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.372394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.372394Z digest=sha256:2a7124e0a63738d43537f9039400a66f56f3ae79f7a9777abc8ef84259457a9e

Observation f9447c53-292f-4a83-a700-71de2682f56d · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications VideoChat: Chat-Centric Video Understanding

Reference 146

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.456281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.456281Z digest=sha256:8b2ccaa23f75ce07688e85921d4b81a543ce72d36cb9b86b4733f5b7243f82d9

Observation 52152dbd-8747-428f-ad2a-f626394b2450 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models VideoChat: Chat-Centric Video Understanding

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-23T05:45:28.365231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:6efafbee88a9dc12ae02aebedb3e5ea004cf6b954f7fba9f254e1efdb9e6816e

Observation 0218fc29-b73d-4169-9e62-bce6235bc050 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos VideoChat: Chat-Centric Video Understanding

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:39:22.552419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:0e8e6706695b1b9067809c4781caa841be6dfe779bd52ad8b6fa1d66e19ef74c

Observation 725677b8-7fb1-41e6-9608-fd9506f47da7 · inbound

H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving cites this paper.

H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:11.565112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:11.565112Z digest=sha256:97103919895c91567254a50274e58a9fc51a7dac8cb21e90a14da5fbe0343625

Observation 1730d6a1-7a2d-4750-9cbe-f79c1e5e14d0 · inbound

LongViTU: Instruction Tuning for Long-Form Video Understanding cites this paper.

LongViTU: Instruction Tuning for Long-Form Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:23:57.923377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:23:57.923377Z digest=sha256:7b3b2d77721046143bf0d60663c684bda03934b6ae70ee7f3306fbb18e5f6c8e

Observation 08b776a3-e1f5-41b6-88f4-15a3c3c6c2a2 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:02:37.497349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:9f214c046c19e2edfa738571faf66a35fc565a553649b4b8ef49ae582e74c979

Observation 2c5330c3-8e67-4e3d-a3e0-c4252774ad4f · inbound

VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning cites this paper.

VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning VideoChat: Chat-Centric Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:13.259565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:13.259565Z digest=sha256:1f9f6ee4ae3844c857808f5c6a9a82255efdceb472195674a6a3abab9da8c0d4

Observation 2e33504e-cdb5-4ea3-9099-cc2a3bfd6e84 · inbound

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness cites this paper.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness VideoChat: Chat-Centric Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.040489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.040489Z digest=sha256:b713916bd64e75f562b08a52a3e487bb2e95f4f7260d45b9bddb27721d627f8f

Observation 567c6c6b-9aa7-4829-8019-bd2f95cc1b56 · inbound

LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking cites this paper.

LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking VideoChat: Chat-Centric Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:11.725520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:34:11.725520Z digest=sha256:329be0a97f6a07facfcbc2601e235c83b03d368e19a2166ec8ebcd576df6e36b

Observation 6cbe9747-b6f5-453c-befc-6c2420fc9e32 · inbound

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis cites this paper.

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis VideoChat: Chat-Centric Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:03:12.859584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:03:12.859584Z digest=sha256:6fa26aeb7870cc86bde737d39fabd547bfe11c1251df599eaabe171a591af869

Observation 34fe1b59-3cca-4a37-b96a-1ea327675c95 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.879549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:8644d480669502c59a2545a889c84ed04658b96efd2c71118a7a8ef18369fc8e

Observation 74613560-8f9d-42b4-8ac0-78b3b5004867 · inbound

ReasVQA: Advancing VideoQA with Imperfect Reasoning Process cites this paper.

ReasVQA: Advancing VideoQA with Imperfect Reasoning Process VideoChat: Chat-Centric Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T15:56:37.565527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:56:37.565527Z digest=sha256:cdb97a72624d97889a5032ad1d5bda0c05ad2ca058a8c6e54a35b3c7cd95906a

Observation a4755349-b04c-4fda-8bd6-7c526c9e1545 · inbound

Temporal Preference Optimization for Long-Form Video Understanding cites this paper.

Temporal Preference Optimization for Long-Form Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.172948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.172948Z digest=sha256:7f9ac1572ddfea9585abf69db097544fe47a3524eff4982919ef84fd4c3aa8ae

Observation 6b1dd5db-994e-46ee-a79a-44ee7d9636a0 · inbound

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding cites this paper.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.240225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.240225Z digest=sha256:4b8602671394e1c3f8e40686c07fcd858716b22d39bcd12a114e8a4636ae8b67

Observation 39898d34-2355-4ed0-82d8-51ca383af8ed · inbound

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler cites this paper.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler VideoChat: Chat-Centric Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.605226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.605226Z digest=sha256:cd488ad4f83d910ce109fb688039eccf3fd816b08cc4cbb2a5b9149b5e673983

Observation 7813284f-12ab-4ccb-bca4-624e0e947a5e · inbound

Understanding Long Videos via LLM-Powered Entity Relation Graphs cites this paper.

Understanding Long Videos via LLM-Powered Entity Relation Graphs VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T13:53:26.184775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:53:26.184775Z digest=sha256:d445178da745dce4e1e6498ef05c919c2c7277788edcd48795216a201abd8fd2

Observation a29aa8fa-f608-47fe-bd84-ffa7172425ed · inbound

$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation cites this paper.

$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation VideoChat: Chat-Centric Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T21:21:45.654533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:21:45.654533Z digest=sha256:9986ef3c836abc2ee06f99fd622afc5e3283930fde4b16e05533b7d80abc1bd2

Observation 58b706f8-00fe-41a4-8b9f-decae703afd1 · inbound

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective cites this paper.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective VideoChat: Chat-Centric Video Understanding

Reference 139

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.281075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.281075Z digest=sha256:db1caf1713bbca4257d2de90af92ca9a233f17a7fc49a4de3ca80f517199f185

Observation 9573671c-4511-4233-b5bb-cc2db817550c · inbound

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM cites this paper.

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM VideoChat: Chat-Centric Video Understanding

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T21:12:22.802431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:12:22.802431Z digest=sha256:8381ed59aecc44dfb7d8f4e668c05be3e89df27e899aac5b7e856f605d461c07

Observation 7ff34654-d8bb-444b-98b8-ecc56fb6f257 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey VideoChat: Chat-Centric Video Understanding

Reference 198

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:18:53.359724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:186f10c78ae5aca5f380cb3ebfda27e3eab6c29bf70e04007c382a434c63622b

Observation 61c4b2fd-55ba-44d4-aff9-e9fe527c97fd · inbound

MusicInfuser: Making Video Diffusion Listen and Dance cites this paper.

MusicInfuser: Making Video Diffusion Listen and Dance VideoChat: Chat-Centric Video Understanding

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:32:15.645719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T23:28:51.767845Z digest=sha256:556e227c25d1a85b8e8ef6b5bd2f675150cd179f74226c1a61693f8719c30978

Observation 660889a1-6a46-4982-a916-8f1ef2e247ea · inbound

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning cites this paper.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning VideoChat: Chat-Centric Video Understanding

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.809487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:3f25c834a8f1e3aaf3958b40bafd0b7cdc55f73d2483bf661fecf1bd651557fc

Observation 96dde0f3-38d0-406b-b44d-5595cf56c437 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models VideoChat: Chat-Centric Video Understanding

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.879549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:2e81537aab0663f0059feeb513369b8ec6badaa46e98620a420c6e4d2db2c894

Observation 3f9b5974-c0ef-464f-93fd-7d8253ab8b18 · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoChat: Chat-Centric Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.546538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.546538Z digest=sha256:f107fe1e0afd4cf5950ced246e5c1bb72bb664374d98ce66ba813ffd5c4936f7

Observation 98ed9a7a-80ee-42db-9d19-2b2be5e63aed · inbound

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language cites this paper.

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language VideoChat: Chat-Centric Video Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:12.678804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:19:12.678804Z digest=sha256:5d6c2caa028310e3f23dd6f3fe3d5bd88c1cb6eddffd4d4d89a8d2542a69c063

Observation faaad85f-8620-4213-b570-1c8ea8c4ffc9 · inbound

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion cites this paper.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion VideoChat: Chat-Centric Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.563235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.563235Z digest=sha256:02abb9512f2150b6929224d5599c9a818de6efd914863c8ce79afdb5f2cadf73

Observation 8de95b1e-1fd5-40ea-b53d-1fe02be8f7fa · inbound

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models cites this paper.

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models VideoChat: Chat-Centric Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:08.311582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:24:08.311582Z digest=sha256:ad89272639627178356e5fb9ebd679c3d91a9d74cd3eaa23a16c6a3032ee25fb

Observation 982636d1-66c4-4956-ba0f-a16b661028ac · inbound

Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation cites this paper.

Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation VideoChat: Chat-Centric Video Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:05.900446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:18:05.900446Z digest=sha256:4fa999b9a47b383bad671480ae6f5ec6873277cd953bb1ac40bfa77d8f811088

Observation 055477c3-d512-49c2-84aa-4834b945b15b · inbound

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought cites this paper.

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought VideoChat: Chat-Centric Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:08.616938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:08.616938Z digest=sha256:56d8c1e3d255369b68a8aa96d1a9db00da34f18765e42a287109c7f48aaac1bf

Observation 57916962-adad-4b7f-9fd7-09b260c88146 · inbound

HuMoCon: Concept Discovery for Human Motion Understanding cites this paper.

HuMoCon: Concept Discovery for Human Motion Understanding VideoChat: Chat-Centric Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:13.144440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:13.144440Z digest=sha256:2cad820fa426d56c55974f0334d827d4dd37a4bca9210d21e3ddf928b2ebf47a

Observation 5bf2a2c0-36b0-4b24-abcb-c95a924b766b · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? VideoChat: Chat-Centric Video Understanding

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:40:55.999728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:06e394283a8a6f4aac80be02224acb698eb96cc5e1f07be5560164d3eb78bbe9

Observation 23b3972f-bd30-42bf-b519-bc4af4101549 · inbound

Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment cites this paper.

Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment VideoChat: Chat-Centric Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:52.839627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:52.839627Z digest=sha256:d216a0187c2df07e8892dfd6750544b3cf1b2becd883dbe8417ea287c28b3b58

Observation cfb42be7-2600-43d1-bc3f-806577ce11ec · inbound

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos cites this paper.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos VideoChat: Chat-Centric Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.512570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:46.512570Z digest=sha256:8c0fbfc7698e2f4712a1b3f8398b2379d9beee5ca356a392c8cfd1a5ba9fa4b0

Observation 3a7c88d5-139f-40ae-9010-2a4162c88189 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence VideoChat: Chat-Centric Video Understanding

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:34:36.917365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:a735a78c2f5aae433a7df9690f78a28f6d38093d9a3f8ea43d1ced51fbabdd05

Observation e02ec366-64f2-4304-8fe3-353a6c33a29a · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence VideoChat: Chat-Centric Video Understanding

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:00:51.411925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:00989a79dd2f52f9c6221244b9c515a8394fe1d36b091e0e256b82325f5787c5

Observation 60b66c9c-95ff-4c99-aa4f-806c7132fb93 · inbound

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders cites this paper.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:19.138328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:19.138328Z digest=sha256:b8728ecc471c52cc0193950b6d56fa7760a292308901431c3aa429af7a75e78b

Observation 21d7e694-12db-4d82-a923-8fecfdde71d9 · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models VideoChat: Chat-Centric Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:51.944782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:51.944782Z digest=sha256:f8cce70d563619886e0e4d1f9d854b13b18a598c064c11733c1b6ee9b8e72d94

Observation 7ddc9ed9-fd2d-4530-ba5f-289c8eae72c8 · inbound

VUDG: A Dataset for Video Understanding Domain Generalization cites this paper.

VUDG: A Dataset for Video Understanding Domain Generalization VideoChat: Chat-Centric Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.451229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.451229Z digest=sha256:e3b8b331ed9dcd714c74d7f3a4a254bb7d08441b7374dfea5fbe5d354f5610a9

Observation ac08881b-48a1-4659-a175-f4a53b64c0d8 · inbound

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering cites this paper.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering VideoChat: Chat-Centric Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:31.859767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:31.859767Z digest=sha256:19acf44c31c83c7d2f2616eda397dec9aae8b25676e02349cec3a541d4ac8c26

Observation f13594f2-885c-4c3f-9901-677686589a63 · inbound

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model cites this paper.

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model VideoChat: Chat-Centric Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:26:37.478169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:26:37.478169Z digest=sha256:733b6c988d20649329ba6265e5873aafde21adb3d3c21a0b5b80b817576ef269

Observation db66c58f-dd38-4da6-85ac-4bbd1d2bfdf4 · inbound

Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues cites this paper.

Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues VideoChat: Chat-Centric Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:21.080473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:59:21.080473Z digest=sha256:dcdcab58ee471a815f8a9c361127636d8dcbc05d5804613f60c3a41a40b8b9b7

Observation 096c6815-5245-45ce-baca-b8e6e4d6cf63 · inbound

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding cites this paper.

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:02.781775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:02.781775Z digest=sha256:f3447bbbeff68d4607831186903f23161ecee6b292d3eea9acacc314381af129

Observation fc6bae13-b3ab-420e-981c-09a913f748a7 · inbound

Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models cites this paper.

Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models VideoChat: Chat-Centric Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:30.708631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:30.708631Z digest=sha256:1de46ef5e090994d6b83fb0ee3def9d18de75ca9f2ebe69a8ae5dd74838d66ad

Observation eaedcbe4-f741-478f-a9e9-13c0ef7ab61c · inbound

Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review cites this paper.

Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review VideoChat: Chat-Centric Video Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:35.768607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:35.768607Z digest=sha256:25753b767ae76c5b901ea8083d460a0428ba2d7e3fcb95d864de0a74adb3bde2

Observation 360fda69-e337-43f1-8982-249ddb5c9bf0 · inbound

FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents cites this paper.

FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents VideoChat: Chat-Centric Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:46:58.704703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:46:58.704703Z digest=sha256:3e3af0d691aea840cc8b80e17ea85a2bfb02e0820afcd208b43b2a851a29dd0f

Observation f4f70289-08d2-46c8-ae70-8bcb2c2b402d · inbound

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency cites this paper.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency VideoChat: Chat-Centric Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:45.945951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:45.945951Z digest=sha256:42d4e8f4ac176aa070c88a804ec363defb185b5847e517e3fbc1f8b51fa712c3

Observation 5affa91c-1097-4287-9971-fc2a790710a4 · inbound

Vid-SME: Membership Inference Attacks against Large Video Understanding Models cites this paper.

Vid-SME: Membership Inference Attacks against Large Video Understanding Models VideoChat: Chat-Centric Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:01.880477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:01.880477Z digest=sha256:f9cc6b3b3d14980619b7f4de81c721d9ddaf642153f62b465acfcf56b6febd28

Observation fb72792c-376e-40b7-80b9-4ed91c6044b2 · inbound

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding cites this paper.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.327297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.327297Z digest=sha256:55cacea76856210456ada2603db268d1f12c4cdc3b86bc3cd2dc224da5a0f873

Observation b5bc3b82-eb86-4917-8094-c69ed13df073 · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs VideoChat: Chat-Centric Video Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:52.131174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:52.131174Z digest=sha256:313078b4361874e9b452c7f250b9a27d575758e84382350731b478172a5e86ed

Observation 2fdc5dd3-0d57-469d-bf4a-92a2918ac617 · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? VideoChat: Chat-Centric Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:09.058999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:09.058999Z digest=sha256:f4088391787c82002d8d96fd76ada86cb5a9034a0cc9d1ba40132a107d60fb32

Observation c9dc851b-57e1-45f1-978b-2206b840f82e · inbound

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning cites this paper.

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning VideoChat: Chat-Centric Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:50.440440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:50.440440Z digest=sha256:d3932474c77312b45df608f635133eabccd6e443351b674d8fc519a93b45fd97

Observation f7621811-7c51-449c-ae3c-e6cb18deddc6 · inbound

VideoMolmo: Spatio-Temporal Grounding Meets Pointing cites this paper.

VideoMolmo: Spatio-Temporal Grounding Meets Pointing VideoChat: Chat-Centric Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:27.739443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:27.739443Z digest=sha256:6a394b9cc13f2fdbad3c0f26b319f69b1dd2e9157ea50f996362c12b530ba074

Observation d7856dab-999d-4b83-a02d-48a4ba13866f · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing VideoChat: Chat-Centric Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.645520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.645520Z digest=sha256:71f98ac10c34c7cc56d7a006ae8b714c87d9b9d4cda6c8958e941da9d40bed0b

Observation 1796d316-0616-4672-8205-6ed41d4acc80 · inbound

Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models cites this paper.

Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models VideoChat: Chat-Centric Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:20:38.170335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:20:38.170335Z digest=sha256:8de39ef555391bb0b736a470348c20745e75b55d1bc7d3d5756b059df5245e27

Observation 3f837d44-eabf-4ebd-b8c0-9ed41b93dcd8 · inbound

Proactive Assistant Dialogue Generation from Streaming Egocentric Videos cites this paper.

Proactive Assistant Dialogue Generation from Streaming Egocentric Videos VideoChat: Chat-Centric Video Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:44.040534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:44.040534Z digest=sha256:d83011170f33156b2fd56c2cfd2e4981deb3a372993350db6f8ac593484516ec

Observation d0161f92-a0c5-4a8c-a35f-243b89c77a80 · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding VideoChat: Chat-Centric Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.898389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.898389Z digest=sha256:55eeeb1adc2a0d36be5efe6c4a5b8062a3707e3431060ed286b1e51b2b71a680

Observation 20fd0780-27b5-4a90-872a-f5ccf4089a5d · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoChat: Chat-Centric Video Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.397026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.397026Z digest=sha256:016d75ba66467cfa1ef2768d2920e8c51a501b97827eb094a19b713d768f4dda

Observation c19432be-2e3c-44ab-82f5-9ad99d1453ae · inbound

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding cites this paper.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.832358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.832358Z digest=sha256:4cfe1c776aada68a74b65c95caa1fe3044c6a5142cbb2d9d7db2a693e861133d

Observation f4bd1b8c-c516-46d8-a784-fb266f7e1a4f · inbound

EgoM2P: Egocentric Multimodal Multitask Pretraining cites this paper.

EgoM2P: Egocentric Multimodal Multitask Pretraining VideoChat: Chat-Centric Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.478451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.478451Z digest=sha256:7b073233e19504df9de328195d3a5d16a37a82cf4a623d82c3399c5da5df7550

Observation 43752052-9380-4f73-bef0-41ec3dc11c36 · inbound

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning cites this paper.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning VideoChat: Chat-Centric Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.979169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.979169Z digest=sha256:aad19488a8e8b4debbcd57899c8509befc648f010729438161da69cb5ba4e8a2

Observation dafab7fb-eddf-4b09-a8f5-610c60bdfd5c · inbound

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision cites this paper.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision VideoChat: Chat-Centric Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:22.296324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:22.296324Z digest=sha256:2437e61a2ba22967307b6092ba5431cb99ca54fec81181040203faa48dc3c12a

Observation 0bca8184-4516-4ef8-9cd2-32287af706f8 · inbound

Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos cites this paper.

Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos VideoChat: Chat-Centric Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:43.345468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:43.345468Z digest=sha256:573b703b7f4ba7c0596359be24a8cd1254fa6daac532de6296e1c03589ca3294

Observation 382a5726-de98-4256-b7eb-0cd9cc1efef9 · inbound

SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models cites this paper.

SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models VideoChat: Chat-Centric Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:51.274375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:51.274375Z digest=sha256:43c79ef1f35693ffff07c393ce1db0a83d9a3d31cab38743a9a1e10564ec6972

Observation bd4ad7ce-5871-44c9-9b2f-0bc008ee470e · inbound

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model cites this paper.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model VideoChat: Chat-Centric Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.316833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.316833Z digest=sha256:a36bb692b0941a60955ce3cb1dc4379b5519883750ec994a7cae4c8ad983e308

Observation d0abc37a-0e0c-421d-b03e-a92d76fd7b81 · inbound

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning cites this paper.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning VideoChat: Chat-Centric Video Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.713603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.713603Z digest=sha256:081c609e1e74803be00975ad810e03283c1a4c5d8b3c9c4cee83b69d65b35697

Observation 5ff9199b-effe-478f-8564-e0160dc4a7d3 · inbound

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations cites this paper.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations VideoChat: Chat-Centric Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.248369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.248369Z digest=sha256:32f5ded433d933bf3f14be9a2890ca5fd4ed639302b4031fb67f846f87f6d0f3

Observation cc5d28cb-ac01-4903-b3b0-4cb656dca795 · inbound

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering cites this paper.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering VideoChat: Chat-Centric Video Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:29:29.218410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:29:29.218410Z digest=sha256:882954d1ff3270009d595a85a98e2f077255834b5a1fe55c4327726a67278e4b