Pith. sign in

Paper Citation Record · LEDGER

ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 89 inbound Pith citation observations for arXiv:2406.04325.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.04325 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 89 of 89 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:36.459440Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.677235Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 69acf894-d236-4b03-910a-a38fbbe25c9d · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.521910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:7215124aa708b1c2e431655fa286dab2efdd4dc1c4de3cc6d0f5fa1054cac8d6

Observation 9c1b1b1a-f9be-4f69-9ab9-2203b31acd9e · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.714646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:8b261b88ed18f2148f9f315be130c3747e9c2f0e15a17b5907bd50c74f1a1709

Observation 13c3cee9-ad9c-4bf5-9fc0-a5837391ad38 · inbound

LLaVA-OneVision: Easy Visual Task Transfer cites this paper.

LLaVA-OneVision: Easy Visual Task Transfer ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:23:49.613523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:26b79ed4539ed9cafd94a7138484ba2b7acc093586dde0b67611159193906d16

Observation 74c89a1a-f1f0-4b83-971d-a272a19e5b5d · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 203

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.347366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:d15a3bc3da03f6cb31c58527944ba16771799ef3ea369091e36b5df0c70311e1

Observation 37de3bd8-d552-455b-8ae0-d85ac3b9b3dd · inbound

LongVILA: Scaling Long-Context Visual Language Models for Long Videos cites this paper.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.468836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:44809d19f1b4139aecb20768a0b33c2d495393fed558f1b429d6cce64610193a

Observation adafa0e3-6f81-4449-ad51-6dc21c67c500 · inbound

CogVLM2: Visual Language Models for Image and Video Understanding cites this paper.

CogVLM2: Visual Language Models for Image and Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:27.726864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:5964a60d121a9ec3e45647d41250fa875da5536d4f837ee2a7af2d5da9c64d55

Observation 39825b21-8a01-4bc2-bf4e-a2fbf81ed070 · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 160

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:32.495104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:df8435c568c56e9bf5cebed01766811315a02a12d9ce111a7d5ce4da10982699

Observation b2e0239e-d167-4485-84a0-1d2e663f2c9b · inbound

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction cites this paper.

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:12:14.756329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T12:12:14.613620Z digest=sha256:a4632a354eb707fd093f9bf115b52cfba52882dd4407cbca5149dfb3322cfa70

Observation eb74b757-d340-4bc8-96d8-8ae242490ded · inbound

SimTube: Generating Simulated Video Comments through Multimodal AI and User Personas cites this paper.

SimTube: Generating Simulated Video Comments through Multimodal AI and User Personas ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T20:34:16.148369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:34:16.148369Z digest=sha256:2a39c82144935cba2efc21e92b3390adb774e4885da21a9e2623b04904a795cf

Observation 81927198-2aef-4d3b-ba16-d860bcfcb006 · inbound

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation cites this paper.

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T16:42:55.681332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:42:55.681332Z digest=sha256:7903c6040855fe2ecb27e8c952d1a27d9ed9ac01929e485264a31dbdd2862455

Observation 411d20e4-f21d-4d13-a7f7-a189a9d5d80d · inbound

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection cites this paper.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.507092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.507092Z digest=sha256:495c1cb942c39e5228074bd0c6b65e9fd528a19b069be307487db4174a8e7409

Observation d2ae6eec-bd68-44f2-ba7a-00b1f27df4a0 · inbound

VideoOrion: Tokenizing Object Dynamics in Videos cites this paper.

VideoOrion: Tokenizing Object Dynamics in Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.353218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.353218Z digest=sha256:1672823c34dbd438f0d7bfb28d9cdf7661f932c886c9f5a80e8cea155f9d7459

Observation d115cd80-01c9-43a4-be4d-04a3b491d210 · inbound

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis cites this paper.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.016872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.016872Z digest=sha256:95ea697747ec389bdb6624f71126c004c17d02992897265d8442c8d00924f70d

Observation ac0f485a-7c69-4684-b804-b43b6a35248c · inbound

Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding cites this paper.

Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T12:47:21.195627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:47:21.195627Z digest=sha256:a060782cf5ae81fc73ab27a52a95d98ba4131a03932c7ef003d248bfc5ba9e24

Observation 130a81ae-75d5-47e6-afa2-7b1689c72968 · inbound

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing cites this paper.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.819643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.819643Z digest=sha256:53366e4870573528f3745d7e3e7d95b8993d473e4af90f78805f2807027998d0

Observation c023be49-699a-45cb-a054-d045b7b2efcf · inbound

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos cites this paper.

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T05:53:34.377247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:53:34.377247Z digest=sha256:53e90e183ac439a1ad11f1cc764ffdb376bd365b6d2fbaf0207037b465e8f9e5

Observation 039c3b39-f01a-4617-964a-34ba3e6b9ec7 · inbound

VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation cites this paper.

VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:56:01.190716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:56:01.190716Z digest=sha256:41b5b2e2ffd7b63ff10c3f7071c1d3ba808ce5e47946919448b0ee78560a398f

Observation 721f1b58-4bc5-4a1c-876f-64120cd3f48a · inbound

Progress-Aware Video Frame Captioning cites this paper.

Progress-Aware Video Frame Captioning ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:58.299593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:58.299593Z digest=sha256:9e50c00f45476fad06d66b578c3da7c313c0aa139c933d43b4b5ec2dbe316d10

Observation 111ec42d-5624-46bc-9f4d-494adccba1d7 · inbound

HunyuanVideo: A Systematic Framework For Large Video Generative Models cites this paper.

HunyuanVideo: A Systematic Framework For Large Video Generative Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.549699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:4664d295e5633db28b3af298ef25ce2f40905c1d496d46c9b19c6a7a7778781a

Observation ddbd9cb0-6eba-4ab4-8f0e-b6cb83285c48 · inbound

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios cites this paper.

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:00.165171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:00.165171Z digest=sha256:e8843ef1ff9df7e82364b12b8658f88d5090930a662ed6df14309e11ef3f86c2

Observation 885d6d3d-6697-418d-8e0f-df75883c3c6d · inbound

LinVT: Empower Your Image-level Large Language Model to Understand Videos cites this paper.

LinVT: Empower Your Image-level Large Language Model to Understand Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:15.963977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:15.963977Z digest=sha256:18f4ed82f94b192db2b846f860e122d2cc3940a568ae3581ea98fc6e34b294fd

Observation 235ff406-5120-49e3-9bd1-921fedb6817a · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.967610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:7e94eb09247943a544da19f51fdb355cf382f04b4ef0b8b578f53908bd8c0163

Observation 514a8eda-260c-4fda-a60e-170b9fd0e581 · inbound

Exploring What Why and How: A Multifaceted Benchmark for Causation Understanding of Video Anomaly cites this paper.

Exploring What Why and How: A Multifaceted Benchmark for Causation Understanding of Video Anomaly ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:32.726429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:09:32.726429Z digest=sha256:91dc643dbb493317c638d75f94e6a3ef43fd0d79c925aebf4f2e3ee492d164ed

Observation 4ebfefe2-7ddf-48c5-b74f-a80380af84ad · inbound

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption cites this paper.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.897243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.897243Z digest=sha256:d75d5e7711a2d8fe1004a6283c06af596bf640f80ebc19dfd728ce23bd42569d

Observation a2949096-c0f9-4c23-990b-a3306ac34c6f · inbound

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM cites this paper.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.705603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.705603Z digest=sha256:84a14ebd4496b578f4f8c77ee9909a7f520300209baee8b6993466616f959a13

Observation 6f00bb27-f81e-463f-84e7-8ec8086f4590 · inbound

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions cites this paper.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.030029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.030029Z digest=sha256:9896d70e4d1608851b17d36fb31dfd388a8c83a2f1eaf2a77f1694fda10e47fc

Observation fd22d433-36fd-4676-9c06-eae22651bc0c · inbound

TimeRefine: Temporal Grounding with Time Refining Video LLM cites this paper.

TimeRefine: Temporal Grounding with Time Refining Video LLM ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:35.748209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:35.748209Z digest=sha256:972450486314964683a647c3af0cc2c99bedecda6cb2f4b83dc6f90c3dfa83ae

Observation e60fdd45-4d6b-417e-ba35-1b90bb74b59a · inbound

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models cites this paper.

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:42.304316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:42.304316Z digest=sha256:bf271751e30235dcaf90d43ca86a77e29faf880875e7660b22bc662d16824912

Observation 17a07dd3-0067-4586-a54f-f922941b0114 · inbound

Can video generation replace cinematographers? Research on the cinematic language of generated video cites this paper.

Can video generation replace cinematographers? Research on the cinematic language of generated video ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.292949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.292949Z digest=sha256:5fb60ad5bd67eac28cb986147a84b0baa5b02f8106aa52410f5322c17c7f1341

Observation 7d254a1c-332f-4cae-b067-0cf07920621b · inbound

ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries cites this paper.

ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:48.029586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:48.029586Z digest=sha256:1c6ffaae54549af6a23c0d6220b55aff67a604fc35c93e587879796e16aec073

Observation e07e8417-78d9-4864-b5c2-a8077e335d3b · inbound

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering cites this paper.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.004298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.004298Z digest=sha256:58a6afc97fee8cac69a566c9dac3f7370df7fc5d045460d2866c15eee7a9c0fe

Observation b28a0493-06ed-4eaa-bb9d-23cb296d62b9 · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:17.079003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:17.079003Z digest=sha256:e2950e79e6c77cf112f4519b9b912bb0152b10710c007b2a0a14f65f97a5f36c

Observation 88c47207-6c7b-49ab-bc00-ebfba7a9d09e · inbound

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback cites this paper.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.947500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.947500Z digest=sha256:d178b554ab4c70ee316e213a3ef4f7085826ee183d028d7db86460cafad3385a

Observation 7e65a34e-54fd-4d6b-99d1-356c8031ba8a · inbound

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks cites this paper.

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:05:29.181506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-23T07:05:08.716223Z digest=sha256:98b5767647191a48474492bdf7971550de0c8a5b2b2e9b712ad662d80979eb4a

Observation e75adb01-ee7a-49bd-b510-13c046284058 · inbound

Enhancing Audiovisual Speech Recognition through Bifocal Preference Optimization cites this paper.

Enhancing Audiovisual Speech Recognition through Bifocal Preference Optimization ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T01:05:25.154177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T01:05:25.154177Z digest=sha256:8a1f2985e56e1e7ec5be6d4ae25bfa822998d13d19d8df52bbbe958936d24315

Observation 6c4c4aab-e7c2-4666-a51d-e9e39f8c8b7a · inbound

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment cites this paper.

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T00:47:12.087492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:47:12.087492Z digest=sha256:8111986d29497f1058e59fb2a56b13d4ba20815fe37acda7667cfbc6ca8dd9a7

Observation 29bee933-1e99-43d5-ba38-f71ccc284572 · inbound

Online Video Understanding: OVBench and VideoChat-Online cites this paper.

Online Video Understanding: OVBench and VideoChat-Online ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:39.830271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:39.830271Z digest=sha256:cfe2589ee2b8d4309336d8aa018e65699203922a0d5ee7d0edd3a241d18d0036

Observation e551242d-40a3-4331-bc06-41df8990ed44 · inbound

LongViTU: Instruction Tuning for Long-Form Video Understanding cites this paper.

LongViTU: Instruction Tuning for Long-Form Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:23:57.849584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:23:57.849584Z digest=sha256:f7a39b45904104d56d4c1b69f1747c164e48a96d6de82cfb643f7eb9a9034ec2

Observation 55563787-0898-4098-923d-ce3997e27789 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.376961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:3e8d1f040a591b87ba1b7c957660b08f5e6b31fad0c8ec4b154f7cd1edf294b8

Observation 7ade2e87-b39e-4fb9-a3b9-2a3bacd370d0 · inbound

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness cites this paper.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.918678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.918678Z digest=sha256:0bf030626df925603e09d7928603fa44974879abc0e9a752f649f27287a4e379

Observation f76dc3b8-f695-4eaf-9c72-ed9ca8af8034 · inbound

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model cites this paper.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.016166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.016166Z digest=sha256:2919d89ddd459cbc4dd7f66c628e29703084fd8f84405e43d1266d30d43f78a6

Observation cfd511f4-7199-41c0-8ae9-99e0872b6a71 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.849967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:c41b5eca7c084faba72ae5669adb7e47d614333cb5d8b1513cf9671aab710c0f

Observation faf5498b-79b0-4d6f-b1cc-31a0ebcfeb59 · inbound

Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning cites this paper.

Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:55.381407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:55.381407Z digest=sha256:f906f89120ba63f6dbd2ddcb718768d9d82f6a86f57e379c058d5909b463304d

Observation 431f0539-0b9d-447c-9712-5c0ea0be68f2 · inbound

Temporal Preference Optimization for Long-Form Video Understanding cites this paper.

Temporal Preference Optimization for Long-Form Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.082895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.082895Z digest=sha256:b0b7176c7c30411807fdbffa292ba81f020fe0717dc8dd9c1684dbb16b471bbf

Observation 9bc1d275-3d71-404b-9f56-69b217851a25 · inbound

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding cites this paper.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.172325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.172325Z digest=sha256:6ed07fd7a2ed4b6bc55c39b483d15dc564f687e326b3651ef702e3b218f99eb9

Observation d0610ac8-318e-447b-a2a3-8986869d718d · inbound

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler cites this paper.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.464986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.464986Z digest=sha256:a60c8ea8bebe2885da94b2dfef207054132e32700b003598b5a33041d4a542b7

Observation d94dd3c6-0fe4-4d42-9c67-1d61c89559f3 · inbound

$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation cites this paper.

$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T21:21:45.575570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:21:45.575570Z digest=sha256:a99b25df8b43e6a6dd77f54a012b4af51e489fc3a108397385a261df173abcd1

Observation 84909031-5f6e-4f8d-8289-8991b2a172e7 · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.062734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.062734Z digest=sha256:d0ba52e4662bec93364d481147e7ef0445b7d9c749d097ac9e04ae9b31cc591d

Observation c48653a7-17e7-4649-b1f8-6273adf43fdc · inbound

VideoRoPE: What Makes for Good Video Rotary Position Embedding? cites this paper.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.097848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.097848Z digest=sha256:ea2a75f5645413965af997418aca2d044758334429a597ae22cfd7bf2afcfe0a

Observation 299b4d8f-4f58-4ea2-b784-d23ec64689c3 · inbound

CoS: Chain-of-Shot Prompting for Long Video Understanding cites this paper.

CoS: Chain-of-Shot Prompting for Long Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.900888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.900888Z digest=sha256:485ca4bde40cd2901ea5c1dad99cb5306b921e469572e07ce2c2e094974c3ffb

Observation 6a3ad532-9edd-4f91-a7da-72196dccc641 · inbound

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering cites this paper.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.138227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.138227Z digest=sha256:09bb4d33995a38e5f765f89d9dfbc9925fcd4678e911aa1758b6eafa32e9cd7b

Observation 8f907f7d-2050-48be-abc9-27c2a2d25941 · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.679159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.679159Z digest=sha256:5947156ef23db6ec36bb5dc448e55b62a87f2ea029af36247b4091ec5c8d1d9a

Observation c44d3654-2579-4e20-b1e4-fb6f56135dc3 · inbound

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding cites this paper.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.459440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.459440Z digest=sha256:6e08d61a40cac83c9a135860f143264d7009483755efe84f05578e20c684e459

Observation 8a3d83d8-9cee-4397-a0ce-37873dbe3792 · inbound

Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis cites this paper.

Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:52:25.260944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:52:25.260944Z digest=sha256:0f2370484e14637a72a96f3f585267c91fec96c9a2a8212a36388314931c33b2

Observation 1efd48a1-dfb1-4daf-b9c6-e74ab80eb44e · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.415013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.415013Z digest=sha256:9134a5f0f69df92f68d9d03c624670e88f45a7e6cf5432e7d07c19672c69fb63

Observation e9cba1c9-195a-4afb-93b7-0eb6c5e9d4c7 · inbound

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale cites this paper.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.823292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.823292Z digest=sha256:89c021933f2561572264357d3058977885fcf54b4d5819adb3923c6eeb00c469

Observation e182584c-e789-40c6-b5c7-1d808da72705 · inbound

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding cites this paper.

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:21.372047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:21.372047Z digest=sha256:1068a1aad0a20a6d016fa8ebcd0e12acac6017b19edad3f59dd78bfdde4d8dc9

Observation a610f35c-47f6-427a-bd63-3ff7d1c30a28 · inbound

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension cites this paper.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.440183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.440183Z digest=sha256:b2a70c76a5927e9ba0b09a85482795e5c7e757b5a441f3c2f8918be9cbf0c198

Observation c7dfa941-13be-4f45-b591-a253d06c7c01 · inbound

VEU-Bench: Towards Comprehensive Understanding of Video Editing cites this paper.

VEU-Bench: Towards Comprehensive Understanding of Video Editing ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.737303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.737303Z digest=sha256:feeae0427659add3bc4b110426cea7e571ac63d889bd2576c02103cadc238107

Observation 329b37c3-0136-49cf-a6d1-28dd703339eb · inbound

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding cites this paper.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:35.458263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:35.458263Z digest=sha256:97779506bb70b78938625950708ebf33cbcec0cec39f8211927095ba52ffae67

Observation b28086b4-d914-4043-8d35-b1de08144eeb · inbound

HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation cites this paper.

HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:02:11.077076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:02:11.077076Z digest=sha256:8acc67404afc1bf2d4a557f5edc313d46e7507aa41895be4767dad342e1e97c4

Observation 3a20c8e5-9658-4394-8b3d-e13cbc5bd47c · inbound

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models cites this paper.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.366150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.366150Z digest=sha256:1dbfea3c744e59c9764c80daadef973dbdda7047ad6a6713df32a20d50e7f8f0

Observation d7edd419-47c9-47fb-babc-346f8edd4be0 · inbound

Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency cites this paper.

Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:03.468248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:41:03.468248Z digest=sha256:c2b4c38b40778c7fd596590e78f106b055cce0dca3904fab96d64282c8010236

Observation b005e4b9-97e1-43e2-939f-20470ea8ce43 · inbound

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion cites this paper.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.747310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.747310Z digest=sha256:6fc605964be935aaa47b119661dfa42d94465d61a48fcd06f048e39b750308c0

Observation b07fa147-f2c7-48ef-9f80-c9173bc1072a · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:59.821440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:02:59.821440Z digest=sha256:eb9cb8e5d7816b937502b0d2175eb18195d6a08b9e5ad240d83ba63ede6bd49a

Observation 10c02465-053a-418b-b21f-3b8822bafc83 · inbound

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding cites this paper.

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:25.703132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:25.703132Z digest=sha256:1da33fb80bd6dc85ddee356213f617fe85cce951713b9ceba4e0b9adce2cc3e9

Observation 2ef9d158-91d2-4043-b920-430cefab504d · inbound

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking cites this paper.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.778162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.778162Z digest=sha256:262d8a7e15e456de89267f1d6a4e824e80b9c988dc67ff367b52c82fc31dedf9

Observation c2782dec-d4a8-494b-bbe4-8deb106d887f · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:51.015037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:51.015037Z digest=sha256:350e256a155e5357996e01a36e7a68bbdbf84b8b67275b486d5aa14e263f7e5d

Observation 7381ecb4-c8b7-408c-aeab-a652d13bdaa3 · inbound

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation cites this paper.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:12.861061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:12.861061Z digest=sha256:aae645dd2f44d6834d2db8ab57df023ea88ce586e3d09c551c2b1ac029a8f75e

Observation 3c9cd646-4ba3-44ce-8828-ca0f1aeea3ea · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.613827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.613827Z digest=sha256:1cac2b7e41cfbe77ff2d4d90b5a8d92f94f9cc29954167e405c523873fc9b509

Observation 9c6f1df4-7b18-491e-9890-07aaa8a4c393 · inbound

OutDreamer: Video Outpainting with a Diffusion Transformer cites this paper.

OutDreamer: Video Outpainting with a Diffusion Transformer ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:53.460393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:53.460393Z digest=sha256:2c67651ad85fb0c1768c56402b32b56b58a5fc9bb4071613791e46f2ff823802

Observation 67515cb2-5b29-4a06-be65-585b37f24353 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.489965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.489965Z digest=sha256:5caa16e765564b31b1019a842212edb059f733c05aac0e715ec275b78a5bd2fb

Observation 2da34952-7e06-42c6-ad0f-7a60dfe52fc0 · inbound

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs cites this paper.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.432529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.432529Z digest=sha256:d862f52dd3fc0c10f8a52c00fb560a11b853d211becabccb0f745e8a5cdbf1f2

Observation 56eef71a-7d12-400a-8d99-5b6bc27ffb9a · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:41.519124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:41.519124Z digest=sha256:5d7e4a46101188bc5bb1acc49d5e306b42f5efad3bff272efe678b9d931f23d3

Observation 7255e154-2e35-4724-8474-3cd9e6b63110 · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.079542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.079542Z digest=sha256:bfb1e72644b287b39eb25f5df1993c7db4406c4e3b894537c70be9b74fd7f511

Observation 457862a8-afb3-48f7-85d2-efa803b70330 · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:37.464967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:37.464967Z digest=sha256:bd5d429984828684128fbc710e7638eb14c9d203d1020281f9053eed4eadad83

Observation ea14db05-ebb2-4935-96bc-d8559c0b8a65 · inbound

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark cites this paper.

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:05.984791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:05.984791Z digest=sha256:fa72e2430ca5ea9c5d21933b6258650a3faa9cced52b0a48746e2c402540d7e7

Observation 0aa4d782-5645-4153-abea-5c2f8d9af55e · inbound

Sample-efficient Integration of New Modalities into Large Language Models cites this paper.

Sample-efficient Integration of New Modalities into Large Language Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.385958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.385958Z digest=sha256:02749f2805a2edddb41f4cfcb13ef976daf9bcef3cf9e6d317f970d0d3aee7df

Observation b8bf05d2-7dc4-4c8a-942c-c4cdef0977f2 · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:15.112945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:15.112945Z digest=sha256:6dd98a65cb7610b373c8c4995c2996485c8ba7fd62bb79b4b0b87e365908de4a

Observation 642d51a8-20ff-417c-9e70-b963f6116528 · inbound

Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos cites this paper.

Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:48.013222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T05:26:48.606759Z digest=sha256:e61ccf60dedcd94c55cea7519555971c9241ead4a5c42873500355bf239ed862

Observation 6e744474-39a8-494e-8bac-3deef550e724 · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:52:12.619855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T03:49:58.240883Z digest=sha256:d889bc8c3c4aba5edef84262be5046a2a2face0fa6cc4d5c375eba152bbfe790

Observation c858b49c-ac43-48b9-b779-a0dc22e0868f · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:09:50.452757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T06:06:20.658030Z digest=sha256:21fbc67e04b69bca123587d95a11fa44a99f2f9c34b8a011bdcce0e811c9e2ed

Observation 6241f389-bd1b-4315-b1d4-789929997f9a · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.313346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:44dc5d06d3cc1c85476b5149b642bfc25a7bb145aa1921ae8508d7f6c5557df9

Observation 124ed411-0beb-4cc4-8191-f166e6c26330 · inbound

UNIVID: Unified Vision-Language Model for Video Moderation cites this paper.

UNIVID: Unified Vision-Language Model for Video Moderation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:07:09.294420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T22:56:26.674841Z digest=sha256:0946f8de68de5db5ab8bd4074fa4b2e8c9b82baf6520e1bfbd43c0519a14de14

Observation d9cb0997-5ccb-4b5f-89c3-df9657029bc5 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.970978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:cfc3b69adaf39e9ea346ef8308b7f3bb4b28045e91c4b00af8c1333eb01df5d3

Observation fab0ff6e-ec31-4328-acb1-0c8a0f266401 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 134

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.678650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:c9d18117f55fee7f8b1b51ea2d167cd21ec9c891ae93b26020944c03de29342a

Observation 9268b81e-016b-4aa9-9175-b18cd1c6f956 · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:388fbfb4531c6c9de09679db70c1f9a4a7dd06ce99737acb0e260038daa94868

Observation ddeb7df9-80c5-4bf3-9afe-bdd40ebea329 · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:50.075651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:50.075651Z digest=sha256:7fa8f2a7ccbc423b39d2b17a4b9479b358afbef196d066e49de2ef7d5fe1b6fb

Observation 581d0e6d-922b-493d-957d-1ec240bfe891 · inbound

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence cites this paper.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.108407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.108407Z digest=sha256:fd8a139a897e44fd941cf7c0bce269bbc106a942a2728876a19544b38d487e8f