Pith. sign in

Paper Citation Record · LEDGER

LLaVA-Video: Video Instruction Tuning With Synthetic Data

As of 11 August 2026, this Paper Citation Record lists 100 of 190 outbound references and 100 inbound Pith citation observations for arXiv:2410.02713.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.02713 v3

Coverage vector

measured 100 of 190 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T23:20:32.330351Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 100 of 284 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:29:51.267504Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T21:36:34.348434Z

Reference resolution

100 of 190 outbound references displayed

  • verified exact9
  • verified fuzzy68
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch19

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4301052c-2903-45d7-9e0d-cba4e2288d59 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:31.147880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:254e04b096ca90126c8669efa455bdb90b35814f66e59255335d2a662102de7a

Observation 66b0dd14-3493-4daf-8f35-012140a2a899 · outbound

This paper cites Localizing moments in video with natural language.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Localizing moments in video with natural language

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.631164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:e7e0843079c8fbfae464e942f79c04335304a9f8b3bafe5ce748fcd2151cac55

Observation 655bad6a-4350-4479-b131-88e785e2e9f5 · outbound

This paper cites Localizing moments in video with natural language.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Localizing moments in video with natural language

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.633942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:5f5803b2ecda360405d93407961992e604df820c08e348ecb274f46ae925edd5

Observation 23827f82-95b0-4124-9dea-f53285e95fac · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.636531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:f4aed437109d978ce1c6dd801f05b9df217feb626396525ad90a5223ad36ed97

Observation 76d02b7a-a755-42fe-b659-ce2066e3eec5 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Activitynet: A large-scale video benchmark for human activity understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.639075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:548ce542380c11a959593acb03843ef628426da41a7fe7b06dc52d218a0f6a63

Observation fbae0af6-2579-407e-a892-124f161217a3 · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Collecting highly parallel data for paraphrase evaluation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.645673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:cad5f833c910afb51925be889eb8135d5358462e114ebb4a31ca1ad3b5c3b9f1

Observation 502ca65c-a6f0-457e-9304-9af098457771 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

LLaVA-Video: Video Instruction Tuning With Synthetic Data VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.738794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:854c5bc4ba553df3cf2a6a89c916dacd3659f13f629c4e674854a1566866949f

Observation 00a4b2d9-fc92-42a8-a72f-f26fb623fc9a · outbound

This paper cites Slowfast networks for video recognition.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Slowfast networks for video recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.652396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:e1f012223108afe9a24bdee9cda9bddea64750016e3c4c04f260031511b8b1b7

Observation 31902d00-b60d-4d16-83c7-29f32c0fd786 · outbound

This paper cites something something.

LLaVA-Video: Video Instruction Tuning With Synthetic Data something something

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.655505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:6e8e0eda3327ebb00a6fe59d41bf82d88f39598e86495cef29585caf8f23cc07

Observation 73279b0b-7784-475a-835f-2be857f8aed9 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Ego4d: Around the world in 3,000 hours of egocentric video

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.658407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:7f5f3ecac4147d91e47c20b161e50b6fe5f11d61a223bef83cbd9fe48e920673

Observation 02bbcd81-0380-4091-8d7b-d1525f9018fd · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Agqa: A benchmark for compositional spatio-temporal reasoning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.661057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:139d2a9ade604caf9a93561f918255fc3059f4e3dbddfaa24674e0e35f29c733

Observation 8b0bd1c8-7485-45f5-980a-f4846de20a99 · outbound

This paper cites Acav100m: Automatic curation of large-scale datasets for audio-visual video representation learning.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Acav100m: Automatic curation of large-scale datasets for audio-visual video representation learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.664571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:089b2f35d2ba3bc31b603f37d7bdde52b79169789956883a3e8256284db14cbf

Observation 230286e9-5613-4870-a364-f26a277a7dd8 · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Less is more: Clipbert for video-and-language learning via sparse sampling

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.670379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:7785d8ad4d3dc53bd69ca67be9504a4041c45a211f68de7da1a419fbb1dd13fb

Observation 06030416-0d62-4587-91a3-5c3df9d660e8 · outbound

This paper cites Llava-next: What else influences visual instruction tuning beyond data?, May 2024 a.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Llava-next: What else influences visual instruction tuning beyond data?, May 2024 a

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.673506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:eeb1c3727d30df3e641f1a7be4344b2598848017029b6bcd3e53d1e1ebd07fbb

Observation c17b7a02-819e-4e02-a2c8-aaa5535a4ee5 · outbound

This paper cites Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024 b.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024 b

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.676462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:ca246bf09b84eea0efd562b6660453a80bf6afff582a1fe144928d62537740bd

Observation c3d63041-360f-4226-9e92-71874183a63f · outbound

This paper cites Multimodal foundation models: From specialists to general-purpose assistants.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Multimodal foundation models: From specialists to general-purpose assistants

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.679896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:f2352499b1355698463aa5169a6e6fec4a211aacfced3d5fa27823554f9180a3

Observation f9f5c995-439f-48ae-af2b-edd84930ef81 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

LLaVA-Video: Video Instruction Tuning With Synthetic Data BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:10:50.405848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:ac2de30e7c830ea8c80a220ba8743f5630e2cc12d8c5dc41935fad3f95fde127

Observation 945090e3-98eb-48ff-a09f-1ec94a9f1394 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

LLaVA-Video: Video Instruction Tuning With Synthetic Data VideoChat: Chat-Centric Video Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.879549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:ec14e7d96ef1ab11da1e8a1be580e2956aa0ab46d9f64ce7c35d6590e79132d8

Observation 1a439b78-4e3e-4342-99f1-27e05dc28378 · outbound

This paper cites Vila: On pre-training for visual language models.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Vila: On pre-training for visual language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.684018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:5a8ecca5561e23c0f609efb3f5d51a2ab776bc17e5a03f0c210f5568a66a1347

Observation d78ea22e-5e74-4614-be36-1dd0f7fc7107 · outbound

This paper cites Visual instruction tuning.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Visual instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.690864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:b867477a57f4d5619ed22b7df8c9e0e262155c5af4628a8af39b6ef106cbc9ad

Observation 17e17dbd-6315-4b2c-9073-1492eeaf970b · outbound

This paper cites Video detail caption.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Video detail caption

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.698073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:6a051d45f9b8836a71d5d39d23b1d965c0fc0d1d288e7261619e57ac2e00e423

Observation b3105ae8-a139-46f9-b45c-ebed021ea7ff · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.702860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:832f24bfb4b983316a42c69373390e77360e9e8d2643280c0d00c05ae35d1afc

Observation 759fdf4c-38fb-41b6-a66e-290f5ae01270 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.709118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:63c7a3ee3778076f91fcbe635fbd5ae56c0d24d009ee4bf239a8bbc34ee61dc4

Observation 471d5b74-6463-4ae3-9a81-01fb4e080289 · outbound

This paper cites How T o100 M : L earning a T ext- V ideo E mbedding by W atching H undred M illion N arrated V ideo C lips.

LLaVA-Video: Video Instruction Tuning With Synthetic Data How T o100 M : L earning a T ext- V ideo E mbedding by W atching H undred M illion N arrated V ideo C lips

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.715166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:c350cd0cd45b371d937e5afda1fa47a1cf38ddced57c1477829092f77f59adcc

Observation f3bd7134-00d2-47b8-af6c-0d0d1a4c37fd · outbound

This paper cites an unresolved cited work.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-10T23:20:33.717985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:a93fe66de34310b22e7f64687c237941f3661c03893880dd4b015f1eecc0b714

Observation 85fd65e4-af8a-4e1c-b389-d5de12345a19 · outbound

This paper cites Hello gpt-4o.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Hello gpt-4o

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.720765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:72f7612c375aec5ac8320e527b6a88d2044e45a6cade596157f3b629f2e5fdd4

Observation 659eb492-fffb-4a12-9c03-e000901435ef · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Perception test: A diagnostic benchmark for multimodal video models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.723455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:043598aa07ab70993f10a69e646561cc7b9fb26e00797fccad8d6e0ff7dcea18

Observation c185b1b6-8ce5-490a-a6ad-bcf64fd3ec7d · outbound

This paper cites Learning transferable visual models from natural language supervision.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Learning transferable visual models from natural language supervision

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.732308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:60d1ad61ba861e102165d5c2172f1cf0444e6487e2e016e59c895d8d14f8f9e3

Observation 0510af63-23de-47e2-a2ab-9d33ba745e87 · outbound

This paper cites Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:32.996852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:88c036099311a831b256bac47e9a2594c794c59914e3bdbbbf3317a8381d328f

Observation 6828b804-2371-4c0a-92bc-c2a141f862b8 · outbound

This paper cites A dataset for movie description.

LLaVA-Video: Video Instruction Tuning With Synthetic Data A dataset for movie description

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.738181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:c372876d1830929ad0d61656c6efcff6dd3ca95a21250b8ec71b2c43891758f2

Observation 1044fd9b-52f3-431a-aa5c-7a565aed18b4 · outbound

This paper cites Annotating objects and relations in user-generated videos.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Annotating objects and relations in user-generated videos

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.746891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:a909250a4ab63fd16e8d250c3a3d59272e87cd45e33e0ef7d20e14d33bb32247

Observation 881e83fa-9b05-4b89-ae6f-f6e0a7292f14 · outbound

This paper cites Hollywood in homes: Crowdsourcing data collection for activity understanding.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Hollywood in homes: Crowdsourcing data collection for activity understanding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.752336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:5b896565053c395604b59b74c5a545a48c7f83a10464a5188d02dacd56638c14

Observation 0be139c6-460b-4b87-ae5c-c34e807f64df · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:20:32.962559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:6f8df5cd5d01ab3c29ea68ee939fb8e7873fd97d10784d9074fdba5f53d8550f

Observation 95d1cba9-dfb0-440c-9a95-63c8b7808591 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.755253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:d756f539d37eac6752ba53d58dd7d35f940d1182ddce0158fe4254e9f15c8ee0

Observation 13c7293b-8856-4dbe-af2f-33d45d80e6b2 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

LLaVA-Video: Video Instruction Tuning With Synthetic Data LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:30:12.519784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:f56f73b03446e987d50d04fd5e92497d34d3dce7e850f064f00a52595904d358

Observation e48d6f90-c0d8-4c76-9fae-0e74d949dfa9 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Next-qa: Next phase of question-answering to explaining temporal actions

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.758033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:eb0684dbf939cb48c97ef6cff10262412977b9d593baf279c4cd5d7d0f3d57c7

Observation 1b10e2c7-aa1d-43e2-a878-7e361e573fd0 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Video question answering via gradually refined attention over appearance and motion

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.763103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:23a4c2dc254f401b1d8916d0e2034a057b0eb2a6fa51048dbbce1a38430021c1

Observation 7650c80a-2a37-4b39-94bb-361ea1bdb894 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Msr-vtt: A large video description dataset for bridging video and language

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.767009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:1daa0a1f568c99ecd33ff8d06d33959650c686f3bc587e74e3a2e8ce2b6791ae

Observation 247feb16-86fd-47c0-a678-b8d9db22fd23 · outbound

This paper cites Advancing high-resolution video-language representation with large-scale video transcriptions.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Advancing high-resolution video-language representation with large-scale video transcriptions

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.775120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:c349434fcde503b622d4aeb6a8e9716a3ecca8919917fa79373473a253c8530b

Observation 8076cbd7-5638-48d4-b610-64ead95d16c3 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.778232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:d08a089b376872c4a3ec00ba367a3a2c381913762cfbab8c7d7ffcada5c5d9f6

Observation d876a09d-f8a6-44ce-9a46-70e84343ad8f · outbound

This paper cites Social-iq: A question answering benchmark for artificial social intelligence.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Social-iq: A question answering benchmark for artificial social intelligence

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.781008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:ff4b27e0134af2e3e12e10d7fe88b73feb0e9044e6509cc247fe7de01d6c4884

Observation 5d2c33ba-49f1-4ca1-86c9-9915819579b5 · outbound

This paper cites Merlot: Multimodal neural script knowledge models.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Merlot: Multimodal neural script knowledge models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.786126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:09b434e3d055805ebaeefe57ead2faedc08735a9ed298cd0e1966ab94133ab7b

Observation 978b6f3c-4301-4ef1-8d03-d9067de561a5 · outbound

This paper cites Sigmoid loss for language image pre-training.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Sigmoid loss for language image pre-training

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.790395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:0fa86c28099a12e6778927b4f2f5204f66a7c25beb08dfd611cdcf20545cc096

Observation cc265d68-56a3-4e44-be49-5edbc41c42b3 · outbound

This paper cites Direct preference optimization of video large multimodal models from language model reward, 2024 d.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Direct preference optimization of video large multimodal models from language model reward, 2024 d

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.793622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:9cf1ddfdc8f663873c6188daaf784d2ca461066108e76aeb96ad53146bf621aa

Observation 989f9f34-5446-4f2a-bed0-7bcefbbe4e4f · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, April 2024 e.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Llava-next: A strong zero-shot video understanding model, April 2024 e

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.796162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:c3e8b858c55b720aeb9f1da4fc737c744880e90ebe3374aa031c6132d42c2f6f

Observation 038e60f8-da92-414b-9ef3-356f2041dc95 · outbound

This paper cites an unresolved cited work.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-10T23:20:33.798484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:7cd6f8775885c7a62cd68b433970a67414f99915115bda6a2d116dfadcb46102

Observation ac8ac4ed-4906-4c74-9462-d53e945408da · outbound

This paper cites Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment, 2023 a.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment, 2023 a

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.801244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:150a1c055c08bbc16202b5614438f417d0eb8da24d196e2d7bb7d42e7eaa5ecf

Observation 815f9ce0-8cdd-4806-a4f9-961be6afd4e3 · outbound

This paper cites Visual Prompt Tuning.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Visual Prompt Tuning

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:33.113505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:7ed3f9cf20e3218fa1b26d6593440ebc9431542f43cf6e86ce9a4718a035dea6

Observation 851caa27-a1fa-421d-9265-2d9f1a246e0b · outbound

This paper cites International Conference on Machine Learning (ICML) , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data International Conference on Machine Learning (ICML) , pages=

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.803516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:400791ad253f9333892ecda55533332ccd605c95bc3ee0b6f716f9f9177ae390

Observation 12e5c4b7-7708-480b-a5db-e7190070fa95 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

LLaVA-Video: Video Instruction Tuning With Synthetic Data LoRA: Low-Rank Adaptation of Large Language Models

Reference 76

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:20:32.505718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:0a09ea7d429dc859957500c53fa5e80c0c5e3b12018384e260d482a06a137925

Observation a7721538-11c9-406b-af25-5b10904daabe · outbound

This paper cites Towards a Unified View of Parameter-Efficient Transfer Learning.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Towards a Unified View of Parameter-Efficient Transfer Learning

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:20:32.551150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:79a49c23b8489155232ddf52ede74611a4a413812e2e3fdb75b1f678a123f836

Observation bf8445e4-ab2f-4cda-869f-ba74c271fd6a · outbound

This paper cites Factual Probing Is [MASK]: Learning vs. Learning to Recall.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Factual Probing Is [MASK]: Learning vs. Learning to Recall

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:32.671954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:38167b0dacf5dcf8ec51439502dba3766bc40e402bbde50d311900d034e0d880

Observation b8f9d25e-f5a3-4096-b93e-d1a79eb07e64 · outbound

This paper cites Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:32.683322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:2df8ed6f619f75a3713721f7104929115d19090a2b3ba71d8dcbcb535aba3843

Observation 13ac991c-6321-4026-bd09-d1972bb5af00 · outbound

This paper cites Advances in neural information processing systems (NeuIPS) , volume=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Advances in neural information processing systems (NeuIPS) , volume=

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.808080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:04747334e65ad1f42db34878c9729594f94d088cffba797106b5b5ba8f198977

Observation 65e560d1-d32c-4388-9613-7d8015d3bbf0 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

LLaVA-Video: Video Instruction Tuning With Synthetic Data BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 81

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:20:32.722162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:2bd2328396963eb37efe18600de65394d54d28a1f3fc2b3aa64f995feb0d1901

Observation 548183ac-c428-4bac-9a42-8f1c91d45264 · outbound

This paper cites Advances in neural information processing systems (NeuIPS) , volume=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Advances in neural information processing systems (NeuIPS) , volume=

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.815081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:379eb34e358243e68e921c1b40d2800bf6ba079698de999a07732b907b54149e

Observation 99b7857a-520c-418c-a08c-269cfde97a3f · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:57:25.623351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:06771d778092b212385481037ecd8fdbe1797d2c4a71d07258f008004ba851f8

Observation c57b2c5c-6d40-46db-8a70-2d7dd357c908 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

LLaVA-Video: Video Instruction Tuning With Synthetic Data The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:34:05.826717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:439252c868d79a7c03999ccf71c80bc69cd67e3305662de091bd7792f59183d4

Observation 0dabb73c-8813-4f71-a429-74a23878f851 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.817845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:40166158d3cd768944fb6a23c5eafb41e661f04918db1a90548d115b96ee6b96

Observation f1cce10b-33e4-4261-92e6-125f6063e2b1 · outbound

This paper cites International Journal of Computer Vision (IJCV) , year=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data International Journal of Computer Vision (IJCV) , year=

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.822720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:76c5a7d2498f1a8f401dc69ef371a5c25312d244235316b9334f9841e8d22409

Observation ed9d954f-8fea-4fa6-a205-fd88e978b7b9 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.825276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:f6a58d0e57bfb7acf5643e8165aabf0b7965511ceeecfd78aaa24a0481b13b8c

Observation 8b1a8a24-e2bb-41d5-82a3-04139c63d618 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

LLaVA-Video: Video Instruction Tuning With Synthetic Data An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 88

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:20:33.070234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:3f77ad6c9781c6058b059267f93f79f0cc083daa77527921684bd029c3da8d42

Observation b65256a4-561d-4e0a-8e13-4fbf232632f5 · outbound

This paper cites UniPELT: A Unified Framework for Parameter-Efficient Language Model Tuning.

LLaVA-Video: Video Instruction Tuning With Synthetic Data UniPELT: A Unified Framework for Parameter-Efficient Language Model Tuning

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:33.080613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:5cd9a933403f025750fb6dd6f50beb81f2f1252505d3b7eb42b7778ac3956a5f

Observation af0df508-c6f1-432f-a240-b6d7eadb743b · outbound

This paper cites Neural Architecture Search with Reinforcement Learning.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Neural Architecture Search with Reinforcement Learning

Reference 90

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:33.090752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:8894f7a25e1aa39550471060eaf1c2f0464d6b652dc1e0486b208dd57da03596

Observation 44187c4e-6bee-4547-b0c5-c33f95d3041a · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.827651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:38d1b98477bc3ec1cb84b5305fc5619b51aa17d43e15ad6f6a5496a79cfc2531

Observation 0b8cfa88-2e5f-4dc9-9118-c979bc5ed0a1 · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , year=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , year=

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.829922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:80565caaa5860d592b73f5b6f9662ef47bb5644f6f1ec644d5beb78069aa54a8

Observation cff8ec30-f885-4feb-8412-25c83ffc151e · outbound

This paper cites International Conference on Machine Learning (ICML) , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data International Conference on Machine Learning (ICML) , pages=

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.837518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:78f33713e76330e0ecee502fedbabb151cba76140e0c109e8eced7cd3176ae10

Observation da59bf03-ebb7-48cc-96a0-65743ee8bead · outbound

This paper cites International conference on machine learning (ICML) , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data International conference on machine learning (ICML) , pages=

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.848698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:150215cfe5b6184500663dd8a5ce86b5b71afa9878551b6146eed7660039f2ad

Observation 9d6537fa-6fc8-4b2e-b057-07ce0c2cb512 · outbound

This paper cites DARTS: Differentiable Architecture Search.

LLaVA-Video: Video Instruction Tuning With Synthetic Data DARTS: Differentiable Architecture Search

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:20:32.657992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:58857ce61ef45319711fd88905af09eaa436c1918e1888d7fe189d71c4aba2d6

Observation f4feed9b-5d73-4ac1-a7cb-507cc782c787 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pages=

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.854265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:7aeded28cba598289cb91802f84d5f64a5ab8a267ce297e070dfa81e6df3c123

Observation f2488925-6126-4f6d-8399-b4319aca5bcc · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.856889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:b58ebb4c3ae893e119ae9731394e3bd594784dbb25071da396b0906c29974468

Observation 7cbb515c-9b65-4e04-9fc8-a7d8d0718e9e · outbound

This paper cites A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark.

LLaVA-Video: Video Instruction Tuning With Synthetic Data A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark

Reference 98

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:12:05.852900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:a33a313d09e89bbcf6b73ea4dc911f19182f65722ea0fcdda4f10cb2feb3d9a0

Observation cb77a37b-9a32-43d9-b740-b43eeaf694e8 · outbound

This paper cites DeepMind Lab.

LLaVA-Video: Video Instruction Tuning With Synthetic Data DeepMind Lab

Reference 99

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:32.704663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:64215f1cebcc72a2147517020293da42b64e4f895da24292a152a7feddaded1b

Observation 079d7a4b-b31b-487a-a4f2-96cb863087f5 · outbound

This paper cites an unresolved cited work.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-05-10T23:20:33.859166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:d65a2b953372074d92a1b1684ca44b84e0f6b1e15810c87a283e79eeaa18713e

Observation becd0eca-edd6-48e1-a4cb-0ac702bedc2c · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.861606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:47c95c278c5910a47f2f34a508baccc3be5f0c268838ce0f01a5acce4d1cdcc6

Observation 386789d2-454e-4bc1-864d-34a9202ebb3e · outbound

This paper cites A Generalist Agent.

LLaVA-Video: Video Instruction Tuning With Synthetic Data A Generalist Agent

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:24:50.082909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:fe67936649680cef9fd0077801911cfe617fc5e5354ec9fd4098f653f75934ed

Observation e151b766-d10c-4e28-8602-7d59752067f5 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , volume=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , volume=

Reference 103

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.864909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:566e0eb845d1b4836e949c5d05e79dadac47c5ef9b73988ad9152b416a6cfb77

Observation 6151fcf9-2484-4cc6-99a4-9ea3ac4333f4 · outbound

This paper cites European conference on computer vision (ECCV) , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data European conference on computer vision (ECCV) , pages=

Reference 104

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.868968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:8d5dbf2463b4b43e846891d6dd54767ab10add143e985c36704a2baab5a662a2

Observation 15138065-4b00-4e76-bf61-966ddadb9054 · outbound

This paper cites Domain Generalization: A Survey.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Domain Generalization: A Survey

Reference 105

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:32.827044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:c981066b73f39c7f332f9bc1d359e96b61ff0366b552d8df85a6f50571e73f6f

Observation 3f589d9b-43fd-44f4-b4d8-c27517ce5713 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Florence: A New Foundation Model for Computer Vision

Reference 106

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:38:09.598269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:d2f69e2e457318c0509fa06d5f2949a2558fd53f28352dab24f398d8f1f3e331

Observation 3fa01dc7-494e-42cb-8ae2-601bb5a9a1f1 · outbound

This paper cites Bamboo: Building Mega-Scale Vision Dataset Continually with Human-Machine Synergy.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Bamboo: Building Mega-Scale Vision Dataset Continually with Human-Machine Synergy

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:20:32.843466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:475a0f0e4a3f7e038a091a78fc26fe73480ee8725eac15d1c2f36c9703d9315f

Observation 2031154a-0a67-4953-9281-e3ff212fcf98 · outbound

This paper cites The International Journal of Robotics Research , volume=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data The International Journal of Robotics Research , volume=

Reference 108

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.177342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:95ac617a8d0117a005254e56bca5124eda00ba22f8274fd9c15815cb09572780

Observation 6468b77b-5232-4c11-b814-83c8a6f35115 · outbound

This paper cites European conference on computer vision (ECCV) , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data European conference on computer vision (ECCV) , pages=

Reference 109

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.180846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:cd0f5baa2b65cbcef42014158f500ffb85a70da7328e85f66fb0e3f80e012da9

Observation 6699bcf1-ff8e-4feb-b031-aefe71b4b598 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , volume=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , volume=

Reference 110

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.191997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:333d8a38bf04f4de131f7fe21e76926e8d20055c6ba6074eaf8f4e94d1360aa6

Observation 6d4769f5-2af0-4189-a26e-0905173bd6c6 · outbound

This paper cites Proceedings of the IEEE international conference on computer vision workshops , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE international conference on computer vision workshops , pages=

Reference 111

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.204880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:f40d8745162036951560ae2ab1f579243f2a4d4c0cc67e83632fc6919652608e

Observation 177e2e96-0336-49b6-b186-406a605a0d42 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=

Reference 112

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.215197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:513a8f95fd6fc03f99eca1d1b8cf457ecb53909be4ff195e50ebfc2c60f825ca

Observation 1bc02244-54b3-42ed-bc46-4ebb6ea780d4 · outbound

This paper cites Fine-Grained Visual Classification of Aircraft.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Fine-Grained Visual Classification of Aircraft

Reference 113

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:41:07.065951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:746d565656c3de6e9b5d1754751672f9a1a7d1cbe3234c5c06fde56d0c5e3b20

Observation 11037149-2c26-4f9c-ac01-070d7d6b68e5 · outbound

This paper cites International Conference on Machine Learning (ICML) , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data International Conference on Machine Learning (ICML) , pages=

Reference 114

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.231663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:af63ca336dbb99a62c9dd3aa8b7c0175d5bdab9be9b91dabfa9980b0fdf01bcd

Observation 1fe8658e-da36-4b0b-b222-812b1063b19f · outbound

This paper cites Advances in Neural Information Processing Systems (NeuIPS) , volume=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Advances in Neural Information Processing Systems (NeuIPS) , volume=

Reference 115

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.243801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:4c5d3e52650fceb4640c43bfc95b9f42dd289326fe396d4e4ac6d6048501fbfb

Observation c12945ab-129a-4116-8b49-11ef163ed4ad · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=

Reference 116

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.248943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:7b5b216712274bb8c8ecacc88f4594b24634ebd1e84563c13b42c5510546c98e

Observation 492ef995-fcaf-449d-847e-1f697b9d95e2 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pages=

Reference 117

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.253338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:5529a0626b289143df53de4d4340d9aa888102ee80b2da2c45b4a5434f04bfdc

Observation c4c79ab6-9f45-4398-a24e-defaa9e5fae1 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

LLaVA-Video: Video Instruction Tuning With Synthetic Data GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 118

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T21:24:15.760169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:10ea138ee19164ea2d5c67b0561f9529eede142e282db6f3777ec224af71e0e8

Observation ecb94d70-e088-48f9-85ff-aca8c816f96a · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=

Reference 119

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.258796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:d432ee610e7c933b3903a208a7e2af7978957548f6f7bc19bc4fe2fd5562d05f

Observation da8c7e21-26a4-44d9-8fcd-85d1c68b5489 · outbound

This paper cites International Conference on Machine Learning (ICML) , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data International Conference on Machine Learning (ICML) , pages=

Reference 120

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.265599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:97e458df1a2720a2c0329b7be38ed3482f25e1415a39459b589ed16a9d504f4f

Observation ba6f137f-115a-4c0f-9c18-fd65ee659570 · outbound

This paper cites Masked Autoencoders Are Scalable Vision Learners.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Masked Autoencoders Are Scalable Vision Learners

Reference 121

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:53:57.787532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T15:38:38.318698+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:2fec9b7de87a9c4933d2e5e484c8730bf09fa1910932d9dfbd1fd5feb42cb270

Observation 8eab20aa-1df1-4fff-a0f6-48c1e6598ab5 · outbound

This paper cites 2009 , publisher=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data 2009 , publisher=

Reference 122

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.271777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:332357519020d9340142575da471c666f82541d5f1ccc50435c02f5e3feb1193

Observation 823fbbd0-89ce-4027-9d77-4e4df0d39b4b · outbound

This paper cites conference on computer vision and pattern recognition workshop , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data conference on computer vision and pattern recognition workshop , pages=

Reference 123

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.276208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:31261dffbb44938bed57b4545b9124bf29fd09208cf2351d40c60ea7b308db5f

Observation 868691b8-0914-4978-bde2-a6ffbf9a5ed5 · outbound

This paper cites Cimpoi and S.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Cimpoi and S

Reference 124

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.290469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:aebc9cf30cae0fee44b74fa1de0176c166b6ee125a0ce51ec678c1a7abfbbe3a

Observation db0b8a0e-8d0b-4442-a52f-fc3ec12764d5 · outbound

This paper cites an unresolved cited work.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Unresolved cited work

Reference 125

Resolution
unresolved
raw_fallback, observed 2026-05-10T23:20:33.293151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:328fa473433b77a0c432708c0da57611d3127b50f0228e769a011c0a24284ccb

Observation 1cc1e7cb-037b-4b85-b3c4-d45c773c5660 · outbound

This paper cites 2010 IEEE computer society conference on computer vision and pattern recognition , pages=.

LLaVA-Video: Video Instruction Tuning With Synthetic Data 2010 IEEE computer society conference on computer vision and pattern recognition , pages=

Reference 126

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T23:20:33.295995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:f54c57f56694b0c6edce89f96293bc5b504b959abe667d409389b49c8f63794e

Pith citing papers

Observation 5a3218de-1b04-4090-986d-e0b103dd39fa · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 147

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.192285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:78b78ae4adfe2a42c4243ef197e650fcb49493d5caed2b5a1d068a70d3c3272a

Observation 1e36f211-c965-4223-a32c-3496cfddc48c · inbound

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces cites this paper.

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 104

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:27:44.083800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:27:43.919941Z digest=sha256:cf6a02c0b288c2bd6b8330ec093e422f9a1e60da7e442cac59b5633bc84b2e12

Observation 976db303-747f-4c80-bb8d-17d02eaadee7 · inbound

VidCtx: Context-aware Video Question Answering with Image Models cites this paper.

VidCtx: Context-aware Video Question Answering with Image Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.267504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.267504Z digest=sha256:3cf392d2c1b72641f30e2b51167dfd23df144047a0da0529c3f4fd7322f1be22

Observation 84022ade-c828-472f-a4f0-b9028aa05dfa · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.632932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:dbfbc02511f05f13bdbbbd8411eb506cc0999918e366404430b08e36197444ed

Observation 10a626aa-374f-43e8-b0d1-20146cd3af0c · inbound

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding cites this paper.

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:56.504052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:56.504052Z digest=sha256:d67857a665fd24266273926bea82f5cd671f6731a40f5e8fc022772fb5d65afc

Observation 3c991fea-0792-4d41-a4f4-08a52fb229d2 · inbound

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models cites this paper.

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:25.171264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:25.171264Z digest=sha256:8fb7de3a5e4c0d996e588a4be9513b98ec85afb06c232e92a3ef0f97eca64165

Observation 2daef667-331a-46dd-8bfb-59fbce9c1109 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-23T05:45:28.399033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:a4fc7fc11e98f7c68c296886de66c48b505a27b53d263fbc23c9438fcf11732a

Observation b5267eb6-df57-4449-8bc5-3b6215c980dc · inbound

LongViTU: Instruction Tuning for Long-Form Video Understanding cites this paper.

LongViTU: Instruction Tuning for Long-Form Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T21:23:58.063067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:23:58.063067Z digest=sha256:e1941156e63d7cba5dcc9f19c4894713ecd13a45e0a7bd00145e8bc5c6f95fad

Observation e86efaa9-45dc-49a9-bece-0435d3fd93e1 · inbound

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model cites this paper.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:41.330419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:41.330419Z digest=sha256:59be667f1b2ef16697fb1d2e7c9b059bc1577b61bc5def7923e65e46bf4c91fe

Observation ccf57c0c-21c5-4ba8-9e87-41ac7bf780c8 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-11T01:19:59.773347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:7c55a3c88090ebff734978a517e9fd156c405bd7ae49b64f4f55ae2422129abc

Observation 0e5a4491-57c8-48bc-8d32-d175db17ac2d · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:32:41.197198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:0473f97d6d56c17c2a76932056bc4db313223cec64406890fd902cb368781102

Observation 10ba334a-bb4c-4f00-b439-b1f38cf09b18 · inbound

Temporal Preference Optimization for Long-Form Video Understanding cites this paper.

Temporal Preference Optimization for Long-Form Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.317413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.317413Z digest=sha256:bb6b0b956a3c21b4ff4a33fc67dd9eb1e4fb987f72c7d1ff3c67b6755c8bb34b

Observation 7ce12156-d8cc-4ea6-ba88-ee60aab2a8c3 · inbound

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler cites this paper.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.086229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.086229Z digest=sha256:e9395724776b4d5c639f800b7e77038c0e1bc9ede1937e6960e5efebc3df379e

Observation c0aa3591-a22b-459d-80b9-f05cf52d2255 · inbound

VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos cites this paper.

VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:36.856590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:36.856590Z digest=sha256:4d48bc6270fdbe93ae394dd415388e14d50457b2ccba7034820625d22f7feadd

Observation a1d6fdb8-f008-4e9b-bf41-016f8e7b30e9 · inbound

HD-EPIC: A Highly-Detailed Egocentric Video Dataset cites this paper.

HD-EPIC: A Highly-Detailed Egocentric Video Dataset LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T23:26:20.596294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:26:20.596294Z digest=sha256:01619051b21ea988fd502338e1d9610d65a4e1d836c1430d9106c5a96f20523e

Observation 4a483b44-5e09-4120-b675-e950839ff857 · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:53:26.245939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:82a56a48afdef2a4653ec480e74597a29eec34fc02bb91d8da7ed8ff6520fd8d

Observation 73da6a09-7f01-4c25-b393-89a6b3c6cfe2 · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.551232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.551232Z digest=sha256:670648a9dfa605e25622ba25618ff1dacc2c8e64aeaefc75187ca23453478752

Observation 07ed3bcc-cb45-4f90-b7e0-6f2b455b14b6 · inbound

Unified Reward Model for Multimodal Understanding and Generation cites this paper.

Unified Reward Model for Multimodal Understanding and Generation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:44:30.925258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:44:30.558048Z digest=sha256:b7d2a4f514aafa4e0d835596ddb42fcc8687972934b4f012447d1c946ad2ae10

Observation 28953e7a-4738-41e5-8848-730399223951 · inbound

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness cites this paper.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.151308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:ca5d6e4b858abee8453fe22603c39eb3b7187ffc5cdc6a8895607734f6c11904

Observation b57506b3-c9fb-473e-94f1-db33721ee035 · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:23:51.750015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:bd8f9110ab74a4b5da8a6b18680a862224adec3acc3e1dab07e6441a2752a445

Observation 3caa9af6-8bcb-4af9-a24a-5ae85ccc00ba · inbound

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning cites this paper.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.764632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:643420e7425874f0a6cb7250aa1eb285adcbdeff6db3876f54a3a0505e8812da

Observation aece33d9-7349-46da-b737-66bdd2dbf49b · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.495843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.495843Z digest=sha256:fdd686f6886052d51e018e51cce0372d7403270691cbc9294c80fb9736c44ca4

Observation 37f6e468-d31e-4e85-a9f6-ac9bc1835124 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.910000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.910000Z digest=sha256:07ad14e872adb2a558efda9724eee8b165bfa6f4d6301f68d1c0ce5cf8a7a699

Observation 75ba5317-b5ea-4f21-b12a-517e7916add4 · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.633738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.633738Z digest=sha256:0b15eeb9a1fdc4b13752c2f1820393739f6785d003c2e5b81578cda22c509b42

Observation 8eacf7cc-13ab-4b30-9181-67af9725b7e1 · inbound

Clapper: Compact Learning and Video Representation in VLMs cites this paper.

Clapper: Compact Learning and Video Representation in VLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.883282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.883282Z digest=sha256:3bc746d001aeff83079380209d3d34abeaec2266fc8895f462438bd69c9515a8

Observation 331ed5d7-2084-415c-bfeb-bc831b564724 · inbound

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning cites this paper.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.403500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:c70bbd48026345d498285ca24bce650b5435a00776c4aa5fdf8bc7dcf2141479

Observation 08d2a24f-71fa-421f-afa5-f205f726ef01 · inbound

Inference Compute-Optimal Video Vision Language Models cites this paper.

Inference Compute-Optimal Video Vision Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:45.034743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:27:45.034743Z digest=sha256:b0b8360f5ec2169a1ab4981e847ffb62d1c7f59a5d13b6e954831e872c74c2d7

Observation 1acf10f4-d309-4cd2-9127-00201219f46b · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:47.117724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:47.117724Z digest=sha256:a844266eab3632c250a2496b2f602fad854bdc414040a8054a1ffb2d0653322e

Observation 98a25c67-145d-45dc-be12-5c0ffaeba027 · inbound

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought cites this paper.

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:16.030499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:16.030499Z digest=sha256:eae7402a4d46fecb39878842630f7eb03e05e930130a4ab8c473d5a31fb2bd22

Observation f3a496cd-ad53-4c95-84c7-41cd139ded15 · inbound

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models cites this paper.

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:10.697724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:08:10.697724Z digest=sha256:b8ec8e7500818768755870a41dd625ff836286748c5799785aaa58793db6ce4b

Observation b1a20b4b-0c71-4321-a18f-87146e3c31f0 · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:05.687068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:05.687068Z digest=sha256:695537dc5f0f33044565f953e1865736cd5c27d249dda5bcc474fa6f82cd5d55

Observation 9f3976c4-30c8-426f-a89b-b2cd0d54fcb2 · inbound

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction cites this paper.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.940393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:88cd6c141cb1f7f94d3d6513a2b640842a506ae6e67002725e777ba9cb35b92b

Observation 123279b9-48f2-448e-98d8-9f2e0c7aed8f · inbound

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts cites this paper.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:16.849356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:16.849356Z digest=sha256:0655cd6b990498968f206981c601a6ea00032689adefeba48fd340bd59ca3022

Observation 62ee6a8e-3681-43ed-8861-7810bfeeed6b · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:40:56.015162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:0a0462b68d3ecd6539eb736a508e7b695a2db2d09c4246d6c034390a7d2ae600

Observation 5065b583-6453-4089-8ddf-b0e1f36b6f9d · inbound

Fostering Video Reasoning via Next-Event Prediction cites this paper.

Fostering Video Reasoning via Next-Event Prediction LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:46.144299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:46.144299Z digest=sha256:cb9b7c68df78e0f031944885f23b44a345e0598b3e765f804ee7f8323a9bc210

Observation 120b31ac-4054-4861-882b-b452c3ddd294 · inbound

Universal Visuo-Tactile Video Understanding for Embodied Interaction cites this paper.

Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:26.216112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:26.216112Z digest=sha256:884064b90fafce5d813d97fb09a0bd3c364b8d03a4aadd3e748624e678f6fdef

Observation 604cc1a4-c2bd-4554-9788-ffd5f7ee5bf6 · inbound

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation cites this paper.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:03.349403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:03.349403Z digest=sha256:05cf36a183b162c641c50cc8d2186eaf3646d16c3df578af751361c05533bbe5

Observation d0c641c0-38e4-4aeb-a236-fcdb37168342 · inbound

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory cites this paper.

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.881781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T12:54:31.765909Z digest=sha256:585ede93e5b2764deeff7310cb195f1dab3b3b0adb6dd06ce00be97bacdddfdb

Observation 7b318f2c-c5ba-4ee7-b673-9e291a49eed9 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:34:36.995325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:c73466c0f8b10a6b8e71f683aaebfc6cf8eaefc34cb39940542e277fb210b4c3

Observation 0051c6d8-b97b-495d-8d5f-7257e36b7dad · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:00:51.308224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:5476ff452f99f5d517c5adabdccce935e5de2fb89b3c8d28d95358687225a87b

Observation b30ebedc-b77a-4463-b3a6-ef8c7a773f6d · inbound

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders cites this paper.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:24.258273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.258273Z digest=sha256:ab38168341b90a98a9fc072aeff418578827bcb17cf3621ffee92eec755bdeb8

Observation 5762f209-5a99-4b11-9c1c-55cd14564bbe · inbound

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times cites this paper.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.604138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.604138Z digest=sha256:c392c5abdccfab8eaaaa2fedb7e6e69c31da28a5d682941094893c8df8f37afc

Observation 7f6d8b62-bb23-49b7-b3fc-4b9f9fd99e74 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:11.136490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:11.136490Z digest=sha256:bdd10d53518fb4c240c9aef01a47ae99828d79144f8f3dc9a41c0243362164a3

Observation 9410b52a-a52a-4521-9fcd-8ae683fdb1c7 · inbound

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding cites this paper.

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:29.980105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:29.980105Z digest=sha256:d8184e038d4beff4622ee05b5930c59545cef625fe69c389945e84c3a57c2f37

Observation 759fd701-e7b1-4f1a-86f2-762e8626f470 · inbound

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding cites this paper.

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:05.475245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:05.475245Z digest=sha256:8340d0eb6723526d03958539c5f49c841a0b399a19934020ba7fc7ae26182fb7

Observation 19ab9372-e3a4-457d-8ffd-1ef64bb7114f · inbound

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking cites this paper.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.687621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.687621Z digest=sha256:7a44ba16756fae67ece561d70d5e6325ec83ae373d5a19b7e1b186d19c38748d

Observation 3078cda2-ac9b-406d-9bb7-d6b3f3c12bea · inbound

Is Extending Modality The Right Path Towards Omni-Modality? cites this paper.

Is Extending Modality The Right Path Towards Omni-Modality? LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:40.962864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:40.962864Z digest=sha256:bcf8ca8a531eaa8c67ba441bbf1ce5d239a6a037bcfddd7acd1fc58b6d39362a

Observation 500a1ece-fc59-4481-a16d-988122648280 · inbound

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding cites this paper.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.418612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.418612Z digest=sha256:025797d4ba00592c0d2b7288fba61d87d31c5613166755c241c7ee4b088e623d

Observation e09aaaa5-b745-416b-afaa-963f8e4c80b2 · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:46.132778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:46.132778Z digest=sha256:11d08b845f4fd8c9ce10d24b9499b99f1462f31146dcbd74d2c441f44cb7ec8c

Observation 02d568b3-97ea-4d33-bb01-cfcee5cc0e46 · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:42.916824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:42.916824Z digest=sha256:7664261ebd8d282ff61165a0ffa6f5c605998780a55aa2b3648c844facacf2f8

Observation c26d1035-6bcc-4792-800a-cc0fad81dce9 · inbound

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos cites this paper.

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:36.347575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:36.347575Z digest=sha256:79871e180e60bc472ed43faf4b94873f9e7d22be4c2bb31fc0d7dfdf5754c9d0

Observation bc43a348-2f3f-474c-a8be-147f491a82d5 · inbound

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning cites this paper.

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:37:15.762639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T11:36:36.687324Z digest=sha256:f513111860defd18931160f09b87c81964aea63966cc401ca9c823ce89c372f0

Observation 94c33dc7-c577-4677-a974-ecd3400ac393 · inbound

Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models cites this paper.

Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:20:39.039310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:20:39.039310Z digest=sha256:05adf3e43f55e258dcbec552ecc04bc99536d3a51c0d8029004a7fbf8734c1a6

Observation 3b014608-1e4e-4c3d-bae1-f2fe67a62cb5 · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.024929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.024929Z digest=sha256:f1a52876fe56b5b03b3149909553c18df2c63e340c65b73586c61dfba3500cf5

Observation b6f15c9d-18d7-4e73-9bcc-4a2501bf2bc6 · inbound

How Important are Videos for Training Video LLMs? cites this paper.

How Important are Videos for Training Video LLMs? LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.923531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.923531Z digest=sha256:8d41d9ba802d9ef561257bcb744d93dfd74de201b456b77c752107de7d46f731

Observation 0ac43119-00d2-4316-80ae-1ac7c5cda771 · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.315568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.315568Z digest=sha256:af11fcf1e7a8b029c479d4f0b5ddaec4c0381c86a155b3397a817655cbb36a50

Observation 5a25c88f-a587-46d3-8cd0-b7f0467bacc4 · inbound

Audio-Sync Video Generation with Multi-Stream Temporal Control cites this paper.

Audio-Sync Video Generation with Multi-Stream Temporal Control LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.527178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.527178Z digest=sha256:92fd442594017caa7bfe5cf1549043b9898e8e7139f6edca5ab72eb4abb20cdc

Observation 21d0b0a7-8220-479c-b170-bdf75b947399 · inbound

Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing cites this paper.

Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-17T04:58:10.283650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T04:58:10.202784Z digest=sha256:40e0a53668ba951f7704a6b3e555c108ad8b27839dfc5d606b0dc167f5c9c67a

Observation 28c45976-fa18-49ff-ad2e-50198112175c · inbound

GigaVideo-1: Advancing Video Generation via Automatic Feedback with 4 GPU-Hours Fine-Tuning cites this paper.

GigaVideo-1: Advancing Video Generation via Automatic Feedback with 4 GPU-Hours Fine-Tuning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 48

Resolution
malformed identifier
no resolver link, observed 2026-08-07T04:29:39.770200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:29:39.770200Z digest=sha256:fe334733cfba67d8a08fe26543d0f52f1f34ea4896f3c5ba6378ab51a1ae4576

Observation 9abb6d21-375d-41a4-a060-60b64bdc1499 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 145

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:51:15.968478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:0dfb9c505ab5c20d78dd9642bf75c418498c5b7a7c4b7e4510c848e4466d7098

Observation 2d883cd6-d552-4a7f-a565-86d5b4dc3436 · inbound

MMSearch-R1: Incentivizing LMMs to Search cites this paper.

MMSearch-R1: Incentivizing LMMs to Search LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:27:04.354589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T15:27:04.228144Z digest=sha256:5a8f9600470d679e642fda755b366a8953929529c1f4c4142f642a0d182e472b

Observation 0c307f2e-7cdd-4c08-b992-c4fc22cea03a · inbound

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs cites this paper.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:05.019965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:05.019965Z digest=sha256:ea1883077bd984bb53891728fd6933e22723741b906e113e7662420d5a4a6180

Observation d6e43932-032a-497b-a74f-48c42b77f00f · inbound

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning cites this paper.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:22.971564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:22.971564Z digest=sha256:cad4c5be3a4c4677b404c92abc78987bfeda8e0d24b5c55d95fe9f5192a06f5b

Observation 113072f7-b149-48b4-98a8-8e9f4e1254e5 · inbound

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams cites this paper.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.967515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.967515Z digest=sha256:a64dddaeb8033122c9d13b2002ed66307fddf0ce6dad32f84919705f1fee5728

Observation b3f8ab5d-e1e0-4dc3-a967-4662dd2d825b · inbound

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs cites this paper.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.623119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.623119Z digest=sha256:1c28aa39c33c69de2d939f7aa06adeb165026d5af04af4011531deeca5b2f55d

Observation 9d1e3cba-0d00-4e50-a1e7-d1699e39c8d3 · inbound

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization cites this paper.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:21.033536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:21.033536Z digest=sha256:39335402fb60002af53a126bc59dfc87ed928693a434917143d8cd1786d694b0

Observation 80aecfca-840e-43ab-83a4-dc31abeba65e · inbound

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models cites this paper.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 22

Resolution
malformed identifier
no resolver link, observed 2026-08-06T19:41:38.650060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.650060Z digest=sha256:c70ae4352ab1b748d37e2387b6d0cfd266382ae2ce33adc097e6c8fd5dbfed29

Observation f86124a1-b5ac-4052-a9e4-c56c69400990 · inbound

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling cites this paper.

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:37:06.449859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:37:06.449859Z digest=sha256:66cf5cf4bba001ff20039732ae60f06dea479f0e62e238d99bc8cf806c5e7550

Observation dd7c2631-a2ca-43c3-be08-e90da971380a · inbound

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs cites this paper.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.312891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.312891Z digest=sha256:a01dc08eac5dddfbc1304fd9a1a8a901523762b2958ddc2434f64f272f6758a6

Observation c59fd0ac-0659-49c7-88c2-f56a3efe1277 · inbound

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation cites this paper.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.749240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.749240Z digest=sha256:7db1b09ecdc40775213052886cd349311da015e680c6a297f70d4d630d69472f

Observation 904869a4-664e-4648-b434-788a107d910f · inbound

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs cites this paper.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.891685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.891685Z digest=sha256:34c213c5f872f0283c0f166f33295f441d992c7d0d100aefe3de110fbe63b5ba

Observation da2a2cfb-ce36-4457-8938-c5ab46742040 · inbound

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent cites this paper.

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:38:45.527724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:38:45.527724Z digest=sha256:2e8d4969cc1629ba56255944b3f746dacde79c6cfdd1c5e0e9212d8361823ca4

Observation 393861fd-ebd2-4293-882b-fc6cb6c5ca62 · inbound

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning cites this paper.

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:22:00.935931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T03:18:14.655384Z digest=sha256:46770c43c9dcfdbccfd7479a3d99cacaad7c0e605b723e34ce6da8d9c21937bb

Observation 8f43ed20-4d5f-404e-a453-983c4777fad8 · inbound

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos cites this paper.

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:17.441895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:17.441895Z digest=sha256:0915f3f49e0c8169b3211671192c24c9836d3a05f0c41ec07cfe95f03e39a9a4

Observation 67cb2719-8b81-48f7-adeb-ca73373abd9c · inbound

Object-centric Video Question Answering with Visual Grounding and Referring cites this paper.

Object-centric Video Question Answering with Visual Grounding and Referring LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.606964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.606964Z digest=sha256:d158ca01fe4b213ab3b585943cd8dc544f0887f56728e79a7fed26577bbfe241

Observation 542037b8-ea75-435c-8d9c-f7c41a471fe3 · inbound

LAVA: Language Driven Scalable and Versatile Traffic Video Analytics cites this paper.

LAVA: Language Driven Scalable and Versatile Traffic Video Analytics LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:32.511719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:05:32.511719Z digest=sha256:67a1fd6cb701286ce094dc819ec2ddd7f6271b61797f97b4a2537c9aefbe7895

Observation fc61fe64-0088-4c2f-b95f-349400158007 · inbound

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding cites this paper.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:37.791480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:37.791480Z digest=sha256:3f116dff50665b0bb43bd67d59d584c442b18950d988848a5d98c61c3bb5a7f5

Observation 8afe1ff4-893c-4b1b-a864-e129ce931309 · inbound

IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&A cites this paper.

IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&A LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T05:20:59.194993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:20:59.194993Z digest=sha256:83b066ac2723aa2e3f349bb239f833419df7d4dd410b0a5f52ab85f8f0a00c22

Observation 3ed91865-da65-4d6e-a1db-a804263ec9bb · inbound

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model cites this paper.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.759463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.759463Z digest=sha256:1a3ca50c719ab86e4bf7c39adbe204bce59381a589450c51d9bf74373caab48e

Observation 5a81c394-c805-4be3-bdd9-a4156c2916be · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-19T00:12:54.179224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:a26cb29d5f4a6d47efc598a32bc97282493ebe316f1e04f4065cea2f04e43fff

Observation c5df480c-2700-4203-8757-002aa20bbbf0 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:05:30.772069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:11e4bc9d844da0a1b6ac999ee5f22b23c4a5e52711e275b0adf86ffee0ac190f

Observation 90743981-e64d-4adc-b237-3e7eea28453e · inbound

CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model cites this paper.

CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:10.273170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:31:10.273170Z digest=sha256:3cb2147caef49963a8fcd4acfeb38b3a9e7c7e1bdfc2b7cd5802bcba75bf255d

Observation 22aecef3-2f8e-4ee9-b934-8a0d42a3adde · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:03.428316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:03.428316Z digest=sha256:69d2d38208758a591337b1f333d2264d9fe3c30c1025bdb20356868e2052f7f6

Observation e03a14a9-e0c2-4e9f-8a77-f03e31f886ec · inbound

DiscussLLM: Teaching Large Language Models When to Speak cites this paper.

DiscussLLM: Teaching Large Language Models When to Speak LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:10:42.320593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T22:09:03.740109Z digest=sha256:647029302484cf64d9e547c5ead8f3be0a95e012ca0dda429890406560e2f345

Observation 495fcdb7-43d6-4704-b7bb-7734549ef176 · inbound

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models cites this paper.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.536185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.536185Z digest=sha256:bf9217cd354e86555e768b3e56a76b65f5bc295129a052d7402b0d867b9c978a

Observation 96eb1452-6673-46ff-b315-def636bb0db6 · inbound

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment cites this paper.

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T13:54:52.875020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:54:52.875020Z digest=sha256:ea4d4d8571b55d7672cb243a39e459a82cfc97a98ac21a1702a569a4fc8da514

Observation 2394c300-a708-4382-b347-33c74c3f59ff · inbound

ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly cites this paper.

ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:06:50.183542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T20:03:26.144452Z digest=sha256:cdc393609910c067994fddeef6594f25099e7cdd45a3135ef2c3500c8dca3a9f

Observation 61073f3f-8553-4916-b6cb-ec7067afd79e · inbound

CAViAR: Critic-Augmented Video Agentic Reasoning cites this paper.

CAViAR: Critic-Augmented Video Agentic Reasoning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:36.408216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:36.408216Z digest=sha256:9cfd830a16baae4b77cc30261e26ede84da3e2f375e28491141575a73632ff22

Observation cc89babc-b975-4d58-9f7b-b9b38196ef01 · inbound

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models cites this paper.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.016518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.016518Z digest=sha256:28cb8572eb410edb4abb5f94a3888f2dd382549d06a0eb7e2df906bb9582b247

Observation 94e48f6b-027d-4227-996f-575df6d5d810 · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:36.989099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:36.989099Z digest=sha256:4f7f3cb7af7b4bf35776137cc4d00fef5bacfd3a5a54618a449d9beaf9e2c9f1

Observation 2a6b7a38-761c-4dc1-80c4-79e67f6e1854 · inbound

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding cites this paper.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:33.485739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:33.485739Z digest=sha256:d84f9738f849e2f51ac98450ff082cfa96eba87919a3957dbe8c5046d30a6966

Observation 6eaca711-a83a-4a5a-a29c-ad30b18f7463 · inbound

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding cites this paper.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:09.070998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:09.070998Z digest=sha256:0c53d076aeb44eeb3c86d6d012b64ae325354075326aa8eb95d65412ff66dea3

Observation db8bfc4c-c46f-4564-aac7-aeb8fe4723b5 · inbound

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding cites this paper.

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:20:22.888478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T22:19:36.366837Z digest=sha256:b69fbc00139d38def8cf75c66f40127eac73147b03bd2c3ea9b0227e4e8141c8

Observation 3f1a1a3e-e5ca-41fa-9fe4-c25a10aaeab6 · inbound

A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models cites this paper.

A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T21:31:37.065333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:31:37.065333Z digest=sha256:4a1a4d4a88dd2656bf805ec4f1fae72a3eec7595733ee2ab967a8a02f08b2c03

Observation 6a320d18-c23d-4e6d-8f39-76695c00133f · inbound

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs cites this paper.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:09.803468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:09.803468Z digest=sha256:2d1c52c0375f31023112d17a471d77670cf9f558eaeb8fefc4bb04015d6de12e

Observation 4d30f68c-5231-4f13-94d8-2390bf3c4d29 · inbound

Vision-Language Memory for Spatial Reasoning cites this paper.

Vision-Language Memory for Spatial Reasoning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:37.221295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:37.221295Z digest=sha256:f4a54af5bad16a6a03843fb8f5dea6979391bc97e7af77a5fa7fef0d567d1890

Observation 7dfb4cde-dffd-4ca0-800e-78e26826345d · inbound

Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos cites this paper.

Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T19:11:53.418782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:11:53.418782Z digest=sha256:228d823b18d7b1be7ebef6413f94b794374b8ba08eb8e3eeb399cc3b60d05fc6

Observation 67934772-bf1f-4c7c-b7cf-dd600c3842b5 · inbound

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding cites this paper.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:32.845533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:32.845533Z digest=sha256:215db8550f5fe5c36e68d0b8686b129b3ba4365efc80e292d9c2c3c09037616d

Observation 129f71a8-9fbb-4def-97d4-53c46bd94599 · inbound

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding cites this paper.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.601891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:9dd03ce9b3c570a9cc656ac987dadd0b4a0786a42c3b430c6f1467b9b392892d

Observation cd94a8c5-6131-4c9a-85df-3dc92179f7b5 · inbound

Efficient-VLN: A Simple yet Strong Baseline for Efficient Vision-Language Navigation cites this paper.

Efficient-VLN: A Simple yet Strong Baseline for Efficient Vision-Language Navigation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T17:18:20.894768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:18:20.894768Z digest=sha256:b7f5a52cee26b4a2351b6cf6be52b67e93da504b1117d3621fe7f686f295e4cf