Pith. sign in

Paper Citation Record · LEDGER

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

As of 11 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 56 inbound Pith citation observations for arXiv:2501.00574.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00574 v4

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T04:02:43.261543Z

measured 133 of 133 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 56 of 56 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:35.314074Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T16:49:57.265026Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact43
  • verified fuzzy32
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5ccd813f-5326-4801-b799-fd04a1a4215e · outbound

This paper cites Ht-step: Aligning instructional articles with how-to videos.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Ht-step: Aligning instructional articles with how-to videos

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.823544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:9181ce6c87909a8cf6b9b4dc72d32ad4c3cf35067216e65a6731c5568df98cb4

Observation e2f55cb3-6f44-4e36-bba9-2e1ad7668ada · outbound

This paper cites Qwen Technical Report.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Qwen Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.344967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:41e360b338c8dca141172b1fa9707abc69182079c21fd45d81b14d0f6f97dad1

Observation 779244d8-d237-4e0e-8210-a9eda9ce2e6d · outbound

This paper cites Qwen2.5-VL Technical Report.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.354940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:d087c1cb4cb2be0ff900720e85bfe0e3df0eb05362583f7fe15a57dde66f4cba

Observation d33aac72-3fad-4d6f-9401-54539ca3cf5d · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to- end retrieval.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Frozen in time: A joint video and image encoder for end-to- end retrieval

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.723681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:f7c3dfc3f19800fde5e2f1301991b7506ce257a300a7b48e0e50fa4a64631cf2

Observation 0d0606ad-ee11-41ca-8cc8-d55eb4b5395a · outbound

This paper cites Fuyu- 8b: A multimodal architecture for ai agents.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Fuyu- 8b: A multimodal architecture for ai agents

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.728025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:66dc556e9ecee7220e3dc64b88b7268ea86cc958d429e40e05eafc6db3d84540

Observation cd0461f2-3750-45fa-9588-9ae33a149ad8 · outbound

This paper cites Token Merging: Your ViT But Faster.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Token Merging: Your ViT But Faster

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.658088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:b310730dcbc16d87f1284762720a84f2ae4fd0f701e4760f986afe8e9c7d7bb9

Observation cbe8f332-5b99-4c76-8ab7-33c1155b89e5 · outbound

This paper cites HourVideo: 1-Hour Video-Language Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling HourVideo: 1-Hour Video-Language Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.666934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:7ebdc349f9a88a2e10f6fe095690a22cdc3d4f83a305ca8ad3f8aaf3264e76c6

Observation 7311a77c-640e-4ac4-9681-d6e43ef63cfe · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:2c028c225f23f6d1a5d44430c2806382befc97d4e626f173b000777523328def

Observation dd31a8fc-f16e-4dcd-bc52-ab91dbd3186a · outbound

This paper cites Efficient Large Multi-modal Models via Visual Context Compression.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Efficient Large Multi-modal Models via Visual Context Compression

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.682863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:b260886afc928b9204223f3a42e90dc57d1964bf6272420c21a7419c34f2883d

Observation f7fbad86-8795-4021-b276-98d6a6a2e4a5 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.759097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:516aa43f52ab6c23fb9a35e5d566cb79dc22d44c2ecb1441a6e81c5a85db5760

Observation fcd427c2-aacc-45cf-8243-6151d72161a6 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross- modality teachers.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Panda-70m: Captioning 70m videos with multiple cross- modality teachers

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.766772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:4929ecb960962aad5b58958e3cabe5e74d25f9bf3e7f21b6ab356483d6f02d9b

Observation db4489fd-e6fb-47cb-98ff-2ea6d77c59d0 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.692536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:51388d054351834e5b25ba1e73031f5448022c48105f51b1b087957eef7c28f5

Observation 600b90e3-8d40-47d6-b370-789f7f70ca11 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.700354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:366ce8eb90993dd456230595ba972b0313d344fa98e23812dfbdd622daebbe7a

Observation 2df3594c-4236-4286-87c5-991fdb42bb55 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.708478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:c94014c1c0577fc5854ac213fa1022cb0b8e8842c68a5b072cf731e4d991607e

Observation 23d85641-bcfe-4bb7-b319-6c101585e709 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.716663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:fbe972f4ec9baa87125efde39d1dd4b11d745bedb5101766eed2fd7b60d6dcbb

Observation 0afb140e-3b56-496b-8add-35ee60a28ad6 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.365637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:afae387895a9ff04cfe3766282f7e42cd742522f7e94e598c0eb09226e8d246f

Observation 09c80962-1b2b-4976-89df-22ce0a95b191 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.374291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:1c01fb35026229c34e0af9a987341e2e8ab4913c0c28b8e1eb4eec338e7d53b6

Observation c4f74501-f968-4cdd-bcb5-ca61ca3d2d20 · outbound

This paper cites Tall: Temporal activity localization via language query.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Tall: Temporal activity localization via language query

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.812808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:2724409d17731345fc65e430150961a18509faf4a31538a7e1beab30354ff697

Observation 5a6407dc-e93e-48ee-93ae-b696bfda9549 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Ego4d: Around the world in 3,000 hours of egocentric video

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.817923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:6c5f83b1e14689180011b9dc3d909b52c1f6f92ac660df450f10212086c44d3a

Observation afaf9f30-f354-46ec-a3e9-4f62459c9405 · outbound

This paper cites Online Video Understanding: OVBench and VideoChat-Online.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Online Video Understanding: OVBench and VideoChat-Online

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.383612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:e67a9b09bac3b46d9474b8c01cda2e50bee3dd5122a5270a4dcaee45dc231d76

Observation 8f02a67c-db75-4ea4-a44e-9ed3ff269651 · outbound

This paper cites Video recap: Recursive captioning of hour-long videos.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video recap: Recursive captioning of hour-long videos

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.827894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:2faf02190d2d98ec064b78fb8f36a9627c09c6052d4a9ac66afe11932347387a

Observation f41c7ea8-baff-483e-98b6-afeb880af81c · outbound

This paper cites MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.396104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:c5aec7f9989153ca80fe65944ebb60ae670f20d7c1c1890b4447d38686132334

Observation 77257d2f-6ebe-424c-99c5-ee5a0d83ca59 · outbound

This paper cites The Kinetics Human Action Video Dataset.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling The Kinetics Human Action Video Dataset

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.404573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:4a568ed19dbcb98cd01d9466698e9cfb61d754c4e8bef34648d74c7a50b6cff0

Observation b3a99a9e-d760-48bd-9bcf-c4f21aa785d9 · outbound

This paper cites OtterHD: A High-Resolution Multi-modality Model.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling OtterHD: A High-Resolution Multi-modality Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.414580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:f4739eb80fced5824b4d73144fa394d807bd5bc01ccd8c6020833e7feedccd85

Observation 06e36acb-10d6-4385-93c0-b600a5db02d5 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.422956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:4bbb88724b2cfefdbe4d6b6073b3b7a354f79d9bd80c7b70d10aefda64da093a

Observation 82223398-adaf-490f-b246-cc2e9b3989d8 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.430481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:a06e24f94a0b920e431b9d564261ee78891923b559fa3b9489958fadd899ea79

Observation 7c3bf99c-f3e6-405f-a78d-f220407e7788 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling VideoChat: Chat-Centric Video Understanding

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.438898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:4dc4b92202502bf65a568f583fd6776b96de2f97216ee67df21a632a36d612a0

Observation 05f1833f-7345-4679-abda-3b63faa6ddcb · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Unmasked teacher: Towards training-efficient video foundation models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.875177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:3bf48fec931b24a19f14ea0eff71a1fd50745fd8df527ec84e6d0d908cae6877

Observation 39366faf-09cb-4205-8c62-99262af4df53 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.879774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:996173a7f2f499e5e05d4100ea703a90bae0edd7ebf372d07670e4b9951e51d0

Observation 6a3411f0-f35b-4861-8d54-10f3e21d3b57 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Llama-vid: An image is worth 2 tokens in large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.884346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:4bd91a00568351b8ccfb66bf80efe5ae3add3f71f12499b6e861b3c89819e4f0

Observation e03994eb-65ed-4c9b-acc0-f9a966a39006 · outbound

This paper cites Video-llava: Learning united visual repre- sentation by alignment before projection.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-llava: Learning united visual repre- sentation by alignment before projection

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.889267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:495c0a2d153007b5199403ac8c5f7262131b6508580b69230a06c72d5945bb26

Observation b19cb18b-dc47-4443-aa35-f12b1b8382c4 · outbound

This paper cites Microsoft coco: Common objects in context.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Microsoft coco: Common objects in context

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.894123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:54aacd4c9ab10204f6daad5563d322c322dd621391af3940675dad673f0abe73

Observation ca620c43-9dda-49f8-8438-ded2f33821b0 · outbound

This paper cites Visual instruction tuning.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Visual instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.898560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:9d4dd87329388077613d213625c488d4a0da0309b9dcde75a4473b8bb84eaccb

Observation fc2efdfa-89c7-4199-81f7-5ccd0e43f5b6 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.448546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:c683d215ab62cee71d0bc444920d4234ef7735d0d93a22e17686f11b9b160efb

Observation fb0f0cfd-d02a-46eb-a989-78fc83bdc244 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.456872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:7e48f8a133ee691f892bc864be2a9062b4bc1c4760c0fff5baa7c4f00a3a513c

Observation 0dd1d071-ab76-4ad0-b384-85327ae7f995 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.912859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:5acfbf8ffdbad1d1f755ff4c377ea9918405c5b30654c8691d6ef6fd7919cb46

Observation b450f8da-bc02-4493-9f13-e6e97c94b1eb · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.733562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:bec3c98a4f07703bc4e1c4f64a31fca997823d2422a5dffac1e6e2568b66be71

Observation 106110a7-88ec-446f-8e39-e87aa3da748a · outbound

This paper cites Spoken moments: Learning joint audio-visual representations from video descriptions.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Spoken moments: Learning joint audio-visual representations from video descriptions

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.739205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:5c1322a0da01b6faaa2e2bfd010742b39addeaa5d5dd8ee62bb8693143c398a3

Observation a653d899-4963-46a4-9e17-1233011a2c27 · outbound

This paper cites GPT-4 Technical Report.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling GPT-4 Technical Report

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.465573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:dfca2b6f2acf27d8c145242e96b8eaa8f438ec8a99baad59e4cf596a6511017b

Observation f899955c-8817-46e0-8d7e-26390ecd5e8c · outbound

This paper cites an unresolved cited work.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-18T04:02:43.752165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:1b2b21346f7f2d15b721c71e79b8a40ff870749dd73f1c875c2062e744206f47

Observation 9374501d-dfe7-4cda-b36b-7a36fb7f9513 · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Perception test: A diagnostic benchmark for multimodal video models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.774711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:31bcf7f892a86e9e4f4d80dff5577ff5eb27ae94f0877f4692b02a8ba8b0b0fd

Observation b843d29b-a458-4567-8ccd-fd94e0d8e4d6 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.780423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:b7badbaa407b79b15071a135d2987cde5427b6676df7fe361ea90a759531b6ff

Observation f9b9e478-9dbf-4b84-83e1-faebb67969f9 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.474946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:0d7da5848f19f9522b43f7dfc54d63b220fbe75d550f98d0d114714926e0f7c5

Observation 00d5b85c-c553-4069-8408-12eac93fca39 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T04:02:43.484634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:d453d815c2f5b55288919463354a9bff74e84edaafb288124f966e7b4d3d55e9

Observation e74b5432-69be-4af7-8dd0-32eccc895a00 · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.799560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:c0002aa9c2451a2e3ca3a7b5f145ced35aa32c82d2c8d2ae7278d4c7460258dd

Observation 69a01eff-aa45-4564-9c4f-28685cf5b52c · outbound

This paper cites Sharegemini: Scaling up video caption data for multi- modal large language models.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Sharegemini: Scaling up video caption data for multi- modal large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.808123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:45abc03662ce8dfd0e5ce1ebc8c64f12ac14aa7258bc1d7ded976362d611631c

Observation 424c255a-4d59-447f-87fa-c36613e5708f · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.493663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:00b6a4ef270dd4f98461d92aff85594c4ffb109bc78c78967a5046f20da18263

Observation 731ff6fb-eb7b-48c7-b925-bcf2ff8d0414 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.503594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:be05c4406b06af52bbd28cab5c381b0643d93e03f753555a397ff15bd1cfea3a

Observation ca6154b8-b0b9-4640-9600-6a93a4c81256 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Moviechat: From dense token to sparse memory for long video understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.847968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:80b2c316fd6c4aad9905fc939abe15fff2bc14f4a006740353e6e411820ec2b7

Observation 450a0c10-27fe-488c-addc-f2015cf948b4 · outbound

This paper cites Koala: Key frame-conditioned long video-llm.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Koala: Key frame-conditioned long video-llm

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.853509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:74bdcb45e15c442caa7b7077c4b83d211d3948ff68d9c8b0f368f2238525c2b2

Observation a6425a9b-0607-4e84-91d1-a274210e554b · outbound

This paper cites COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.513310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:712d66f1d00c59b9fa6aa0df169ef2ce6e9a8562c96458ff7d849993d1ac3b56

Observation 5f8b2268-7e4a-457a-8f70-84c0d0615742 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.522302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:7c4df9a9c7cb8ce802322cc1a6404e5e2a69edc0e63e0fba509f385f2454c4f2

Observation 5760c3b7-d10b-43f6-85ce-d9520ab46306 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LVBench: An Extreme Long Video Understanding Benchmark

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.239348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:67af433e474a2d669ece6576da4dad2f02ef69ebc77ee8d0b2cbd81245e9acfc

Observation d47e1344-6c36-4fc0-b57d-eef20f1ac2e4 · outbound

This paper cites Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.537223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:3d202d6effec26bcb75f87e2aeda0eba995367b5e9431ae08dfd477fd6fbd632

Observation 9489c8d4-380e-4e33-aee0-489b66a22283 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.544903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:43d116a78c6f4794a279cdfbee8d41aa2696a31a39d850aeb4a15ecb5577d9fd

Observation 600ba87b-ff0d-4e3e-a1e7-027435d4e235 · outbound

This paper cites Internvideo2: Scaling video foundation models for multimodal video understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Internvideo2: Scaling video foundation models for multimodal video understanding

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.744790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:1c5b07b9bd0016487bf7332085e094cd7779fd9fb2589ea1aea356e526764bb5

Observation ba361a7a-a74e-498e-8ba7-1aab43771bf1 · outbound

This paper cites Visual Context Window Extension: A New Perspective for Long Video Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Visual Context Window Extension: A New Perspective for Long Video Understanding

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.552982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:60898ff916b2ad68425fc113b0bd187983d92596546080474400978106f21a29

Observation ac284730-eee2-40f8-ab43-a6799a557d74 · outbound

This paper cites Longvlm: Efficient long video understand- ing via large language models.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Longvlm: Efficient long video understand- ing via large language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.792603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:6017ede6cbfac14bcc20462e90f949532fefa9fa3dd413f565d0b291602d5f93

Observation 4de7f155-5d02-4109-9a9a-0c302c4cdc64 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.332710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:101d23e36747a232b5e5ac140021c0f6e2e792721261cb5a17cf0f968b93c788

Observation 6feda575-e467-4932-b44d-2da9d80ef809 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.561759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:7e3bad966f707a3937b83371301cd435492b1ec1d8dadd29143408c13e5c2a18

Observation a100e30c-9c51-4c45-9148-50e2b78a265e · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.569786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:057568ae1604a754900190ea68ae076611225ce96d26b18f2328bb77fc56520c

Observation b9332801-9f36-4c56-b77f-a8364c0f149c · outbound

This paper cites Advanc- ing high-resolution video-language representation with large- scale video transcriptions.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Advanc- ing high-resolution video-language representation with large- scale video transcriptions

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.864787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:54632a5d067d4437272dc8f04457cbb16008cf7551b70be214b71a01b755726a

Observation 5cfa7a2c-3401-4e26-b69e-c5672311255c · outbound

This paper cites Vript: A Video Is Worth Thousands of Words.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Vript: A Video Is Worth Thousands of Words

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.578826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:201a1800b009babbbdf29c4b73ccc58fd420acaeea34442c2a2e5033d7d75d43

Observation 97ab7361-9b0b-40e7-9f20-7f65a83d98de · outbound

This paper cites TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.587545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:df0d060518391a035d0f86cff842a0e29515ee427b787d753b63327d8f0d37e8

Observation 6d06eef6-2d31-4fd3-a6bc-2736bdb5bb57 · outbound

This paper cites Sigmoid loss for language image pre-training.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Sigmoid loss for language image pre-training

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.907153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:82e6777df0e0f6938eecf54712efd89062c8c7a87f4e03673c21a3bcac353bb7

Observation 47e71157-dafe-4069-b911-66e8c6f76dcd · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.594914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:ad789b1b53860ec24dcd60e2ff21459257b5842ac225886189ebbfe659797ac2

Observation 428835ff-04fc-425d-834c-bd2a36c4e492 · outbound

This paper cites LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.607191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:e721e98ae46461854171132e068b19ddda002a05b43581652d1a6feaadf41c8d

Observation e272bae0-47ed-44e8-9638-54b01db80967 · outbound

This paper cites Long Context Transfer from Language to Vision.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Long Context Transfer from Language to Vision

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.616047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:2c5e753f8bad23a8f559a58d661530225c13101b40e507dc62808ddfed70f9f3

Observation 869fdb9a-b101-4c3a-b5bd-5fdbd436cb60 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.624687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:db48e9f1606c3907575bf2f0c9c23a02b42f80af61f20fcd7b8fc2ddd94b470f

Observation 53a85dce-09a4-4b2a-b5c5-1a81b15478ab · outbound

This paper cites Llava- next: A strong zero-shot video understanding model.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Llava- next: A strong zero-shot video understanding model

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.902683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:3f5500fef1aa6554955d7015a66941178e32ba3abd705bf1314eef8c5c99df45

Observation 84022ade-c828-472f-a4f0-b9028aa05dfa · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.632932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:60b8b0e440557bfa53a7b05a9b9953e2953e5eda6010170b44c345b81bdd884b

Observation 93d93d03-bac4-4091-959d-50669fe75564 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.641934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:bec75aa6e6168b0569b9527709cce72b80bfcbca852668dadd8529b37e64987f

Observation 9d9007bb-eb1d-400d-ad90-cecca5f3ca5a · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling MLVU: Benchmarking Multi-task Long Video Understanding

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.650265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:032373cf3cdd90c2525cd4f63fb70fad8d2142aeb87390117ae4a08930ff680a

Observation ed625540-0776-4336-b00a-cf1276dcd26e · outbound

This paper cites Visual Dropout in LLM Visual token redundancy in LLM inference.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Visual Dropout in LLM Visual token redundancy in LLM inference

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.870297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:79a5475df88fe1b6f1eaf78507bf4e9aeef49053a02532f68c302e422f231c11

Observation 0cdb5f83-0a33-42e7-abbc-9a728e983447 · outbound

This paper cites Video-Language Connectors As shown in Fig.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-Language Connectors As shown in Fig

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.785923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:c79c63843811aa320e1b5f4dd94255d385dd599cfb24ed03a837a88b69a0581a

Observation fb7abffb-5425-4f12-8bfd-eaa32f4854c0 · outbound

This paper cites We provide details of the data construc- tion pipeline for each dataset as follows.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling We provide details of the data construc- tion pipeline for each dataset as follows

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.835348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:b2f8828edb90e1fa3b2f940761e6e422a1422f3f3ca535ea3c199841fa6fd02c

Observation d0c7ab25-93cb-4bd3-9c63-697861feefa1 · outbound

This paper cites 11 and 12) and long video understanding ( Figs.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling 11 and 12) and long video understanding ( Figs

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.858779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:7ccb045dd7a45e035f1bae9df2ade33bd74b5a61928de621b2460bcfcc36c3f8

Pith citing papers

Observation 59801fb1-50fb-4317-8a32-1acef7e4adef · inbound

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding cites this paper.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:32da9edd3cc8c17ced1eccbfc96cf8720af370ffc7213899430785e68254f50f

Observation 304f2a4f-3440-41b9-a78a-5c5e63b912ff · inbound

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning cites this paper.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:49d2c1445fd3f307c9a40b9d96f1c6466176caf26dda01e358801479481a3a65

Observation 0e55e230-5270-4ea2-ba9d-10a36c347294 · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.968075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.968075Z digest=sha256:fffad2fb19ecab30593ab00c90c8f3a1c83e596f6f1501b745a65422ab939596

Observation 7cbb5882-cc3d-430f-b987-5b744e65f3ef · inbound

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought cites this paper.

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:08.944838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:08.944838Z digest=sha256:be7d684deb0767265ab91949040cf12956defb03fba502f495a840540e9db248

Observation 1746ded8-6795-49d0-871f-472260bf428f · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.621901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.621901Z digest=sha256:ce7cc3194a7584516ccfa2f16ce3092bcd2e9a67964c1e42edd69c308e908e32

Observation bdef7f07-c1c2-49ed-8e95-b25717da792e · inbound

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs cites this paper.

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:07:15.433789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T11:03:59.222849Z digest=sha256:b32f715df5164c80f4d301c1e4528d76cbc5c19aa004ba56360ff3cdffac9a36

Observation 464c5e5e-07f5-46f0-ae02-fcf6e2b9ff09 · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.894485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.894485Z digest=sha256:aa5b1e1b0ce85208b7f9f69676a67f133d96ede5f896d9a9193fd59616986a53

Observation bdf7c58a-339b-43ed-9661-ce81ff93e999 · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.438610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.438610Z digest=sha256:9383ff158f6386c7e5ee66fc0aa52bde8a15020484cabdac8c62c1f74b79929f

Observation 632099fc-9a4c-47e4-b513-325dddaaf3e7 · inbound

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization cites this paper.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.138597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.138597Z digest=sha256:2399be0e96fefc8187c5b96d9aab2c0070ab7bfc6b46e7c332f40db548064a3d

Observation 97ac225c-a44b-48e7-b367-d1e266ecfdf5 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.714751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.714751Z digest=sha256:c0a50a1f66d72071326a2f0e1da0bf28ae3b15daad0b9242ce659cf886a35b67

Observation 215cecf1-67bb-41ca-83fc-6a30a3038aa5 · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:39.008817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:39.008817Z digest=sha256:2c590706f66172d8a22951411485a5518f50d52d84d4d613ba3de0fafedba03c

Observation ceb20892-77c2-4ab6-a4b0-7359bf440543 · inbound

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding cites this paper.

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T22:01:25.950345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:01:25.950345Z digest=sha256:4c9b92d14e602898df7ec8ed9a78e58432b3d8e32a0ea0edf4e383fc2f4f18be

Observation 7022ef98-3fa6-4b6c-b519-e0a1146e13de · inbound

Cambrian-S: Towards Spatial Supersensing in Video cites this paper.

Cambrian-S: Towards Spatial Supersensing in Video VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T03:46:04.363500Z digest=sha256:b25aadbe4e09e418f852246256f79fe6d846af7770f8c107f5b29b4bb7b4d5e5

Observation 001a86da-c99c-44df-80e1-33be32c08153 · inbound

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding cites this paper.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:26.641945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:26.641945Z digest=sha256:7531a0be9cd48d181c62ef53f0cb3012df36b5114dfa7b72f728c60750e6cab6

Observation b8d31a43-6c44-42a5-9ac9-39649bcdf1c2 · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:39.557562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:39.557562Z digest=sha256:1d5ec71470b577f731b9d2cb3e77b25fd9363c80e1a5dadfaa8561db1c293f28

Observation 8b10fd8d-e836-401e-8b70-c49303de9192 · inbound

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding cites this paper.

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T04:21:29.526008Z digest=sha256:d769a742756307ab4da6d523b7772dbb67d6760a2be80830d9efa1d218d5d2b8

Observation 5f773962-080e-4f81-b5a2-6269198da281 · inbound

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning cites this paper.

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T12:17:42.135851Z digest=sha256:8d2dc8cd3e55235a2d07aa9f2e386dead976a1663229c566492a077cc34c2d6f

Observation caf712ff-39f2-405b-b02d-3c7e44b8e3ab · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:8bd241263aa91143ddbd215eb5480ef8028c3936344b596d2e30d77b7cc3c4ed

Observation 755d7aae-8794-410d-9ab5-18f342f58639 · inbound

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting cites this paper.

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T15:29:31.834566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:29:31.834566Z digest=sha256:50a27a403b386d3d79fdbee32ae93f06a79391ef9c6ef45a97b7bc58b6ba3f6f

Observation 9bf6b2fc-d160-4eda-b8b9-fa7e050e45f8 · inbound

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning cites this paper.

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T19:46:16.975267Z digest=sha256:cccd1088bbb0ad9e45daf45fdaaec7dc3dfb7c11709efa012b564b5d8375c8db

Observation 6f1f7f69-f505-4b72-878e-8402334e5135 · inbound

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding cites this paper.

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T17:21:47.439019Z digest=sha256:499af525319fc729e1fbc2f2fa5aeaf81978d74eef4eab1dbd2478ac3adb6bf5

Observation e2f84104-be0e-42d7-a0af-20076b606867 · inbound

Small Vision-Language Models are Smart Compressors for Long Video Understanding cites this paper.

Small Vision-Language Models are Smart Compressors for Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:e797c8123c9bde356456bcacfd875b73bf22dd2533a6b1086670acbf1d363ef9

Observation 67be4ce6-9f37-4586-b1ca-299ab7ba4e77 · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:6c06f4b287609fbda82df6da4dc47cea958670d51f2d3f36c6f6b41f4c28f5a5

Observation eb49f453-1f16-4af0-8773-198ffa504596 · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:b932641a13eff4d56ff9d64562a616ab39feb7f2157772adcf8b0867f6bbe8a5

Observation 2f95654f-aee8-4488-b844-0c8567bbfff0 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:31:26.325118Z digest=sha256:a75d0450622f4334842ef9ae4f977db80ba1cd0f511254a6f8bc20570abc389d

Observation d3a281b8-a4c6-4ec7-9e22-7b5e6f809ea1 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.161913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:4f440ed852e5745d01802f2659e900f6da4d34e7d3aa0d715e53ad47d6af5d4d

Observation 1318ce42-fc44-483c-a7c7-eb48232047fa · inbound

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding cites this paper.

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T16:51:43.783133Z digest=sha256:ed99756ef3476b969630af90f5d1cb8b5d7da46573f8f19fa88a99bfae747740

Observation ce784b1c-a6ce-413a-ac62-537513023795 · inbound

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading cites this paper.

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T16:10:22.588945Z digest=sha256:c3a88b8c18f55c2ab61f108335b61b3638d4889d263929451246c832ba987645

Observation 52e12638-e132-44a5-8b65-f8bfc00e0044 · inbound

MedHorizon: Towards Long-context Medical Video Understanding in the Wild cites this paper.

MedHorizon: Towards Long-context Medical Video Understanding in the Wild VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 97

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-08T12:28:35.008604Z digest=sha256:3b1e0a44a449e75160319b3251f43c65c3b19d704b26cbd9723ca638d4de579f

Observation 4adeab72-40fb-4e0b-b7e2-130439272b81 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:13:21.487431Z digest=sha256:bd5dc7521e044fcffc96e4e39a0e9eb3a1dd3ba6478f49d8fabbe36f3e2ee92c

Observation 6ec7e470-e915-41ae-adcc-28dc93d42da5 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T06:53:42.726350Z digest=sha256:07ac6faef3ba65832a9017e5cb369a734d027199dca40f6c2f6cdf2577a74efe

Observation 66fdcaa3-4c5c-4457-9147-816527f507fb · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:3b2a11c17fa80b5d4744fad26b1be251524718a323210c53f29f4a434b170d3a

Observation fe731a24-2731-4709-a6b4-008686feedc5 · inbound

StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video cites this paper.

StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:33:50.726989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T23:32:38.332828Z digest=sha256:8717db73d30a22001b7f42c51fc77326f6aa66ae3eac747276ddfbaeb953dcea

Observation ec812b9c-1654-49bd-9cb1-e7a4964b0231 · inbound

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning cites this paper.

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:13:24.794852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T15:12:00.408851Z digest=sha256:be53bcd7217b80c61d1bb03b7e1ce339599f20dbe98e0a9f98faf89cca910eda

Observation f2146df8-67d7-4947-ac5f-ccda5531421b · inbound

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues cites this paper.

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:14:42.412095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T07:13:43.716510Z digest=sha256:1f2e2de9504635cece55ab847b95f78f69f1526da9344700f5c63bf0e23a6c02

Observation fb4fd70d-585c-48b0-9d6e-22d13ada3ef0 · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.641616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:a38377fede3d8bfcc382e1c69b611560bf784f2c9875d4ed02c850da98f6ca39

Observation fd7d14f7-d3f6-4de2-91e8-9d9a6a2997d2 · inbound

Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification cites this paper.

Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:33:28.319648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T13:26:21.016303Z digest=sha256:79067215cb68f0e4c1c4f6f51ef5ec182629a92f1ee00fd93972e5472f82af7d

Observation 43226218-c24d-4428-810e-2184c620bc5c · inbound

Linear Scaling Video VLMs for Long Video Understanding cites this paper.

Linear Scaling Video VLMs for Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T23:02:46.240303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T23:00:11.246232Z digest=sha256:2639b58a5db29353a2a370a5ca1e5a1a1f3d5c47131627eb21707ef16f154b10

Observation 7c215330-0ec7-4050-a3a4-90e738110d57 · inbound

AdaCodec: A Predictive Visual Code for Video MLLMs cites this paper.

AdaCodec: A Predictive Visual Code for Video MLLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 74

Resolution
malformed identifier
local_arxiv, observed 2026-07-01T22:26:18.177531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T15:20:48.248576Z digest=sha256:05ff856d62112edd0bf7bf4eafad67b391d06ca84bdb515c0b657014010eb1f1

Observation ce149296-1715-4fee-8bf7-78373f24a336 · inbound

Towards One-to-Many Temporal Grounding cites this paper.

Towards One-to-Many Temporal Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:16:57.727597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T02:11:48.455492Z digest=sha256:b4c99f37257b8b7a9c54c65b17ed906e57d096fb7f02ad444dfa036974857308

Observation 9be2b6a0-608e-4891-b064-1e0a1a0a2cf6 · inbound

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset cites this paper.

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 110

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:26:57.204162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T02:05:47.810096Z digest=sha256:a17c89a095cfd3518b9d1c7ad0224595c512b4cf78ff65e3366975167327a193

Observation a32cd9a1-2731-47e8-8f02-3a0d2d0a4957 · inbound

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning cites this paper.

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:56:47.341319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T06:33:32.090913Z digest=sha256:3381d55f2d5d6c8e3556905682c436acc2bd78b71268af7d9beef2b6199d7aba

Observation ecee5536-b925-44dd-9aa9-f5e695be6697 · inbound

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models cites this paper.

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:17:09.563522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T22:49:18.420491Z digest=sha256:1e14bd9170286ee99723a8795c35efa0a232cc3dbbe7cc136933beefbb36621d

Observation fa29733a-7da2-4aac-bdbf-d5e70755b7cb · inbound

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding cites this paper.

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:29.050040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T17:35:07.667016Z digest=sha256:16acbe0e5524fdf1dae207cb3dcf9904594dd4bc72bb7944f2403bca83dbaede

Observation 4b1a77c9-1b4d-496e-9916-5f4f058923b4 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 168

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:38:28.858320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:aed7eaf618023bdc2f6185ad489cd8f0653cc2504e2b8ee967f5d1c1bdec60fd

Observation a3eb1385-31d6-423c-b654-d54bef28fc13 · inbound

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference cites this paper.

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:19:29.917353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-26T18:17:53.013043Z digest=sha256:48c56dbfc82a064c17bfcdfa46d9ff3c61a3aaac6ec1cc7b1971c68c94338b1b

Observation 2df7dcd3-e675-474f-858a-cee988064085 · inbound

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression cites this paper.

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:49:57.266359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T00:16:56.174638Z digest=sha256:324dced9e83c728cff61c91cbf5a00418f66234c26811e2818fc86662077b98a

Observation 8ecf9277-e14b-4f6e-a53e-1cb3ff4a078f · inbound

Learning to Deny: Action Denial in Multimodal Large Language Models cites this paper.

Learning to Deny: Action Denial in Multimodal Large Language Models VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T09:45:39.610427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-01T06:21:09.996386Z digest=sha256:dbfc542f483ad112f16331ef41e53fab08aaebedfb7e4e746a2ded225cd78d74

Observation cde95483-51b5-46fa-9610-b56e3f11fa18 · inbound

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning cites this paper.

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T05:48:27.255331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:48:27.255331Z digest=sha256:f80c0b95ae7be6fff2bba7e1c026ffac0749f6eca0674e05590a5bdfa7233de6

Observation d9641bb0-bc0e-4aa9-8d06-47f68c5dbe93 · inbound

DynTrace: Tracking Dynamic Object Evidence for 4D Spatio-Temporal Reasoning in MLLMs cites this paper.

DynTrace: Tracking Dynamic Object Evidence for 4D Spatio-Temporal Reasoning in MLLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T06:34:27.665052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:34:27.665052Z digest=sha256:a551f942fe8b1807d1863dbc3ab501cb76a276c02f2e48f07bed6dc66dfbea26

Observation e7e61b8a-07a7-4a88-98f4-a78be078cb89 · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:40.495393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:40.495393Z digest=sha256:cead89e4ab0f17348c295f7c9787fb8918c4d7ba37f1db640a57663004f8246d

Observation b7669212-e82e-4bbf-b512-faae1e162957 · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:14.015999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:14.015999Z digest=sha256:1d90ee4d6a1e8a1a962f1f29a58752e73e9a5f4835cf369039ec8ebfb6f4f52d

Observation b3bd2182-e873-4443-8b8e-d01840d6f635 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.750593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.750593Z digest=sha256:a29ab7e602a936b1e75246e8019355b3c8293d2586786ec5a24bbf7b7614febb

Observation d3b7f910-a70e-4475-b4e6-7f994cd1d712 · inbound

Long-Horizon Embodied Decision-Making via Multimodal Memory Compression cites this paper.

Long-Horizon Embodied Decision-Making via Multimodal Memory Compression VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T00:13:22.601162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:13:22.601162Z digest=sha256:71edd9d31c76f9964297d031de81b37705022054594748ac96ee833913a8e383

Observation 98abc67f-d45e-4fbe-9411-15972855c281 · inbound

CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding cites this paper.

CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T15:43:54.334942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:43:54.334942Z digest=sha256:04e2a0ac5ed35695f00fa132fcbeb59e860c91faa5bd18fe108733bd442ead01

Observation 8651f988-a6eb-478d-9397-9933409b53f1 · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.314074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.314074Z digest=sha256:fcde7eb6962f52b8036bde6bce30daf638c1dd43e66fa55ce26513ad63b60364