Pith. sign in

Paper Citation Record · LEDGER

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

As of 23 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 67 inbound Pith citation observations for arXiv:2501.00574.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00574 v4

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T04:02:43.261543Z

measured 144 of 144 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 67 of 67 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.594042Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T16:49:57.265026Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact43
  • verified fuzzy32
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5ccd813f-5326-4801-b799-fd04a1a4215e · outbound

This paper cites Ht-step: Aligning instructional articles with how-to videos.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Ht-step: Aligning instructional articles with how-to videos

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.823544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:dc3fc03c930ff1ea19307c51472a9413468a66cb3f67b1cdab6f1dcffb070140

Observation e2f55cb3-6f44-4e36-bba9-2e1ad7668ada · outbound

This paper cites Qwen Technical Report.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Qwen Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.344967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:9e8b31fa335853170b00518de340c5d08ab10c3852a15ef4b159610a6bd35230

Observation 779244d8-d237-4e0e-8210-a9eda9ce2e6d · outbound

This paper cites Qwen2.5-VL Technical Report.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.354940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:3f01083b7b38ab31aeb64bed72954bc804bb044487b6b868e538507b16d56233

Observation d33aac72-3fad-4d6f-9401-54539ca3cf5d · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to- end retrieval.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Frozen in time: A joint video and image encoder for end-to- end retrieval

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.723681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:8137a6a028fe07bf934268fc23e9419c6c14c0352888164e73951be07325208c

Observation 0d0606ad-ee11-41ca-8cc8-d55eb4b5395a · outbound

This paper cites Fuyu- 8b: A multimodal architecture for ai agents.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Fuyu- 8b: A multimodal architecture for ai agents

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.728025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:8c70346082e65e82d23b525837ffae2d1d5abe711cb79da29ebca86f0b959c3a

Observation cd0461f2-3750-45fa-9588-9ae33a149ad8 · outbound

This paper cites Token Merging: Your ViT But Faster.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Token Merging: Your ViT But Faster

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.658088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:01ef44807382349efc4012e66f880e889e87ecdd787ac65fe858defe08068a4a

Observation cbe8f332-5b99-4c76-8ab7-33c1155b89e5 · outbound

This paper cites HourVideo: 1-Hour Video-Language Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling HourVideo: 1-Hour Video-Language Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.666934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:5062fd0c9db58d341afa292310626196d6e2994b60a0bb967c22c79ae4fb2dcc

Observation 7311a77c-640e-4ac4-9681-d6e43ef63cfe · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:e31a9812ab91bcef64d0ae1ac66f99027c9f50863b43872d28c11215109c51ed

Observation dd31a8fc-f16e-4dcd-bc52-ab91dbd3186a · outbound

This paper cites Efficient Large Multi-modal Models via Visual Context Compression.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Efficient Large Multi-modal Models via Visual Context Compression

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.682863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:c0495da8ab36c852ddcd27a56df94c749f2fc3d4c78a293bf57b3bd8e140adf7

Observation f7fbad86-8795-4021-b276-98d6a6a2e4a5 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.759097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:9447bdd8f22a3e27ebfbf469f77650c6d1aa3ba862aac56d3a0433f9f2a6bdee

Observation fcd427c2-aacc-45cf-8243-6151d72161a6 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross- modality teachers.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Panda-70m: Captioning 70m videos with multiple cross- modality teachers

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.766772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:459acd3d4a82ae6130bc82d63f92f8a12ba0756a2012c673a6056c06498101f1

Observation db4489fd-e6fb-47cb-98ff-2ea6d77c59d0 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.692536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:88de3fe78599a29bf6d4f3f57787dc57d3a25415e5b3b20ef54ac6458650c445

Observation 600b90e3-8d40-47d6-b370-789f7f70ca11 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.700354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:feecdd025d89d7cdd9cd0184a55ff5a786a41826e7e378f7254e7582d5725780

Observation 2df3594c-4236-4286-87c5-991fdb42bb55 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.708478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:8a18872a02cdfc30ca40ab7e71de9c15984c9054040c0ca9b5d874e49a42ce9c

Observation 23d85641-bcfe-4bb7-b319-6c101585e709 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.716663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:44dec07e3db24f2ca8428bc1321e3468a00436e3b5247001b3f7be3f07e3ca20

Observation 0afb140e-3b56-496b-8add-35ee60a28ad6 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.365637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:7700832697a00dc2e665f13ed6ba76c72ac1635ce7d8ba84962bb832bb28eac5

Observation 09c80962-1b2b-4976-89df-22ce0a95b191 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.374291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:9db37ac66fdd4c235782df4940068e5bdde0711f30d7d2a9e6fdb5e337aea775

Observation c4f74501-f968-4cdd-bcb5-ca61ca3d2d20 · outbound

This paper cites Tall: Temporal activity localization via language query.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Tall: Temporal activity localization via language query

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.812808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:b193fb34d302f3d2c4b389f03efbefed9076086935c99200654fde873687cee3

Observation 5a6407dc-e93e-48ee-93ae-b696bfda9549 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Ego4d: Around the world in 3,000 hours of egocentric video

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.817923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:291ac0e821731ab94537c6da4fdb43899948a70efc327538c7f2da05973249b7

Observation afaf9f30-f354-46ec-a3e9-4f62459c9405 · outbound

This paper cites Online Video Understanding: OVBench and VideoChat-Online.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Online Video Understanding: OVBench and VideoChat-Online

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.383612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:3e496763356e7c3e91fb6f2e816fc8868e73afc061c9c9c6234e92fe97cb9917

Observation 8f02a67c-db75-4ea4-a44e-9ed3ff269651 · outbound

This paper cites Video recap: Recursive captioning of hour-long videos.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video recap: Recursive captioning of hour-long videos

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.827894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:f6076176885d98fbd9da868acf365ae47c63624588f5d704f28db35b1dfd5720

Observation f41c7ea8-baff-483e-98b6-afeb880af81c · outbound

This paper cites MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.396104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:c2fb371a6a4465497171b1146ef726e7075d5f6b891d383f3eee084ae58a7276

Observation 77257d2f-6ebe-424c-99c5-ee5a0d83ca59 · outbound

This paper cites The Kinetics Human Action Video Dataset.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling The Kinetics Human Action Video Dataset

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.404573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:7af9af6c5c92042f1a407ab12844f40b7c49a8b2ab1e2d20fb1c4e0ba0b8b3e7

Observation b3a99a9e-d760-48bd-9bcf-c4f21aa785d9 · outbound

This paper cites OtterHD: A High-Resolution Multi-modality Model.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling OtterHD: A High-Resolution Multi-modality Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.414580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:4d3c73b501e43cfc2fd9a35947dd0382a2e7d6bf70eae2d95a6ccb54737c397a

Observation 06e36acb-10d6-4385-93c0-b600a5db02d5 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.422956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:27f05ef71a49c243ff76ebb300c3724440fd66fb785697da868fc3b8e7ff20d6

Observation 82223398-adaf-490f-b246-cc2e9b3989d8 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.430481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:c86c2454fbeda21802a1720184b307390da9daeec4046f070fce73be9231420b

Observation 7c3bf99c-f3e6-405f-a78d-f220407e7788 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling VideoChat: Chat-Centric Video Understanding

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.438898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:9ba2c2c2288b44aafea2ca9071353cb8bad6817610b064855d3f2e4110bea121

Observation 05f1833f-7345-4679-abda-3b63faa6ddcb · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Unmasked teacher: Towards training-efficient video foundation models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.875177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:97b1d6eaf99e5f32c086b9c0e4e8c35c7bc1d8cc8976238dc8cc0d3562cc45f9

Observation 39366faf-09cb-4205-8c62-99262af4df53 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.879774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:ae67c182f3e49b917f565e2c7a524338be9c3e105b346aa7564b7f373a684df7

Observation 6a3411f0-f35b-4861-8d54-10f3e21d3b57 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Llama-vid: An image is worth 2 tokens in large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.884346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:e221b992e2b74b80d557e037c4c7ed09e63ce7145df68464436d8f02f19b38bc

Observation e03994eb-65ed-4c9b-acc0-f9a966a39006 · outbound

This paper cites Video-llava: Learning united visual repre- sentation by alignment before projection.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-llava: Learning united visual repre- sentation by alignment before projection

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.889267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:40c4b5c2efa1692374fa641623bbcb0f75a92b0dfa7e106ded107c4d6efea312

Observation b19cb18b-dc47-4443-aa35-f12b1b8382c4 · outbound

This paper cites Microsoft coco: Common objects in context.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Microsoft coco: Common objects in context

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.894123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:8b1ea457582b74ab8ce608072e21949bf4a49d13a6c625f91c53d5f5f9e080f4

Observation ca620c43-9dda-49f8-8438-ded2f33821b0 · outbound

This paper cites Visual instruction tuning.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Visual instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.898560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:3a00a2939a76d250cbc8fbd789055d5ac80c94e2fb0aec9af7db722e08282800

Observation fc2efdfa-89c7-4199-81f7-5ccd0e43f5b6 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.448546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:95cd53b236b055004eebff25cc91d10705eb5baca1522d0f72ae32d741f897ae

Observation fb0f0cfd-d02a-46eb-a989-78fc83bdc244 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.456872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:541105245d1a451b932749c4a1148717d6777044bd6beea5fa75d82c40b6974a

Observation 0dd1d071-ab76-4ad0-b384-85327ae7f995 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.912859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:10e42015fa8bd67075fcf31791e8f2851ba500c81c37de79f1e6864a36894ac9

Observation b450f8da-bc02-4493-9f13-e6e97c94b1eb · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.733562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:15def52f059ef3feacef2d151425e3db1a32bfdfefa712e4860184d6bc289305

Observation 106110a7-88ec-446f-8e39-e87aa3da748a · outbound

This paper cites Spoken moments: Learning joint audio-visual representations from video descriptions.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Spoken moments: Learning joint audio-visual representations from video descriptions

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.739205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:8d4cc62cb0ffc46edddf8282eaae3be546d99328bca532ee5422d4cf53f32050

Observation a653d899-4963-46a4-9e17-1233011a2c27 · outbound

This paper cites GPT-4 Technical Report.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling GPT-4 Technical Report

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.465573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:2812dec714a765815c8199c382c6d8bbd9e18e1736c2aa4b425c141b3efdb6ad

Observation f899955c-8817-46e0-8d7e-26390ecd5e8c · outbound

This paper cites an unresolved cited work.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-18T04:02:43.752165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:c4a6f79b3d2ee23daae3b494f9eb30f0aebe3225895f816895c41471be3fae72

Observation 9374501d-dfe7-4cda-b36b-7a36fb7f9513 · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Perception test: A diagnostic benchmark for multimodal video models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.774711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:3216c1ff0d65b1a18cf7890394058dcc7349a27e46a9b84c83483dea05d06cde

Observation b843d29b-a458-4567-8ccd-fd94e0d8e4d6 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.780423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:468fe063442f5d173af60fb585027c2fffb367f1a33ee86bc42459ae5a7f888f

Observation f9b9e478-9dbf-4b84-83e1-faebb67969f9 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.474946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:cbf71497e5fd3b160fabc7b391b114f08f58ce5e9262b67b945ca470d387bde9

Observation 00d5b85c-c553-4069-8408-12eac93fca39 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T04:02:43.484634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:83c64d98c68ea9227b90289268dc78ad0961b322edee8cb49b56c7cb37740a1e

Observation e74b5432-69be-4af7-8dd0-32eccc895a00 · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.799560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:f742a4197f168ddeeffbe7f2d692f7a1dc5a2d5f8d47c9bed9e18640ff98e618

Observation 69a01eff-aa45-4564-9c4f-28685cf5b52c · outbound

This paper cites Sharegemini: Scaling up video caption data for multi- modal large language models.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Sharegemini: Scaling up video caption data for multi- modal large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.808123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:e842452c2bf867a1ad619411047b6c49f5aef4f3b33c358336b476c038935070

Observation 424c255a-4d59-447f-87fa-c36613e5708f · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.493663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:09f5e861c0a19f8ed3f5ae7d102ee98825162cbbbb5e82cf8c5f3dd16f3fd756

Observation 731ff6fb-eb7b-48c7-b925-bcf2ff8d0414 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.503594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:14728cfe0ca5347c8b1aa85a1c53e0064f3f911dd1f1ea277beab6508f89452b

Observation ca6154b8-b0b9-4640-9600-6a93a4c81256 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Moviechat: From dense token to sparse memory for long video understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.847968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:1e10a52f2e636b3f61a3fe68f90337f9b09eeb43e9fedcd1cc67b5cc4a6fd184

Observation 450a0c10-27fe-488c-addc-f2015cf948b4 · outbound

This paper cites Koala: Key frame-conditioned long video-llm.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Koala: Key frame-conditioned long video-llm

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.853509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:f2276b24ed8c47df09e655118319e856e38e6899d6c1f4e5f95d2d4a5a5c2544

Observation a6425a9b-0607-4e84-91d1-a274210e554b · outbound

This paper cites COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.513310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:04edebb488c172da6ecb4a6c3604b88b44144339af053c62f59aa7d34e11363b

Observation 5f8b2268-7e4a-457a-8f70-84c0d0615742 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.522302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:6d89af67147f0bccae69305c7a04ae7d255135543e2664087003d8993c7163f7

Observation 5760c3b7-d10b-43f6-85ce-d9520ab46306 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LVBench: An Extreme Long Video Understanding Benchmark

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.239348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:f07c2b7a56c9dbdef1b2fa583ac81e1501ec6659b89c1c1a1f7a53f7caba2802

Observation d47e1344-6c36-4fc0-b57d-eef20f1ac2e4 · outbound

This paper cites Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.537223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:407ab90a9af5f22ceece69ceb7f52a7889311537da28b35ab4a31bf6e970f586

Observation 9489c8d4-380e-4e33-aee0-489b66a22283 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.544903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:f838bc92a09d777c19b35878c0fbe6cefdf01d73e95e150ff5d6459b92e24659

Observation 600ba87b-ff0d-4e3e-a1e7-027435d4e235 · outbound

This paper cites Internvideo2: Scaling video foundation models for multimodal video understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Internvideo2: Scaling video foundation models for multimodal video understanding

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.744790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:38b3ee4b65a6c774c9d02f8caa10216eee51b107fc475f8aae5ed5a3d451658a

Observation ba361a7a-a74e-498e-8ba7-1aab43771bf1 · outbound

This paper cites Visual Context Window Extension: A New Perspective for Long Video Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Visual Context Window Extension: A New Perspective for Long Video Understanding

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.552982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:c512b82ed92e370a32a30aebab20426170117f6cc8963544f6e86ba562aca287

Observation ac284730-eee2-40f8-ab43-a6799a557d74 · outbound

This paper cites Longvlm: Efficient long video understand- ing via large language models.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Longvlm: Efficient long video understand- ing via large language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.792603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:87e9a0c16f0e612b5825a27db7a191b9d8b088b51b92fb435bcb199271201165

Observation 4de7f155-5d02-4109-9a9a-0c302c4cdc64 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.332710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:0f21d11e7bd4512882eaa995294a98c397b6d5e872902317e78a01bbe53661d1

Observation 6feda575-e467-4932-b44d-2da9d80ef809 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.561759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:0eaa1fb52d255f8ae8e287675e6fcf1d25c8767189aa64acddada287998dc453

Observation a100e30c-9c51-4c45-9148-50e2b78a265e · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.569786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:fc04131344792fd18a8bd4e1d339d200efa8f8c6a7674d3e173ad269f62a84e8

Observation b9332801-9f36-4c56-b77f-a8364c0f149c · outbound

This paper cites Advanc- ing high-resolution video-language representation with large- scale video transcriptions.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Advanc- ing high-resolution video-language representation with large- scale video transcriptions

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.864787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:2a35146d9b22dc98d4e254a65f856eb7a5ec05921d5f3a799dc7dfe3cfc5e4f8

Observation 5cfa7a2c-3401-4e26-b69e-c5672311255c · outbound

This paper cites Vript: A Video Is Worth Thousands of Words.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Vript: A Video Is Worth Thousands of Words

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.578826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:a912b06b329ac104fd29458c20c8b3047a9ee496de86bd9d2bcb2e79d1132abc

Observation 97ab7361-9b0b-40e7-9f20-7f65a83d98de · outbound

This paper cites TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.587545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:ce9edf31b775826d194ce6185860e59fde2bfc6cf426bf7804fe468d4bc3d8ae

Observation 6d06eef6-2d31-4fd3-a6bc-2736bdb5bb57 · outbound

This paper cites Sigmoid loss for language image pre-training.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Sigmoid loss for language image pre-training

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.907153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:c84a51a5f8a211560e09ebff3918f689f3a867c6c15dcf148e879c591e077982

Observation 47e71157-dafe-4069-b911-66e8c6f76dcd · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.594914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:5154878240ef2e1ac518976f104b78f41ccbc5fe077ee8eafcad525eb691d2b4

Observation 428835ff-04fc-425d-834c-bd2a36c4e492 · outbound

This paper cites LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.607191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:b934641d14a22f13340be9ca0c06b5b304af4e4f287eed886fb9a3e490ebaf9b

Observation e272bae0-47ed-44e8-9638-54b01db80967 · outbound

This paper cites Long Context Transfer from Language to Vision.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Long Context Transfer from Language to Vision

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.616047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:e398d4c64ead67d323d0ee506a1654b23795608a975ed30ba2d211c3c6080f50

Observation 869fdb9a-b101-4c3a-b5bd-5fdbd436cb60 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.624687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:7612f368d1258c1fe2eb1e049b84423906f9a0967624dca928ed9ca70a06e101

Observation 53a85dce-09a4-4b2a-b5c5-1a81b15478ab · outbound

This paper cites Llava- next: A strong zero-shot video understanding model.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Llava- next: A strong zero-shot video understanding model

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.902683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:3c374a12e0066e16649be86ab8fd08a4869e01d385e0c7fefdc0681d3657bcc4

Observation 84022ade-c828-472f-a4f0-b9028aa05dfa · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.632932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:a920201561fa8e9a7a8cb1e2332a467afab0175ee899924b7569a950f9727fd7

Observation 93d93d03-bac4-4091-959d-50669fe75564 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.641934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:01f5ea13cd7cd15eb468aade1b7522bac0e8bc22c988ac5147a2cb6ea3d9ac73

Observation 9d9007bb-eb1d-400d-ad90-cecca5f3ca5a · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling MLVU: Benchmarking Multi-task Long Video Understanding

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.650265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:2e8f216b0557a0639fcf15880e3442fb929b08d7454aaea008db2e02ddcc9e00

Observation ed625540-0776-4336-b00a-cf1276dcd26e · outbound

This paper cites Visual Dropout in LLM Visual token redundancy in LLM inference.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Visual Dropout in LLM Visual token redundancy in LLM inference

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.870297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:6dfa46e2c699ee3d87771dda57ba1b3d8f38ba3872996787897bbc72c780852f

Observation 0cdb5f83-0a33-42e7-abbc-9a728e983447 · outbound

This paper cites Video-Language Connectors As shown in Fig.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-Language Connectors As shown in Fig

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.785923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:12af955f7b54ba31a6934a66fa00c6c3cc3d51d3e83535651405f0b307011b29

Observation fb7abffb-5425-4f12-8bfd-eaa32f4854c0 · outbound

This paper cites We provide details of the data construc- tion pipeline for each dataset as follows.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling We provide details of the data construc- tion pipeline for each dataset as follows

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.835348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:886a67c5f94cc9b98c89f4f652d1302f5728c5d5e33b10762d658ace6e97b461

Observation d0c7ab25-93cb-4bd3-9c63-697861feefa1 · outbound

This paper cites 11 and 12) and long video understanding ( Figs.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling 11 and 12) and long video understanding ( Figs

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T04:02:43.858779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:4562b84077d1c15894ccdfa498177acdd113e72f4bedaa2a741bbeba22adf0d1

Pith citing papers

Observation ed0f7410-4675-4230-bdc1-907ebd304255 · inbound

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay cites this paper.

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T21:31:44.439495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:31:44.439495Z digest=sha256:3901dedbd5f32fc8728523b2ed19f0d5811dce0878f45340271e596d36d5148d

Observation 59801fb1-50fb-4317-8a32-1acef7e4adef · inbound

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding cites this paper.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:cdbea6dfe91015d500df2c77d3713098726b06c547cc565ec89b9e62a22c0714

Observation 304f2a4f-3440-41b9-a78a-5c5e63b912ff · inbound

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning cites this paper.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:c2e7e8e3d3e78d2e1a286274a62c4c51330c591c01493bc221de2b97f4b7d421

Observation 59b3a12f-1ae2-4023-b6b1-7ce39e328063 · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.594042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.594042Z digest=sha256:933bd0332e13b23c37061fd4cb83a1ec1017f02c6d5c33dc07dfaf3359a88c1e

Observation 565c56fc-a2a4-4cb9-9369-3c2a3140fe2e · inbound

MR. Video: "MapReduce" is the Principle for Long Video Understanding cites this paper.

MR. Video: "MapReduce" is the Principle for Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.673130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.673130Z digest=sha256:2ebc6e75a6dde94b0bd11c32fb62a2a167a85eeab0891d5a3b1476df1ba20d92

Observation 3dd99b1a-365c-4828-8d50-2430973835a8 · inbound

MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention cites this paper.

MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:26.691907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:15:26.691907Z digest=sha256:5625fdb54d351b4589ca5125d918f150d317707cc380e7ef3e94ecbe1a7ff038

Observation ef6225a9-3859-4964-aab2-ccc56aaf5b35 · inbound

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos cites this paper.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.231104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.231104Z digest=sha256:8a8e12166fa8989a8b892f8cb7702c551415a8b800bd3fc77b51328268bd3cac

Observation 4de4747d-d6ab-44b7-ace2-b89525acc29c · inbound

Multi-Agent System for Comprehensive Soccer Understanding cites this paper.

Multi-Agent System for Comprehensive Soccer Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T23:49:10.088142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:49:10.088142Z digest=sha256:07a46258403dbe1e433edf9f3a8b7e8614aa883275e8dc34faa70bb26b64efc1

Observation c725de8b-4d3c-42b8-bbb1-e77eda694019 · inbound

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models cites this paper.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.622058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.622058Z digest=sha256:3c4b68ce86a243a5f3e64f6e0dc24d214c63725761d127d15bbee833ddd7350e

Observation 0e55e230-5270-4ea2-ba9d-10a36c347294 · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.968075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.968075Z digest=sha256:56e52792d8d1b576393d27870a4d1641a8f9738297178ccd955e07142e32d44c

Observation 7cbb5882-cc3d-430f-b987-5b744e65f3ef · inbound

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought cites this paper.

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:08.944838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:08.944838Z digest=sha256:be7d684deb0767265ab91949040cf12956defb03fba502f495a840540e9db248

Observation 1746ded8-6795-49d0-871f-472260bf428f · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.621901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.621901Z digest=sha256:22085fcbed74ac27678a61a11c3060e4d77e4b7302ca19b4ec7840ec9b1eba22

Observation bdef7f07-c1c2-49ed-8e95-b25717da792e · inbound

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs cites this paper.

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:07:15.433789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T11:03:59.222849Z digest=sha256:d9558474384024bcfcac1b6b0c9062c728564037ed7343f9ac1c1238e0a760ae

Observation 464c5e5e-07f5-46f0-ae02-fcf6e2b9ff09 · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.894485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.894485Z digest=sha256:9c284377daebc52d3b31f91879db389556a3c40ab58e0a242ad2667fd9c4b87e

Observation 2180f84d-ae7c-4eb7-9557-4d947e0cb1ad · inbound

Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the Unknown cites this paper.

Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the Unknown VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:13.842219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:13.842219Z digest=sha256:94744233242ffc3fe93636d42ef8911443b9362880dafc413939da5d8a30e685

Observation bdf7c58a-339b-43ed-9661-ce81ff93e999 · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.438610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.438610Z digest=sha256:93e03a7c8502c55f7769ee0d4c864f437890428886c6f69cb411c6d840501dd8

Observation 632099fc-9a4c-47e4-b513-325dddaaf3e7 · inbound

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization cites this paper.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.138597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.138597Z digest=sha256:602c2885eaed6141c424eb86052b0caadb150a1ba5fecd081bade3449cf07ed7

Observation 97ac225c-a44b-48e7-b367-d1e266ecfdf5 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.714751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.714751Z digest=sha256:ba11b8a26ecec3415a8e4bdf28362b028d6f6b350d3d1eb8d58c0cb2d9974099

Observation 215cecf1-67bb-41ca-83fc-6a30a3038aa5 · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:39.008817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:39.008817Z digest=sha256:469fcc4f7e4e47011a6d2d6ccc9938243619bac8525701c30c96fd92829a01cb

Observation ceb20892-77c2-4ab6-a4b0-7359bf440543 · inbound

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding cites this paper.

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T22:01:25.950345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:01:25.950345Z digest=sha256:b63c189cc77a429df05f9cafcb1c9b1bc9e47e1a4e59225d0dd686b1311d2949

Observation 7022ef98-3fa6-4b6c-b519-e0a1146e13de · inbound

Cambrian-S: Towards Spatial Supersensing in Video cites this paper.

Cambrian-S: Towards Spatial Supersensing in Video VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T03:46:04.363500Z digest=sha256:d3a42eb0d4a80e298b5bb5ef5dff9b92ab6d7f63c140b3c8ce3b26a80f1c32a2

Observation 001a86da-c99c-44df-80e1-33be32c08153 · inbound

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding cites this paper.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:26.641945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:26.641945Z digest=sha256:7531a0be9cd48d181c62ef53f0cb3012df36b5114dfa7b72f728c60750e6cab6

Observation b8d31a43-6c44-42a5-9ac9-39649bcdf1c2 · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:39.557562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:39.557562Z digest=sha256:0f56cd0860c412f109fa71abd1dc7defaa9cd943bb53bea54cf8aa552dbf1da1

Observation 8b10fd8d-e836-401e-8b70-c49303de9192 · inbound

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding cites this paper.

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T04:21:29.526008Z digest=sha256:dae0accc66c74a0acfe14f0661eddf87c270b6e35bcf7aad6bda306cefcf812c

Observation 5f773962-080e-4f81-b5a2-6269198da281 · inbound

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning cites this paper.

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T12:17:42.135851Z digest=sha256:3103956bf371a478fecbdb6e9f2ff603e5de7a89d52e6f599efe73d9cd6613c9

Observation caf712ff-39f2-405b-b02d-3c7e44b8e3ab · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:58278802aede9883701bc8ccc83283949bf804697dc77b44783046776ca2c9b6

Observation 755d7aae-8794-410d-9ab5-18f342f58639 · inbound

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting cites this paper.

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T15:29:31.834566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:29:31.834566Z digest=sha256:ce5263b706531325b273bd0c3c727aeeeb637a899b95aeb7d19025fa9b590977

Observation 9bf6b2fc-d160-4eda-b8b9-fa7e050e45f8 · inbound

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning cites this paper.

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T19:46:16.975267Z digest=sha256:6d9b51de919e1c3b22239dafc9824f543f8ff804e050c69d4ba1019b114c2a10

Observation 6f1f7f69-f505-4b72-878e-8402334e5135 · inbound

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding cites this paper.

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:21:47.439019Z digest=sha256:3981cd170aa10a6c7acd777ac6b54ae7684a3b1fbe43d7602c30bfe92b418440

Observation e2f84104-be0e-42d7-a0af-20076b606867 · inbound

Small Vision-Language Models are Smart Compressors for Long Video Understanding cites this paper.

Small Vision-Language Models are Smart Compressors for Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:fa1ebc4df7f64a0d12a0ffc3fa69091580bbd586f9e1d4631e5d26ebaebb1d8a

Observation 67be4ce6-9f37-4586-b1ca-299ab7ba4e77 · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:9530cb562c6fb4a245e0049cf86b34d46294a4a6f4675abb10418bc5cf9b6e5e

Observation eb49f453-1f16-4af0-8773-198ffa504596 · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:fe80521f01f9f2c4735496de96005ea2b3f4da2cd0b826633b44111fa327aac5

Observation 2f95654f-aee8-4488-b844-0c8567bbfff0 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T04:31:26.325118Z digest=sha256:ae69c5f089ae33e0c2e1a9b56dbe78e6e87b4fa787341e33c9d9bd861f61405e

Observation d3a281b8-a4c6-4ec7-9e22-7b5e6f809ea1 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.161913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:fb128b498a5c1301daaf871a62839674238737ebe1977ea057d1fc2af936628e

Observation 1318ce42-fc44-483c-a7c7-eb48232047fa · inbound

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding cites this paper.

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T16:51:43.783133Z digest=sha256:ae2575c260ca794ddf251510d8e5818d864ea0cb89902ffa70111bfd2a2ad6f1

Observation ce784b1c-a6ce-413a-ac62-537513023795 · inbound

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading cites this paper.

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T16:10:22.588945Z digest=sha256:bf150d6987b586a0ec660aaffccadac9a8d158a5daed6031e31804102a7bd4e1

Observation 52e12638-e132-44a5-8b65-f8bfc00e0044 · inbound

MedHorizon: Towards Long-context Medical Video Understanding in the Wild cites this paper.

MedHorizon: Towards Long-context Medical Video Understanding in the Wild VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 97

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T12:28:35.008604Z digest=sha256:066640c46ecec741888bcdbb6e6d29471da4ce3b0d482eac58397de1f260ebe7

Observation 4adeab72-40fb-4e0b-b7e2-130439272b81 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T04:13:21.487431Z digest=sha256:10e6a332693d5bb74ba22bbf70f14d8a3afd6514f8e9ab43e6c4f2f302543ef3

Observation 6ec7e470-e915-41ae-adcc-28dc93d42da5 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T06:53:42.726350Z digest=sha256:9690eca77efcc76a6a2d061f751e4a0e201c234d6f42cc02898feab6046584af

Observation 66fdcaa3-4c5c-4457-9147-816527f507fb · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:1d7d8e16d144726ea3fd83f21a08f6402e460ac58e0d08a22008e3f832f37d61

Observation fe731a24-2731-4709-a6b4-008686feedc5 · inbound

StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video cites this paper.

StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:33:50.726989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T23:32:38.332828Z digest=sha256:3d5a3e21881f3523d991dc58acbe07f916f0219868090645dd01cfb08f1c6e2c

Observation ec812b9c-1654-49bd-9cb1-e7a4964b0231 · inbound

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning cites this paper.

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:13:24.794852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T15:12:00.408851Z digest=sha256:ee26266dd741d5ec7ea89cf66947327d9a34062cf2f5f3aac14e10431fcb04e7

Observation f2146df8-67d7-4947-ac5f-ccda5531421b · inbound

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues cites this paper.

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:14:42.412095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T07:13:43.716510Z digest=sha256:ff8c0e2ae02a4991103decc74696bb7af254556966ee17e1af454b8badaa7512

Observation fb4fd70d-585c-48b0-9d6e-22d13ada3ef0 · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.641616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:1ef4dc86ddd443606655ea92e8cebc3bd77c9b2b7cd5febdcaf5422f437ede44

Observation fd7d14f7-d3f6-4de2-91e8-9d9a6a2997d2 · inbound

Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification cites this paper.

Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:33:28.319648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T13:26:21.016303Z digest=sha256:8367cb8497411e995bd956172581a3c5866b420e58c045da2f459a4fe040f35c

Observation 43226218-c24d-4428-810e-2184c620bc5c · inbound

Linear Scaling Video VLMs for Long Video Understanding cites this paper.

Linear Scaling Video VLMs for Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T23:02:46.240303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:00:11.246232Z digest=sha256:15afca91125cf5e4b6d098ff25b09bcae19a4e0a63a69f1944c1249c58775203

Observation 7c215330-0ec7-4050-a3a4-90e738110d57 · inbound

AdaCodec: A Predictive Visual Code for Video MLLMs cites this paper.

AdaCodec: A Predictive Visual Code for Video MLLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 74

Resolution
malformed identifier
local_arxiv, observed 2026-07-01T22:26:18.177531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T15:20:48.248576Z digest=sha256:06b7aa4753e6da4d20182ec14c59d22744a78b1dbca87774763ee958c6770b28

Observation ce149296-1715-4fee-8bf7-78373f24a336 · inbound

Towards One-to-Many Temporal Grounding cites this paper.

Towards One-to-Many Temporal Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:16:57.727597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T02:11:48.455492Z digest=sha256:a420ded863b39ba2e4ac59fd9acd6672172d192de701db9330a20de9e3174021

Observation 9be2b6a0-608e-4891-b064-1e0a1a0a2cf6 · inbound

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset cites this paper.

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 110

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:26:57.204162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T02:05:47.810096Z digest=sha256:6dc3e528a011336dc11321a837699444cb576820ca16c9fe9b44d183d73397a8

Observation a32cd9a1-2731-47e8-8f02-3a0d2d0a4957 · inbound

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning cites this paper.

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:56:47.341319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T06:33:32.090913Z digest=sha256:05d088f94a005fb798aacc4d76170a8ab6e011019a16bac589d3de36050af7af

Observation ecee5536-b925-44dd-9aa9-f5e695be6697 · inbound

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models cites this paper.

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:17:09.563522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T22:49:18.420491Z digest=sha256:bed7eb6bfb61b14f7957ac9a7ad2afe92d4b3641bcafea1abdac8057c911bcc4

Observation fa29733a-7da2-4aac-bdbf-d5e70755b7cb · inbound

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding cites this paper.

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:29.050040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T17:35:07.667016Z digest=sha256:f840eef3097db4d837602d452d23b3d14e05ebb57dd84a9737efb882480057a9

Observation 4b1a77c9-1b4d-496e-9916-5f4f058923b4 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 168

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:38:28.858320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:01e3d0b526a8441824e54b19c2d09c9fb4a82c3f583452cd04a41a176741fafe

Observation a3eb1385-31d6-423c-b654-d54bef28fc13 · inbound

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference cites this paper.

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:19:29.917353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T18:17:53.013043Z digest=sha256:4d16bdd6bed32fe367dd3079895b02c0fa2e920948089f2cef0ac0d47e6421f1

Observation 2df7dcd3-e675-474f-858a-cee988064085 · inbound

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression cites this paper.

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:49:57.266359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T00:16:56.174638Z digest=sha256:47c5fac5a1922721803234492f2e5c8443c9970e961bbfbbfc043763f738dac1

Observation 8ecf9277-e14b-4f6e-a53e-1cb3ff4a078f · inbound

Learning to Deny: Action Denial in Multimodal Large Language Models cites this paper.

Learning to Deny: Action Denial in Multimodal Large Language Models VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T09:45:39.610427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T06:21:09.996386Z digest=sha256:61732847f3d7c48720eb44f9cd80042ad9575ba0fe1783a3c561675554606269

Observation cde95483-51b5-46fa-9610-b56e3f11fa18 · inbound

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning cites this paper.

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T05:48:27.255331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:48:27.255331Z digest=sha256:2cb95f6e9e74e0450a759dc38f8e13e7603421f3a2700847eba6a6be4948e626

Observation d9641bb0-bc0e-4aa9-8d06-47f68c5dbe93 · inbound

DynTrace: Tracking Dynamic Object Evidence for 4D Spatio-Temporal Reasoning in MLLMs cites this paper.

DynTrace: Tracking Dynamic Object Evidence for 4D Spatio-Temporal Reasoning in MLLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T06:34:27.665052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:34:27.665052Z digest=sha256:ff041fb6f5cf668bd6e19c4a6785d6b46f8eef696c405427081832ac52ac4607

Observation e7e61b8a-07a7-4a88-98f4-a78be078cb89 · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:40.495393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:40.495393Z digest=sha256:e163cf6fa6910690476f57bc9265304795671325c16484fd493926bf80341f13

Observation b7669212-e82e-4bbf-b512-faae1e162957 · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:14.015999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:14.015999Z digest=sha256:1d90ee4d6a1e8a1a962f1f29a58752e73e9a5f4835cf369039ec8ebfb6f4f52d

Observation b3bd2182-e873-4443-8b8e-d01840d6f635 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.750593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.750593Z digest=sha256:def5a0a92b892c6c11518f075e0cc53234df39e7b263186c4b96d1bdbeab4f41

Observation d3b7f910-a70e-4475-b4e6-7f994cd1d712 · inbound

Long-Horizon Embodied Decision-Making via Multimodal Memory Compression cites this paper.

Long-Horizon Embodied Decision-Making via Multimodal Memory Compression VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T00:13:22.601162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:13:22.601162Z digest=sha256:555affdf0ad0f619804b2a882c269781016c7c8bee36addae50637a5eba0bc96

Observation 98abc67f-d45e-4fbe-9411-15972855c281 · inbound

CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding cites this paper.

CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T15:43:54.334942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:43:54.334942Z digest=sha256:f6953407de7a153e6c5d4b8e56f606faad28f78e695b19ced642348fece028af

Observation ff6a968b-88c5-4d75-9640-05dcc13fcd9d · inbound

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding cites this paper.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.744823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.744823Z digest=sha256:f8d4c3e7e9bc14bfe18556290f397f5fc662d154ec94fdbc7636e3aaa9455d34

Observation 8651f988-a6eb-478d-9397-9933409b53f1 · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.314074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.314074Z digest=sha256:65edd1106307573c6eadeaeb610fcebe835178bdac39cb5a684abae06792fbec

Observation b8e0bcb6-93c9-40c2-8b3f-4bfb5f079b1c · inbound

Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No cites this paper.

Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:13:21.297865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:13:21.297865Z digest=sha256:a79488fd833214582e62c9427551d7a7b745072c7aa6a1af947da73d563fd25d

Observation 012b1cc9-df6e-435f-aa48-8bbc157a53fd · inbound

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs cites this paper.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.437312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.437312Z digest=sha256:eb79f3e3f4f8bc566a098452c22bec43e7f560676676d573c64cbe1f92a4185b