Pith. sign in

Paper Citation Record · LEDGER

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

As of 11 August 2026, this Paper Citation Record lists 100 of 122 outbound references and 0 inbound Pith citation observations for arXiv:2607.14935.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14935 v1

Coverage vector

measured 100 of 122 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:44:49.747631Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 122 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a5ccee9-7d80-4eb4-980c-aedb59733c49 · outbound

This paper cites Qwen3-VL Technical Report.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:40.263863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:40.263863Z digest=sha256:cae1b68f542808a67721f4b4839af42f6b57a6884f24aecba1b2b7be5acd2f77

Observation dbcfc400-3eab-4288-add5-b74a12a64c15 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:40.385436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:40.385436Z digest=sha256:d60bd7dfd11e04c24ff8b9dc41c34cab1b0f461365a0b0d972e4817ab1c5933e

Observation e7e61b8a-07a7-4a88-98f4-a78be078cb89 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:40.495393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:40.495393Z digest=sha256:cead89e4ab0f17348c295f7c9787fb8918c4d7ba37f1db640a57663004f8246d

Observation e667046c-a725-4642-b587-c2328b08047c · outbound

This paper cites PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:40.660990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:40.660990Z digest=sha256:1b8af0e2c85fd87cb44f49f5b2cf9ebf887a86e8bac00ba4d2e6024924dae198

Observation 89442b5b-01c1-4281-b165-e4ae54133cd4 · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:40.771026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:40.771026Z digest=sha256:aae8ee81567611bd98b099e5c708f45cd06337a3b16477b15db019b4b4283515

Observation ef734fdc-7083-4879-a6ff-04ef4bbef4a0 · outbound

This paper cites Qwen2 Technical Report.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Qwen2 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:40.831039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:40.831039Z digest=sha256:bfca83d5a2e3ecbdc49de679fb7cb815806f8e12ec089118993704c011e18880

Observation b3002662-927d-4f41-84b7-8cd3dfe506a1 · outbound

This paper cites Qwen3 Technical Report.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Qwen3 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:40.939379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:40.939379Z digest=sha256:d4e49bebea6b62a3cbc79986db7eb5848afb7cbc38283ab8582629ae1138e375

Observation 48755b6a-f903-4acc-923a-ccfe3364939c · outbound

This paper cites Videochat-r1.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Videochat-r1

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.103975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.103975Z digest=sha256:205c90bcc76de7fa43103065c366cd4309cbc7f55c0d7157c299128249fad8bd

Observation 7e95a7ad-809f-46f5-8825-6b314a59a3ec · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.162135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.162135Z digest=sha256:05a4bb511fd37321a41056c1a47edafbb6d6d1a2c4718e7c91ee00538eef1e2c

Observation 7eb20e68-e4ab-472d-83e9-bc2fcebe7712 · outbound

This paper cites Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.221547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.221547Z digest=sha256:31d6818869ad75b6668c065eed2403cf02e0bc43b653863c00db6de97520bae1

Observation 6a14c559-b0a9-4999-89b1-aa7072fa0657 · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.291479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.291479Z digest=sha256:52a3fc6dd901d123e53ffc2b548fe31d2cf8c74625321be8eb48ba27ae1e82a6

Observation 3971ffea-3617-48e5-a5ad-d0b190556c65 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.382976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.382976Z digest=sha256:8108a641982f1e37a89cbbfc58e7d15bc2ec861ac2be3a68df53875764e4152b

Observation 896580a9-b2dd-4e5f-928b-d905b606de14 · outbound

This paper cites Timesuite: Improving mllms for long video understanding via grounded tuning.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Timesuite: Improving mllms for long video understanding via grounded tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.430131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.430131Z digest=sha256:76a56c05c3799f917c3028f8889af9c092958d9fa7c035de35377a6afca93e27

Observation 9ee61b8d-42f9-49b3-bdf3-4d68bfebf8cd · outbound

This paper cites Timelens: Rethinking video temporal grounding with multimodal llms.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Timelens: Rethinking video temporal grounding with multimodal llms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.491845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.491845Z digest=sha256:4e816bc975ad8d019cf82a6c38f11286bedef176e4df14ffe5654e0a256066bb

Observation a97708be-9b01-4dac-b427-131bd70fac0e · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Videollm-online: Online video large language model for streaming video

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.520198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.520198Z digest=sha256:d69d51ae38600884d1ef9026c6ff6c40ff273ca565a93f884ad0b6dbe5080ab2

Observation 7e5440df-9915-40ca-8717-78f8973438bf · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.589936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.589936Z digest=sha256:85bd704be1921e6b05091cd3882a3da35cce13b146c4bd4ca095eb0abe11081d

Observation 1a369fec-b7b9-4308-81b4-7288778b1d35 · outbound

This paper cites Online video understanding: Ovbench and videochat-online.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Online video understanding: Ovbench and videochat-online

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.682216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.682216Z digest=sha256:7999f3ab78856ad684622831771c74e7eceb488fdb999d54af9ba3a91ed1db26

Observation cd5440ab-3792-4d09-8638-349f095ca67a · outbound

This paper cites Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reac- tion.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reac- tion

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.752189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.752189Z digest=sha256:60e1fbc0dc65ff29cbafdb8304f34ccbb1fda7b9aa36fe993cb1d9ca2f9bc8d1

Observation 98d24d20-5989-423a-807d-fd5b57368f76 · outbound

This paper cites Streamforest: Efficient online video understanding with persistent event memory.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Streamforest: Efficient online video understanding with persistent event memory

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.822857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.822857Z digest=sha256:b912dfaa014f31fda47bf066f44f256e02f35649cc0dd697eeadf3132a941f72

Observation 8e546c26-dd84-4100-afa7-2c56326fc8b1 · outbound

This paper cites Streaming Video Instruction Tuning.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Streaming Video Instruction Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.892453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.892453Z digest=sha256:434d4c6c93934da56dcc3a5b8540decae47d45215c5c8789066d2529babc53ae

Observation 993fd87c-d1d1-4a9c-9cb8-d355b9fc2d39 · outbound

This paper cites Streambridge: Turning your offline video large language model into a proactive streaming assistant.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Streambridge: Turning your offline video large language model into a proactive streaming assistant

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:42.000829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:42.000829Z digest=sha256:375053051198eccc0485b21d32faab919606900e43a5a103bdc16a97c6fad7b5

Observation 84c47551-8d13-4eb5-afe6-0ea131a643b8 · outbound

This paper cites Gemini 3: News and announcements.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Gemini 3: News and announcements

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:42.058308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:42.058308Z digest=sha256:4d9216145c425f806080618944eec13cf11c1a2c41676a6293dc5875858c2705

Observation cdfce8b2-b132-4d59-be38-3009f0e0f2bf · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Kimi K2.5: Visual Agentic Intelligence

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:42.146048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:42.146048Z digest=sha256:d6ca2c457770178af7c87cbafc7e93ec603803dd7d7b14c8b39eea1b01e6bcd8

Observation f292ffb8-0f86-4e74-879e-f18ccb366a37 · outbound

This paper cites Seed1.8 Model Card: Towards Generalized Real-World Agency.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Seed1.8 Model Card: Towards Generalized Real-World Agency

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:42.258052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:42.258052Z digest=sha256:b13dbab53c9aa91efad65a4e1bededfe294e1f1f8ca1ccfc802af0e1c3cc2492

Observation e387c711-e657-4b7c-89db-b3d318a62123 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:42.442107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:42.442107Z digest=sha256:29a14496a7199f77a00f30f6327de3814ebaa38338fae734d1ae26344517c3c4

Observation 68902810-0686-4fd3-a7f9-7b0fbb7bfd67 · outbound

This paper cites Llava-video: Video instruction tuning with synthetic data.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Llava-video: Video instruction tuning with synthetic data

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:42.626276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:42.626276Z digest=sha256:17005ca2da0722d62cc3481b8525af1c0ae7331426a187c89a1c72faea2a61f6

Observation f310e039-88e8-4489-a3f5-4b83df3475d5 · outbound

This paper cites Spoken moments: Learning joint audio-visual representations from video descriptions.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Spoken moments: Learning joint audio-visual representations from video descriptions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:42.674221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:42.674221Z digest=sha256:c9e60f9c6b6799d7d432d1fab1c4c3197000169fd46cbc8bacc8459681489473

Observation 696b1e40-7be9-4046-870a-3400740071da · outbound

This paper cites Vript: A Video Is Worth Thousands of Words.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Vript: A Video Is Worth Thousands of Words

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:42.744544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:42.744544Z digest=sha256:965d5c3de061c28d797cf68ca0a1e25109a2a0b81aaa9409a9389d60bae352ab

Observation 348f95ff-4641-4035-9846-921b4270211e · outbound

This paper cites Kimi-VL Technical Report.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Kimi-VL Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:42.862348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:42.862348Z digest=sha256:a0eb73cac866c9fae1b808166e60916969a2b40b9951aa661313a8f163625dc4

Observation 7ad1ffdd-369f-409c-b585-466a4e9b6ca0 · outbound

This paper cites Visual instruction tuning.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Visual instruction tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:42.961347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:42.961347Z digest=sha256:9465814ffc073556e3aa453a2939d1687d41817afccfb08f0aab15495bfac631

Observation 1fbbbd9f-5d8a-48f9-8681-349e8f704aba · outbound

This paper cites Caprl: Stimulating dense image caption capabilities via reinforcement learning.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Caprl: Stimulating dense image caption capabilities via reinforcement learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:43.030845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:43.030845Z digest=sha256:063aa813ccef632553fc58d83008c56dfdfdbd35c13e43f43a0e8d3c3da95ed9

Observation fc39791f-29a1-412a-8774-f3fba5cbe5e9 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:43.081065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:43.081065Z digest=sha256:9d3cf94ac52dc2e43cb44cb17833844e625072dd4e5865d89ef8159f80334adf

Observation adf665ae-a5d2-46df-a260-54c49dd84a5f · outbound

This paper cites Densefusion-1m: Merging vision experts for comprehensive multimodal perception, 2024.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Densefusion-1m: Merging vision experts for comprehensive multimodal perception, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:43.181761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:43.181761Z digest=sha256:9538dd99a164a8dbdcc0156219aa42d38c894dca80b34c104acf034485589bd0

Observation 2e086231-e4f4-4f51-b110-bb8fb79de085 · outbound

This paper cites Conceptual captions: A cleaned, hyper- nymed, image alt-text dataset for automatic image captioning.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Conceptual captions: A cleaned, hyper- nymed, image alt-text dataset for automatic image captioning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:43.251365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:43.251365Z digest=sha256:ca4fe6f85d3af982310fe9f6c070272de2d3f7f467151f02a7d3940c510c9d42

Observation 04b79d1e-0edf-457b-bda4-415a13500e9f · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:43.345486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:43.345486Z digest=sha256:72ecde89fc00fe2f782ac843baa8db05d60317fc141404209dccc5ccfd08a37e

Observation 5f1b7b9b-5867-4fee-917d-998c90074fef · outbound

This paper cites The Kinetics Human Action Video Dataset.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding The Kinetics Human Action Video Dataset

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:43.450847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:43.450847Z digest=sha256:c6e3f9eb4e7ba3064c8a3050aa93d302a2122d43ce83fa9cd7012f03c02c8864

Observation 9070223b-b3a6-407a-930f-bb3cadac97ad · outbound

This paper cites Perception Encoder: The best visual embeddings are not at the output of the network.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Perception Encoder: The best visual embeddings are not at the output of the network

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:43.511377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:43.511377Z digest=sha256:c6a09f9b7e1bab59779dc627b6b3ed98324e79dedcacdc123870c57a698d8af3

Observation dc8bdbe5-db5f-4970-bc1b-09f5f5777cae · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:43.581545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:43.581545Z digest=sha256:e73bbf7e007c73ba73fe19d10f639361594d6d9e1e4906f12e470c6d1dabdcda

Observation 97b1d4bd-f720-4045-b127-9119a204ce99 · outbound

This paper cites Movie Description.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Movie Description

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:43.651420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:43.651420Z digest=sha256:be97bf067042c78856a8fd9481b922c2ebb0735dd0eaa9a96dc4b8678e3d42b4

Observation 2487f5b2-24d7-41e8-9d0e-57d5915cece2 · outbound

This paper cites Vatex: A large-scale, high-quality multilingual dataset for video-and-language research, 2020.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Vatex: A large-scale, high-quality multilingual dataset for video-and-language research, 2020

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:43.802803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:43.802803Z digest=sha256:c6d465c4b39d2e1e7ade99dbc396238a6c247cb7330acaa2494cca2012be8ad3

Observation 5d1616b3-84c4-4329-b3fc-b43ea5e67a88 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Activitynet: A large-scale video benchmark for human activity understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:43.836826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:43.836826Z digest=sha256:0770223ac2e6aaed2f2aac8df77727503db57df846ec39484f487b9409e78140

Observation b3a2e685-6b4c-491d-acd5-1ad19717cea8 · outbound

This paper cites Hollywood in homes: Crowdsourcing data collection for activity understanding.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Hollywood in homes: Crowdsourcing data collection for activity understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:43.888588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:43.888588Z digest=sha256:39ee94ae2cc7c01b7b440d65491931b641164ab1ca2f4d262173802263c580e8

Observation 2722b1ba-4233-4f03-834d-c6f0f5029497 · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding The” something something” video database for learning and evaluating visual common sense

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:43.944633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:43.944633Z digest=sha256:e4a290fe65f8865618d1b6bd6ac3425d64b7e8f882f2375751dee39675233a74

Observation eb7ea558-f755-4e43-98a0-275f5ac5a278 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Ego4d: Around the world in 3,000 hours of egocentric video

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:44.014790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:44.014790Z digest=sha256:9f34031bf95401443af66eded18153924935e125517485072731717f8cec0316

Observation 5334e404-e128-4616-891d-a8d35781854c · outbound

This paper cites Sharegemini: Scaling up video caption data for multimodal large language models, June 2024.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Sharegemini: Scaling up video caption data for multimodal large language models, June 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:44.154403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:44.154403Z digest=sha256:28a1a57d0d998305caa246f67735ce209bddec2bb0c37a80f03dd6b8ea7a21ac

Observation 434eb6fd-db63-4997-8139-4e4e8e4601b0 · outbound

This paper cites Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:44.322802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:44.322802Z digest=sha256:8c429d2fc47b936965f6611ffcfe45bccd5bc8b516612425e72de0d4470a6b61

Observation 1c212413-8de1-446d-97f7-3302e944f610 · outbound

This paper cites Bee: A high-quality corpus and full-stack suite to unlock advanced fully open mllms, 2026.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Bee: A high-quality corpus and full-stack suite to unlock advanced fully open mllms, 2026

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:44.416164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:44.416164Z digest=sha256:16f4a8403c9a86bdf430b9942f46557b2ab7884ace8e1ec40115ace2251c18c8

Observation 73f7c1a8-e954-4cd3-b94e-18816196ddd5 · outbound

This paper cites an unresolved cited work.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:44.524832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:44.524832Z digest=sha256:2f8700101160cb62e16624cdad50d1dbf3d5d63a5ce7bd13d993e47d2069aa53

Observation 05740471-e970-4649-8eae-b3c488fdc495 · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding TVQA: Localized, Compositional Video Question Answering

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:44.781343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:44.781343Z digest=sha256:ed4bf85dcf069692988da005c1908315c66b770397f1f38a266193da151fd46b

Observation dfe6d3b7-0d2d-478d-aad9-a23b8f8f4043 · outbound

This paper cites Finevideo.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Finevideo

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:44.904024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:44.904024Z digest=sha256:1aa96ae0d7c7b3db113914e5e3e33f04c565514b92d2b96ecf528e7be8491882

Observation 63fe83c7-4305-4236-8c2d-31689e10f41c · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:45.027159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:45.027159Z digest=sha256:ad1fb9696de87510f7216e4a72e377b9025f02546593497ef4dad4bb8d50fa67

Observation 44f6e26a-3390-417a-a7dc-68dcecdda5ec · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:45.094167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:45.094167Z digest=sha256:84ac018fe421998ed2185fda249143cfe75fbaed5609f2fd3046883d22f843c6

Observation d7aad7ba-4663-404f-86ec-3a64470c839f · outbound

This paper cites TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:45.137483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:45.137483Z digest=sha256:5cd13cf735028b8a6485e9234a87b25ecd7adcd4325db1b0ad00912f95f58e4c

Observation 933de6f8-6ff6-4e45-937e-4ccaa01c61c8 · outbound

This paper cites Tenenbaum.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Tenenbaum

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:45.214475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:45.214475Z digest=sha256:a19254224d56b0157379d650d7e16a5314e91e1a5e1311c4303a52fa5d5bbdf0

Observation ddffa66c-6b18-4507-8f6a-462f594e4986 · outbound

This paper cites Exo2ego: Exocentric knowledge guided mllm for egocentric video understanding, 2025.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Exo2ego: Exocentric knowledge guided mllm for egocentric video understanding, 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:45.353879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:45.353879Z digest=sha256:35ef7dcc43080075e49ef90f0290cbc41ea86b7daca60c9ad7785e0ec350470e

Observation 29679ae7-b005-4d36-ab11-651511c0719a · outbound

This paper cites Motionbench: Benchmarking and improving fine-grained video motion understanding for vision language models.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Motionbench: Benchmarking and improving fine-grained video motion understanding for vision language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:45.449671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:45.449671Z digest=sha256:57c8d20fbdb11e71094e69982447a281c256b020ae725899346fbb03ae176e20

Observation 12757dcf-7036-41ce-b4ac-9ceee6d0bd8b · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single object tracking.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Lasot: A high-quality benchmark for large-scale single object tracking

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:45.524820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:45.524820Z digest=sha256:92e8c8f19eede57975cc8142d703d3eda433b4d545c96df3a0fb936f91252b97

Observation 92921bda-1059-4880-aed9-db3eb7094c78 · outbound

This paper cites Got-10k: A large high-diversity benchmark for generic object tracking in the wild.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Got-10k: A large high-diversity benchmark for generic object tracking in the wild

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:45.610802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:45.610802Z digest=sha256:292c9cf6f9c2e8bcd8b06c2174c4bf766d97ebef4de92091e9855810057f2f29

Observation 68d5b2d2-ca47-477e-9453-4f25bba290d2 · outbound

This paper cites Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijaya- narasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, and Jitendra Malik.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijaya- narasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, and Jitendra Malik

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:45.825885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:45.825885Z digest=sha256:70e351c189985512a7cf692e65fd0d776952011886c3178c2db7d2e85f6c99cb

Observation 797fe49c-7856-46dc-a804-f06fab2fc2cd · outbound

This paper cites Egovqa-an egocentric video question answering benchmark dataset.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Egovqa-an egocentric video question answering benchmark dataset

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:45.896415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:45.896415Z digest=sha256:b132e9a21f722b75e01d95585246a9550d2a22e2aa289bad2a9fbb9af3dc3329

Observation 28367e16-b65d-4c7e-84da-1042e0956299 · outbound

This paper cites Learning transferable temporal primitives for video reasoning via synthetic videos, 2026.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Learning transferable temporal primitives for video reasoning via synthetic videos, 2026

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:45.961453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:45.961453Z digest=sha256:f89f58c85af0ab129675dadf1ed9df77fe22623d0ea2cfa63d22c918782a8faa

Observation e21f2666-0810-4545-b42d-eb15d92e4cfc · outbound

This paper cites Tomato: Assessing visual temporal reasoning capabilities in multimodal foundation models.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Tomato: Assessing visual temporal reasoning capabilities in multimodal foundation models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:46.065996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:46.065996Z digest=sha256:219d930b7ee6bfabfb997e13a6b89fbb6134f34b78ec74360831c6b5d7a07d0d

Observation 8e3ba013-2d83-496e-9ee5-82869176d604 · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:46.152973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:46.152973Z digest=sha256:4f450ab3dd34c029018385b32d7750626290feee2862ebe5eb6751affa43545a

Observation cc81c15b-8707-4181-a9b9-f212d9ff8779 · outbound

This paper cites Tempcompass: Do video llms really understand videos? In Findings of the Association for Computational Linguistics: ACL 2024 , pages 8731–8772.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Tempcompass: Do video llms really understand videos? In Findings of the Association for Computational Linguistics: ACL 2024 , pages 8731–8772

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:46.342096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:46.342096Z digest=sha256:a0648d0b4ac313e3d906ddaf7e64a2bd5f4e08b56be573a3021ed714e1f0c003

Observation 57d2e2c4-13c4-4e1d-a7e6-9879a171b493 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:46.436665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:46.436665Z digest=sha256:2a39d6999ff3cae07b722aa927359135574d94b361312989cc33712141e069eb

Observation b1af6863-d03b-4ec8-8da1-c04869640ed1 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:46.573335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:46.573335Z digest=sha256:b4fb218deff0c1fd24a345eecf7744bc1f1f7a9dbbd5b41d5666e1117743b32d

Observation 84ee4e48-c29d-4129-aa2e-0182b72b4e4e · outbound

This paper cites VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:46.835407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:46.835407Z digest=sha256:feeac1fdbbc462c2fe607d6bb8d6eccd58006dbf50ee7654c21cd43c6467245f

Observation 1a0e508e-08bc-41c5-a574-c1a893454cb8 · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:47.002048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:47.002048Z digest=sha256:11f41bf028b10e394a51f41e8eff1b25d33d1db30a7f81e4fb6668b5796186eb

Observation 6c9669d4-2c06-4e4e-bba6-df62de1d7772 · outbound

This paper cites Mmvu: Measuring expert-level multi-discipline video understanding.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Mmvu: Measuring expert-level multi-discipline video understanding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:47.229883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:47.229883Z digest=sha256:7dcc7ae75b2c7267b9d866faed4c9fbbf610b1298557ffa13f19df0b32338c74

Observation 6cf8ffcd-ab76-4593-9b85-c9087da6b0f0 · outbound

This paper cites Minerva: Evaluating complex video reasoning.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Minerva: Evaluating complex video reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:47.313301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:47.313301Z digest=sha256:ed15a25b7c128a5fe9c06d3578b1f2b5954acacbf8d45083ff69d6438e17e217

Observation f09832a2-f0ac-4d92-aefe-e83c8dbc9759 · outbound

This paper cites Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:47.434527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:47.434527Z digest=sha256:31b6c4f1771aa8215a4b94e3a6d120574ae89317672aa3efe7ee991739163c1f

Observation cb923698-d405-4226-9789-6d1f8bdb5cc4 · outbound

This paper cites Tall: Temporal activity localization via language query.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Tall: Temporal activity localization via language query

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:47.600495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:47.600495Z digest=sha256:1c7891f51bd27e973fe5f6715fd15003281a3672577b9cb7340de312a1e13743

Observation 2685caf0-64d8-4b74-8a34-57bfd9f32434 · outbound

This paper cites Dense-captioning events in videos.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Dense-captioning events in videos

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:47.761237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:47.761237Z digest=sha256:f0b8a86e4b7e51506fc005f3145bb3e8eb78f0e50774da96ad6f96f5b7484b1f

Observation 3ec15f71-9897-4104-ab87-3af549c81da9 · outbound

This paper cites Detecting moments and highlights in videos via natural language queries.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Detecting moments and highlights in videos via natural language queries

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:47.805666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:47.805666Z digest=sha256:eeb5a7e8de63fd8d10a95bd6c549920ac45eb2fa4b59ccec918b2c8ef837a4b7

Observation 8536bac6-8a77-49a4-be09-1afb22f38f08 · outbound

This paper cites Vidi: Large Multimodal Models for Video Understanding and Editing.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:47.844379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:47.844379Z digest=sha256:3ef9808ff090d916584039588b363847f1b8ba89e2141b190335e396f42675f6

Observation 40348e98-5b43-4527-8e75-a38ed0aeb684 · outbound

This paper cites Vidi2: Large multimodal models for video understanding and creation.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Vidi2: Large multimodal models for video understanding and creation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:47.884749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:47.884749Z digest=sha256:834d348b9a13dd983360d7da6bde18c381997cac09ec0d67c00c6339c9179288

Observation 54903a52-59db-465b-91bf-80471d354aa0 · outbound

This paper cites Momentseeker: A task-oriented benchmark for long-video moment retrieval.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Momentseeker: A task-oriented benchmark for long-video moment retrieval

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:47.923648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:47.923648Z digest=sha256:f32748e6a314ce15d7c80f0f15de2ce1ef452ebe464669d34ac837cd8d2df4c8

Observation adf89ab0-e959-4a7b-84ad-7c967815939f · outbound

This paper cites GPT-5 system card.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding GPT-5 system card

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:47.963677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:47.963677Z digest=sha256:f7da82c42dc38090f1633f03e40f2eafa57ce186fd1ebf6061d26274612fa6dc

Observation 7e079e2a-41aa-4261-81bf-284bf742459e · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:48.014595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:48.014595Z digest=sha256:4ecea331fef8fb8d1a58770335ef725535a3e94c1a5620e3691713b45ff18aa5

Observation 527104fb-143b-4554-aacb-595ec68ac764 · outbound

This paper cites Introducing claude sonnet 4.5.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Introducing claude sonnet 4.5

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:48.141435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:48.141435Z digest=sha256:4f1df53d80fc958d9583025f974a727d56d2e2e59998d96f5f59e79e2e61a72d

Observation 29f11d55-cb24-4841-99a3-6ac23f236f6a · outbound

This paper cites MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:48.181447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:48.181447Z digest=sha256:ae3b7bea5e1957d418283c9b49c24e7c55f87ef50eb6237634ab6febff735bc9

Observation 4bb62b33-3fc7-4574-9664-dd78ac3c8c3b · outbound

This paper cites Eagle 2.5: Boosting long-context post-training for frontier vision-language models.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Eagle 2.5: Boosting long-context post-training for frontier vision-language models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:48.248621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:48.248621Z digest=sha256:f859faae693bf336501dac3830a921d83b7b7efa6400528b28837422ac289071

Observation 4f6503ae-9746-4fe7-918f-cbccd66be913 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:48.337537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:48.337537Z digest=sha256:2bb7b4a11795c67701cf77ce1563aa497fb3a4d92a3c2502e3255f7f26ff7cc0

Observation 34055a81-c4a9-4c9c-957e-3431b35f15f0 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:48.427210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:48.427210Z digest=sha256:43188a4ea9eda3153f3e90607dda142e93534ce583f58f163b9a7d439982e16e

Observation bb5c00dc-8b5c-4208-a414-9e766e2d3148 · outbound

This paper cites an unresolved cited work.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:48.525948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:48.525948Z digest=sha256:390e7c82fcbe6eef905f616b3110801f5af59136973134a02ba8ba60b729768a

Observation 45268e19-3c2e-44e5-8398-aa21ee8d91ec · outbound

This paper cites Streamingbench: Assessing the gap for mllms to achieve streaming video understanding.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Streamingbench: Assessing the gap for mllms to achieve streaming video understanding

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:48.594780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:48.594780Z digest=sha256:4c4c5ba4d1f060fb886735643b8ab644f606f041c7674acb55e7b3f21d2ca196

Observation d7a17479-bc58-446d-8e99-0bea900400df · outbound

This paper cites River: A real-time interaction benchmark for video llms.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding River: A real-time interaction benchmark for video llms

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:48.759089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:48.759089Z digest=sha256:9113dbb89e5d8007bbf0b0658e8968d488419d56e717e6bff7e12eecd7bc08e1

Observation e8ee7c1c-267d-4f8c-b05d-d9d00aca60ac · outbound

This paper cites Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:48.868441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:48.868441Z digest=sha256:85350da50d9fd8ae9006a2741a14921b815440b489cfdc9a177c7527cae124cd

Observation 7523c331-160b-4a39-b74a-911205ccb356 · outbound

This paper cites ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:48.909159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:48.909159Z digest=sha256:7b451536efb41719f0f5069c3fc2fe141a65e7f1bc8b6ade0334cdfb76b1bb28

Observation 2a95b4f6-12f8-44f1-b33f-20cffae494f8 · outbound

This paper cites Livecc: Learning video llm with streaming speech transcription at scale.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Livecc: Learning video llm with streaming speech transcription at scale

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:49.000448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:49.000448Z digest=sha256:a423925cc1b20798e087d7fff9d0954f06de7fe3ad0fb7428c532f1c43dd9453

Observation eef673cb-ab0d-41eb-96c0-09b8a853e2c3 · outbound

This paper cites Timechat-online: 80% visual tokens are naturally redundant in streaming videos.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Timechat-online: 80% visual tokens are naturally redundant in streaming videos

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:49.159431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:49.159431Z digest=sha256:8bc9adb8d8e1ec8aa4c4594d3ff0ae076af378404141006b48cd074262c02079

Observation f7dbb6e1-fc38-43aa-a75a-89864475aee9 · outbound

This paper cites StreamingVLM: Real-Time Understanding for Infinite Video Streams.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding StreamingVLM: Real-Time Understanding for Infinite Video Streams

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:49.226362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:49.226362Z digest=sha256:e8a8e3649c6938dc2d5a0336de31fc2504d010e014da5813bcb88c4ec0213d35

Observation 1a57edb3-07ad-4d15-83ec-8fb67b370d25 · outbound

This paper cites Mm- duet2: Enhancing proactive interaction of video mllms with multi-turn reinforcement learning.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Mm- duet2: Enhancing proactive interaction of video mllms with multi-turn reinforcement learning

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:49.297464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:49.297464Z digest=sha256:d8a94686cd8bf1d7c8717c9a124e646a19ba8cee41b696943fc0fd9175c37744

Observation 43fc2d0a-4e5a-49ff-8900-0cd58958772e · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:49.338761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:49.338761Z digest=sha256:c11c52eac4f6735e2d4c2e2f347d630602b17bd8a9afd4951570e9a36c3e10fb

Observation 34cdd0cd-7ef9-42b9-8cc2-6c5f5665e243 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:49.393750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:49.393750Z digest=sha256:401f6d2b6c7225a7f8844f1fa0ed524dfa55bc664de83de6996ae3d3be9a9cd5

Observation 52d95c55-6cca-42a4-8338-c243094e56f0 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:49.433839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:49.433839Z digest=sha256:684df46013e83153c9eeba4e4820ea2250fc9da995df432d1e3ee67e6f89a4cf

Observation 5a1eb399-e080-47c6-b0d6-af99858ce804 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:49.479276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:49.479276Z digest=sha256:eff6fd6ebae881f1288fe9341989c8b4a84602e9c50393d5ea10817b460e78fa

Observation b2fcb43b-74dd-4bf0-bd01-d9d6382cb93f · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:49.528523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:49.528523Z digest=sha256:ff497de41baac3957ca429a86b4e2d8e82768727061944da385b69585bb3837a

Observation 916cc133-aeb9-4a69-85be-5014ca2f787c · outbound

This paper cites LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:49.651285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:49.651285Z digest=sha256:63d122260f400df8ab21d90b6c95e130b903ae4407384bf950c49e7eccf5daf6

Observation 5939ccc1-f8be-42e1-91cf-a1dc4dec07c9 · outbound

This paper cites AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:49.747631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:49.747631Z digest=sha256:6986ecf37b2e9caa00ae5b3c3add2f5210a5c75e4f54a2731bdb67bad47cd17a

Pith citing papers

No inbound Pith citation observations are available.