Pith. sign in

Paper Citation Record · LEDGER

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

As of 15 August 2026, this Paper Citation Record lists 100 of 298 outbound references and 2 inbound Pith citation observations for arXiv:2606.07433.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.07433 v1

Coverage vector

measured 100 of 298 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T22:00:28.350003Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:41:18.289865Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T00:36:56.481588Z

Reference resolution

100 of 298 outbound references displayed

  • verified exact47
  • verified fuzzy0
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 470fec84-5101-4207-a561-20e5f4adaad3 · outbound

This paper cites Qwen3.5: Towards native multimodal agents,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Qwen3.5: Towards native multimodal agents,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:5afcf993aa7efe70779b4bddf369246ad60c21a2e3c6ba4bad8a5d1b3dbc5470

Observation b1a03ebf-7729-44c5-91ac-d82454cae18e · outbound

This paper cites Qwen3-Omni Technical Report.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Qwen3-Omni Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.631234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:a28600fa0c52a0c772090f7f9f37275a5541f2fe5faca89fb51f98c0aa7c7192

Observation 7c42e758-6134-4eae-a7bd-475225cd6e4f · outbound

This paper cites Qwen3-vl technical report,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Qwen3-vl technical report,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:88412f2dfca6f4830696e56389ddf8d702342f589f3071f3b53b2be207c5eaba

Observation 6410d80e-c6ff-4ba8-85ad-5d2a89bc9f61 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Qwen2.5-Omni Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.619936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:317380219a78cce3a9526ac8b47e8711edd132581f1ef5abf6e1280dcbafb185

Observation ae7501ad-991c-4e6e-8fac-cec4f55cf572 · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.628146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:ef0c89f2f5bfce7e2a1786c66e93d9fa89e0a57ae0a9c474cf7b733563449c65

Observation dcf79bbe-8e66-4445-b164-6e4cb83da00c · outbound

This paper cites Msr-vtt: A large video descrip- tion dataset for bridging video and language,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Msr-vtt: A large video descrip- tion dataset for bridging video and language,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:efcc894f1200803f897b8bd29cda5a28e35fe92ff27535e4533d1dbc7465c49a

Observation 04f357ea-a740-4cd0-8c2e-085458660645 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video question answering via gradually refined attention over appearance and motion,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:b3658b4a234db0bb881894ea9ace5787a2c2172c49dda55f723082dc0eeb9f66

Observation 70b39ce8-58f1-439d-9f5e-d2abb3c9de62 · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Tgif-qa: Toward spatio-temporal reasoning in visual question answering,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:52dd58dca9a2cf5ae96ce5eea09a62f2a26058158bfdeee2207589ca374b307c

Observation be9e0995-0aa6-46e3-a5a4-d021b08da945 · outbound

This paper cites LongVU: Spatiotemporal adaptive compression for long video- language understanding,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs LongVU: Spatiotemporal adaptive compression for long video- language understanding,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:aa5e74f440b7793bc946021e9c269ce68f99c88f212e2f8d848e5f30bb79134a

Observation 677d658e-a194-4cff-8db7-3ed14280f2c0 · outbound

This paper cites LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.676404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:f27a3b86d563c05aa95123247a02b8dee4e42b7934e03b4e0d60ba15358d0c7e

Observation b9cb58ef-339d-40ea-bf26-e2da5cad6ca0 · outbound

This paper cites Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.494462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:ca53531756ab1e62952c2dba690b1c2e292b5e5bc2acb8a610cc6f959c414576

Observation 85920590-edb0-4d63-8a45-ec6f1905523a · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Ego4d: Around the world in 3,000 hours of egocentric video,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:da938c373e3905e3ffeb26867499f064e354e6cc3fc06500b943b614a1b15f04

Observation f94a8dbe-cb34-452b-bc54-d7321e746e94 · outbound

This paper cites A visually grounded language model for fetal ultrasound understanding,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs A visually grounded language model for fetal ultrasound understanding,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:bee9934aad5d3a80c968b95ea7329f36e90c9645bb42052f21045d8b3bcb20cc

Observation 8713d225-e2c1-47b9-9fb7-02420a9ec188 · outbound

This paper cites Streamingvlm: Real-time understanding for infinite video streams,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Streamingvlm: Real-time understanding for infinite video streams,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:00232ea0d31d024a636a68f5656ec25e6d55bdf751197ee52d9952c8bb3585d5

Observation 99f19cce-399a-45ea-af49-7a991abf30fc · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.504136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:0e92a37b3c6028c524bce193d05ef263b86948ac90e8956455c5a71da2a8189f

Observation 32b32525-956f-4587-8a61-1c3826126fd5 · outbound

This paper cites Adaptive keyframe sampling for long video understanding,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Adaptive keyframe sampling for long video understanding,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:1214c7788608f26eb4fb4cdcdf2d037500c6b9b08f6626d796ed0d8160d1370f

Observation 448d42e9-20d1-4e16-a889-747a316d02eb · outbound

This paper cites FrameFusion: Combining similarity and importance for video token reduction on large vision language models,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs FrameFusion: Combining similarity and importance for video token reduction on large vision language models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:75c9eeeeb56bc4e21a458bb954257eb134c05550304b93fc935e15452f8e8169

Observation 70f7cc58-b488-4f44-92b4-c764110c54cb · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Moviechat: From dense token to sparse memory for long video understanding,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:b9e80c2b5146302256c4f20bcb2c3d84ce8ec7b3060662be998869dd8a16049e

Observation 41a8fa6e-0ea2-424e-bad9-8af29e776310 · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:8c30d8e3703dbabd4dc3263f6946aadb98ae5cbd16e4b772fdbe9642a2fcedfa

Observation 2b2e50ed-d4b3-4c13-8886-42123c4e3bbe · outbound

This paper cites Memory-enhanced Retrieval Augmentation for Long Video Understanding.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Memory-enhanced Retrieval Augmentation for Long Video Understanding

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.682433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:3e6bc62bb5a331c0c42980a472e26c46e96b315a7c2bc559c2e859f4572dccc4

Observation ace06fdb-019e-44c5-aecf-8cefd999604a · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.497535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:8c60d636f1df5120475d143472374cf19d5aa942e32499df8f5c2c8aa94ebda2

Observation 71c727b2-aebc-4c6f-ae89-7fa506eb5885 · outbound

This paper cites StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.536238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:d5f1f4b75e2df03bd313c65e431e2fdfd918213cbf84164d062dbcd7b1751570

Observation cf307e9e-eb4a-4574-a123-362d95f657cc · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.613385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:96c596a4364a2f2e19f0f97201c9896f88f32c3f612f1cc4e74e2982ed238cd8

Observation 86725e39-b5db-470b-94c2-6a5be27362f8 · outbound

This paper cites Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.657429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:c4dfd666a5dc3cbe2600b924d3bd716ddcbae43260b90f9144f6cd8e5138be5e

Observation d9bc7633-e4fb-4873-8808-3c15b431dd7b · outbound

This paper cites Open-o3 video: Grounded video reasoning with explicit spatio-temporal evidence.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Open-o3 video: Grounded video reasoning with explicit spatio-temporal evidence

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.796210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:18fc4bad78dd14b5c69694bcd763b6df0c67cae5c18dde7e4bdd1f9cf92b8e3c

Observation 26674536-e164-4883-88af-f7ed786e4eae · outbound

This paper cites Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.802070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:060c621550be4f9706d318b8a7e60a10da20811bd7bafa6cb4b61943e4773311

Observation 534f1959-5ab9-4481-86a3-16af66200fc0 · outbound

This paper cites Video-language understanding: A survey from model architecture, model training, and data per- spectives,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-language understanding: A survey from model architecture, model training, and data per- spectives,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:b791b8c1aa2be625b9de0fdaecca20019592a728167d137a30486f73f97e65cd

Observation 7256a394-d786-4b1a-97b0-e67e3f4961f0 · outbound

This paper cites Video understanding with large language models: A survey,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video understanding with large language models: A survey,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:b27ae012d5702c1693f119b6a8c8e689a8063e69844d73b966dc21f69c5f0fcd

Observation 49bafe87-e9b4-4b20-b944-ddd57ffc1cf8 · outbound

This paper cites A survey on video temporal grounding with multimodal large lan- guage model,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs A survey on video temporal grounding with multimodal large lan- guage model,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:445fc560564d6723d5f3be66483491d0c3d31a66751dd580db0bad95b945f5c8

Observation ab7bb81b-feb6-4cb1-b6ad-8dd26adc9593 · outbound

This paper cites Video-lmm post-training: A deep dive into video reasoning with large multimodal models.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-lmm post-training: A deep dive into video reasoning with large multimodal models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.640149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:bb892ec190086d0463ce1edda06b0ece7940b624e56970646a1a44224ee5479c

Observation d4c59d4a-3ef9-4e6d-8198-32ccac65e378 · outbound

This paper cites A Survey of Reinforcement Learning for Large Reasoning Models.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs A Survey of Reinforcement Learning for Large Reasoning Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.500849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:8249a4437da65620defb274355c9c2e3719afec79f606c8b566faf798f42a0d7

Observation 65f00e7b-dba5-4b01-9f86-3ff893e99903 · outbound

This paper cites Memory in the Age of AI Agents.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Memory in the Age of AI Agents

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.031096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:a72bd5182b925f73a9254af90ba6cd7efc36ab0acc42d36cc335c9d88f6e21bd

Observation ed615186-123c-4b51-a335-d212b8a0ef71 · outbound

This paper cites arXiv preprint arXiv:2505.18227 , year=.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs arXiv preprint arXiv:2505.18227 , year=

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:27:15.679384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:538b863b331f0031b3f34869cf616a3d7e001330b6f5d5421a4a6982ddbd399a

Observation 3db0dc68-3a88-4d5a-b518-32cb56fcfb3b · outbound

This paper cites Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.037520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:572f89cbc8721aeb22c5532e7a57f4adf5699604a793e783c5bbabcfae3770c1

Observation d6210a4c-4b8d-4fb4-97d3-46ea3dde4d8b · outbound

This paper cites Lita: Language instructed temporal- localization assistant,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Lita: Language instructed temporal- localization assistant,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:422863741c60586666d4f9928af6b85f799578c2f86594907becaed25c53d7b8

Observation e096ad9f-002b-4963-94d1-a56e9187972d · outbound

This paper cites Universal video temporal grounding with generative multi-modal large language models,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Universal video temporal grounding with generative multi-modal large language models,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:f7118ffa37ce359403c11b47e4b9a008db74a9512d35be4337d82d83656b5e66

Observation 3d94bdf2-c212-44cd-803b-289b88d55d3f · outbound

This paper cites arXiv preprint arXiv:2512.14698 , year=.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs arXiv preprint arXiv:2512.14698 , year=

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:37:14.030570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:ad6809e2517f872f78a462948cd32eadbc68591357ed3829795bcc29fcdbf8ac

Observation 6774cf53-d140-4d87-9dfa-02fc6a1bdd31 · outbound

This paper cites Towards one-to-many temporal grounding,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Towards one-to-many temporal grounding,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:9d5b9ff13ef8f9327c219fcf497a1a4b3e5adf36ba1bc03748b144ae064a3427

Observation 9d88011c-296d-4909-bf1d-4faf9afd5e13 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.041253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:4d9ad3b97414b4094780360f76bdeff624b8b71bc7374059652f537798aca61b

Observation 2f565c03-1547-4ca0-a376-663912c12253 · outbound

This paper cites Sama: Towards multi-turn referential grounded video chat with large language models,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Sama: Towards multi-turn referential grounded video chat with large language models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:d601e35c750852b9ff02fdf8733bc043681d1b1d1047a80ae783e8647d80c0e9

Observation e8188500-5fc2-45dc-bdef-1a93e263223d · outbound

This paper cites Streaming dense video captioning,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Streaming dense video captioning,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:46630cc3068bba75f24ae020ca5adeff9b539c5bb9c42adbdd8dadbfda5a8ee9

Observation f46ae04d-e8d2-4d1b-bc2f-fa8ddec060a7 · outbound

This paper cites Do you remember? dense video captioning with cross-modal memory retrieval,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Do you remember? dense video captioning with cross-modal memory retrieval,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:f1f0217a15e928e50bf5c0617361767104b7cdf887f5a9573310a9072260f556

Observation 3e7d640b-f394-4356-885a-76ad781b6cae · outbound

This paper cites Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary en- richment and online refinement,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary en- richment and online refinement,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:18a5241bc3b6e437edc652d6589de8d527d459798d69f3ee0e01ecd3f919bb22

Observation 1db51307-ad73-434f-9901-01772a1abe2d · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.625038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:06c49246a01432c3d4410c4d1dafea4891d14d512292de0e37eaea24e0928a30

Observation 00fc8319-181c-4837-ad2f-97b1f9e2570d · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.575211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:b2e290cf99977c693e17436fad8fb7b0b41ae92fb632fc45ce02db7a7678eb85

Observation d344e5c1-beb9-4a5f-a9ab-a1c298b9f080 · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.563714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:035201b1dbee0c46d52f620f075f222056601bffcf76a129349b255b5ed40be8

Observation a3e788fe-5b9e-482f-a1bd-9f06988f7cb3 · outbound

This paper cites Baichuan-omni technical report.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Baichuan-omni technical report

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.519812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:0e03cbc3d738c43819fc8f8036e0094592ecc6177acb68b325c87120433ac169

Observation 02327fc4-30ea-44bc-9b36-2beebe1677fc · outbound

This paper cites Ming-Omni: A Unified Multimodal Model for Perception and Generation.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.025507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:6a4cd402a7579659e5b7db137aad708016696d208d80cbafbfd5031c086cdc1d

Observation eb3d44dd-5b67-4797-96c8-e90f5dc44120 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.662681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:d803a027b3684634a05cca95732734e45045b06752e6e5a4a4cdd76e654e9e1e

Observation b3a8cf48-e086-43d3-bc8e-60ddc7414b0b · outbound

This paper cites Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.634543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:778e9bdb28657113035a0da887c552aa106f676b0a4bf60c4b5b6ca7a15b27cd

Observation f4e8a51d-2404-4008-af5e-42bb65278d08 · outbound

This paper cites OmniCaptioner: One Captioner to Rule Them All.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs OmniCaptioner: One Captioner to Rule Them All

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.021888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:8ce3465fba267d61bd6e0463d04a89ef092d01a5da43130d895629a2b09b1093

Observation ec4f2d40-b768-4f6c-99bc-dd8fe6a0e163 · outbound

This paper cites Omnivinci: Enhancing architecture and data for omni-modal understanding llm.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Omnivinci: Enhancing architecture and data for omni-modal understanding llm

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.027098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:76a6cf2b9fdeafc381d0da1fdb470ba52889a67022c13f55b11d2f5ea3991167

Observation de4ae604-e237-4b7a-9565-2c18234cc729 · outbound

This paper cites Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.057718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:844670388b18a4e769da0481d11e7f4adc7e61e02b57211f0fbf0581c743ace1

Observation c035664f-04b6-46ee-8fde-057eea9f1ce2 · outbound

This paper cites DyCoke: Dynamic compression of tokens for fast video large language models,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs DyCoke: Dynamic compression of tokens for fast video large language models,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:9868fb4ed003bbdbd6c18b83e84226d8f4a128f3638416bb0e7b90901478e8a8

Observation 009b482a-7cd2-4c0a-a22e-50cb3ec90cd4 · outbound

This paper cites Streamingvlm: Real-time understanding for infinite video streams.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Streamingvlm: Real-time understanding for infinite video streams

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:27:15.041754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:1ca915a385d71a12bb97cce9c1c46a65b5b9ae9afb754f24477b50fdcf42c696

Observation 25a0b264-e534-4b6f-9a11-ec6ff6e28042 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Vtimellm: Empower llm to grasp video moments,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:2c5bc79dac8338b4093c6641b48078e8a410c6e057b4de8aa84f4f7883bdcafa

Observation 122cca4f-cfcf-49b6-b085-ecbf659a35cb · outbound

This paper cites Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:9d3ed27843567013ceea9d06003ffcb818b397b9b8ba73e217f1fd37b46c3f0f

Observation 55f91030-75e7-4d02-9ae3-047ddde04ebb · outbound

This paper cites Distime: Distribution-based time representation for video large language models,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Distime: Distribution-based time representation for video large language models,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:960227d6a0ba9af8e263c488d679c47953936d0fde5a3c730008ab749556a37c

Observation 732cb7e7-3395-41d0-b089-50e7272270c8 · outbound

This paper cites Self-chained image- language model for video localization and question answering,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Self-chained image- language model for video localization and question answering,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:580693a38953dabfac8df6c4f7854bb30a1e49cfdd79cfbeb5b23c6886b70af2

Observation 0b2b1842-8324-47dd-9e39-7e9d9215ef89 · outbound

This paper cites LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:14.990641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:1e4b0fec965e81e8bc535b81dac81cfea09f50fdb116e215fad379c95eb74a03

Observation 84a2c07b-a05d-4151-8dfd-a752d5421d95 · outbound

This paper cites Timesuite: Improving mllms for long video understanding via grounded tuning,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Timesuite: Improving mllms for long video understanding via grounded tuning,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:500211947185b664faeb4ef0564e22aa9c917f75550b64d8fcaa23052628d2ed

Observation 99336aa3-e415-44dc-8026-07c0c700cff4 · outbound

This paper cites Scanning only once: An end-to-end framework for fast temporal grounding in long videos,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Scanning only once: An end-to-end framework for fast temporal grounding in long videos,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:999d28d264fda4a3a9a96d7d83c656c700913ed0439139e6ba35e7c43b078b72

Observation f7e3701d-ac02-4841-86b0-48089a54eb63 · outbound

This paper cites Trace: Temporal grounding video llm via causal event modeling,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Trace: Temporal grounding video llm via causal event modeling,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:4adaa6d9590507d5cfe2784248f0efa75015f67eb84cabf4698ee0aff79009ef

Observation c8c49ad2-98fd-4b20-b646-437a84e9f354 · outbound

This paper cites TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:14.976786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:0ef9be58763612dd74402dd01b7e822c8cfc8475bd720f59b6645a60b9263819

Observation 43060d12-bce0-4242-89ee-08bd37db8ef6 · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:14.973292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:f0d225f190162d31b93ae7a21e3a24f0c2036c0a18edf8059ec4f768a17a2f53

Observation bf592336-9039-4686-9448-fc6a9d60a463 · outbound

This paper cites Momentor: Advancing video large language model with fine-grained temporal reasoning,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Momentor: Advancing video large language model with fine-grained temporal reasoning,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:5dc8b64e0f07ecb4c3906f28fdc88fded94e6f4bf76eeb58bb495732009556e3

Observation 22782f23-b055-40fe-a3fd-480129dee614 · outbound

This paper cites Luowei Zhou, Chenliang Xu, and Jason J.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Luowei Zhou, Chenliang Xu, and Jason J

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:27:14.983423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:d3dd3086df2c251af2e243b59a6ed0bdca6fc0343f4f6270ef017c67e41e455b

Observation cf68d152-da0c-4486-9e6e-bc9893b58257 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.836506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:0eb0b1d43156c759abc2490a51c50cd3569e006c811a2dd3cf4bbe7b4cea9bf9

Observation 34ae96de-f14d-4c2c-a5ae-ac121e656b05 · outbound

This paper cites Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.826646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:003e0a2c444155d8e768930cff94cdd0f67061149eeb827b48c47eed15d56086

Observation 1c4b1f02-8d40-4948-9db3-8f3101b27eee · outbound

This paper cites Videozoomer: Reinforcement-learned temporal focusing for long video reasoning.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Videozoomer: Reinforcement-learned temporal focusing for long video reasoning

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.831821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:c9a34e32ba3ba26ad49b90f605de19499845f47e327e935d5a70f60277c2e4c4

Observation d83e18d7-2dd2-4760-ab08-85c1c7832d6b · outbound

This paper cites Datasets and recipes for video temporal grounding via reinforcement learning,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Datasets and recipes for video temporal grounding via reinforcement learning,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:437f4e3f0f8db55244c32d202b5e6994fdea0ed66a4da56c516dcbc8f4554ebc

Observation 0cba4df5-6108-4d84-b38a-8c1e04efdca5 · outbound

This paper cites MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.834103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:501bc3b7f57d50d4af1d6bba47852ee4c4cd2567a35103285afb36a6591d5943

Observation d24a5721-fb42-4ae0-a17f-db1a69665c82 · outbound

This paper cites arXiv preprint arXiv:2510.12798 (2025).

Watch, Remember, Reason: Human-View Video Understanding with MLLMs arXiv preprint arXiv:2510.12798 (2025)

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:27:15.819395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:ef812ab915c4fe099bbc94960f83417420c62e80c53620ac119df18f079e9d2a

Observation 9f873f25-c3dc-4730-a153-b7eca279b160 · outbound

This paper cites arXiv preprint arXiv:2511.21375 (2025).

Watch, Remember, Reason: Human-View Video Understanding with MLLMs arXiv preprint arXiv:2511.21375 (2025)

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:27:15.821898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:78e9dfbfc69424de13e5fbafaab015a8cb23d02b9f6705358ae947203e888f86

Observation bc8ac887-80ac-4566-b640-d4e7ba88adf3 · outbound

This paper cites Universal instance perception as object discovery and retrieval,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Universal instance perception as object discovery and retrieval,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:3982b8393d4efaf5f856c6e681ec44ac86e0ecb90d00a364ab2aee1589f7bad5

Observation 38760e07-9fad-4d9c-9ad7-1b5480c8b3c8 · outbound

This paper cites Multimodal Referring Segmentation: A Survey.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Multimodal Referring Segmentation: A Survey

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.816944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:d971d5eb6723ebec30d1929c110b69521031441b0d54bdddb3a2dbd53d0df2af

Observation a6671a1c-94cd-4ff9-8395-227a8aacc08a · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.843435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:d07b9d5dee40d782283136b5b946ba5057c6d078c37d06fb0143c23af90a9e47

Observation 8f74da62-f36e-4bda-9e73-0b8dfd4e9c90 · outbound

This paper cites Lisa: Reasoning segmentation via large language model,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Lisa: Reasoning segmentation via large language model,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:d8a574f6b3944babd89e0cb1f06c7747b3b9e878e427d70d2cfb1521e8c60956

Observation e9037922-680a-4b5e-9933-fda8ab7a1e73 · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:372060a834188cabae044d3cf0f1aca94c03a678b799467d9c936fd4ddc9276d

Observation 979b6460-4c29-4ef7-be79-b4af886f6a9f · outbound

This paper cites High-Quality Entity Segmentation and Grounding.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs High-Quality Entity Segmentation and Grounding

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.788784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:53f714ca4d76396e16a26fc010c637227c8935b27ae7039a198f77d2ea56b677

Observation 17791f4d-f18e-4477-91de-7bb3273a8790 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs SAM 2: Segment Anything in Images and Videos

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.838832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:7b3016ad18e1651bcb06f53fa04632733b4efc3f232707e705e75b8cc4818a58

Observation 1a18f4a7-0e7f-4b1d-bee8-ae5063d4a4fb · outbound

This paper cites Unipixel: Unified object referring and segmentation for pixel- level visual reasoning,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Unipixel: Unified object referring and segmentation for pixel- level visual reasoning,

Reference 83

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:475a09632f1ffdda47e9ad96ccaa2b8b20f9db1a52f61bb0bf900b2c19165da2

Observation 89f988ab-d5e2-40f5-9bf2-850cd4b214bc · outbound

This paper cites Samtok: Representing any mask with two words.arXiv preprint arXiv:2601.16093, 2026.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Samtok: Representing any mask with two words.arXiv preprint arXiv:2601.16093, 2026

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.036900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:05576d58c25bcd2b115f27198fdb92c2af8ccd89d2f8bbfb2308231f2eaa5b6f

Observation d572f914-b166-4af9-b5ea-d87b7587650a · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Collecting highly parallel data for paraphrase evaluation,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:f5273b3514e8448133b1071db3524d64dce9b608df1c520c3c501ba01bb4b439

Observation e50f0776-31b5-4c2f-afa7-736e19cab602 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-llava: Learning united visual representation by alignment before projection,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:06beda09cf9f4dc58fad4cd7b3c9f63ccf2b9c9d159d343c045808be8cf683dd

Observation 325f0750-8865-48cf-b56b-997cb7b9819e · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.561094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:1436bfa16b146cb36e92af645fc35512159f47d5aa5f32c78bcd1d46b40afd7d

Observation 1c250328-5078-438c-b812-bd495f678941 · outbound

This paper cites Video recap: Recursive captioning of hour-long videos,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video recap: Recursive captioning of hour-long videos,

Reference 89

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:cbe1e4a9bad0fe3cf0643aa5046f72d7c517b78e91acc3be8f67c980534cb97b

Observation 74bbc708-9bc2-4185-a15a-3c23914329cf · outbound

This paper cites LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.763039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:c239dd582352bf37dc2b865caa2f8254af690c51603e9c4483d988eade30bf03

Observation e9841d58-d9f5-4f29-93ec-bd548fc09fd3 · outbound

This paper cites Fine-Grained Captioning of Long Videos through Scene Graph Consolidation.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Fine-Grained Captioning of Long Videos through Scene Graph Consolidation

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:27:15.765503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:370a95b778b090511b4c5fb6eeb1dd703bb9e63b378674a8676653f7dae52a77

Observation 6909c904-fb97-4a95-aee1-5bbac09e37f0 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.770571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:dca06219ea8235e0c9d8f5dd27d83bad1c21b7fde6f159d09b1c3093e96a9c04

Observation 577b781c-615a-4add-b747-514c1e9853a1 · outbound

This paper cites video-SALMONN 2: Caption-enhanced audio-visual large language models.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs video-SALMONN 2: Caption-enhanced audio-visual large language models

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.539022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:49d32d18c1d3bd74528d0be856ab653fbf90bc76ecc9e2ddb14a81a089f24389

Observation c8439cfd-3612-4e2a-a1f4-fcd335241947 · outbound

This paper cites VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.755390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:fa09f6c5e80aeb0affb6feebada98d72b69f4d2a1e56cb94b202f856980972d3

Observation a29d4349-76e2-4025-9e38-e08a8d8b8536 · outbound

This paper cites OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:27:15.748175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:04982e2477264a573888195ef6fac2962f2eb4e7dde93cb8842fc88e5e3ae62a

Observation 1d92c611-654f-45e6-ac0f-5eeb48ac7f38 · outbound

This paper cites Towards fine-grained human motion video captioning,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Towards fine-grained human motion video captioning,

Reference 96

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:b5e99fd38eae83a2cd6cbb9de0d9b430f79f30244d64dcf48d018cc45b0dd050

Observation 29488250-0ac9-421f-936e-ed56697a307b · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Sharegpt4video: Improving video understanding and generation with better captions,

Reference 97

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:571e398b7863f9ab026b47904ed5c4ee020353e4a08231cd0cd9eb0098214adc

Observation 92c72d79-9b4a-4015-9f16-f2bf4b205e71 · outbound

This paper cites Panda- 70m: Captioning 70m videos with multiple cross-modality teach- ers,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Panda- 70m: Captioning 70m videos with multiple cross-modality teach- ers,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:4c0f924600c08a97b19b2fb3246795d57bda8171c5f555f0a40cbf1286f82ad5

Observation 0be33b41-f6ea-4f99-add3-b7a71a23750c · outbound

This paper cites Vript: A video is worth thousands of words,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Vript: A video is worth thousands of words,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:24c1931ff7e5efbc7a4f0a77c1c39a89916b9aa91482aa9f94361bda047257b4

Observation 695a5423-fe49-4fbe-8a57-6e6b5c895561 · outbound

This paper cites IF-VidCap: Can video caption models follow instructions?.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs IF-VidCap: Can video caption models follow instructions?

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.581109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:a24b6925956e183cf759c6326a711ca14b96df49121274c95f36dd40f89832cf

Observation 280879c3-079f-47ae-bd7f-b70571d57359 · outbound

This paper cites Any- cap project: A unified framework, dataset, and benchmark for controllable omni-modal caption- ing.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Any- cap project: A unified framework, dataset, and benchmark for controllable omni-modal caption- ing

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.745659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:6c6b35935ea5f1aaa129f12c5c5665dbb215569b1f7903c41fb7ec2c7620f97b

Observation af362dc5-c511-4ee3-b77c-0495ce0eebac · outbound

This paper cites Intentvcnet: Bridging spatio-temporal gaps for intention-oriented controllable video captioning,.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Intentvcnet: Bridging spatio-temporal gaps for intention-oriented controllable video captioning,

Reference 102

Resolution
unresolved
no resolver link, observed 2026-06-27T22:00:28.350003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:f40cd60233e8fe8975534821c545136b0001a63b8547ae40abe490d5f8cf6210

Pith citing papers

Observation 2acb1284-9429-4c5d-be3c-714d5ec81a97 · inbound

LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents cites this paper.

LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents Watch, Remember, Reason: Human-View Video Understanding with MLLMs

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:36:56.487289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T00:36:56.179636Z digest=sha256:d3e38c517f6e65c7b5c1644cf6905186de300b32258e908e244438df9e05484f

Observation d613dd20-b8ac-47c9-abcb-98a6b7bc44a4 · inbound

StreamFlow: Dynamic Memory Flows for Streaming Video Understanding cites this paper.

StreamFlow: Dynamic Memory Flows for Streaming Video Understanding Watch, Remember, Reason: Human-View Video Understanding with MLLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T13:41:18.289865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:41:18.289865Z digest=sha256:fb2cb8a3135ee845e7ad6d2130e76215776052e8402dc3a326c56fe4d098ca76