Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T22:00:28.350003Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 100 of 298 outbound references and 2 inbound Pith citation observations for arXiv:2606.07433.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T22:00:28.350003Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:41:18.289865Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-11T00:36:56.481588Z
100 of 298 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 470fec84-5101-4207-a561-20e5f4adaad3 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Qwen3.5: Towards native multimodal agents,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1a03ebf-7729-44c5-91ac-d82454cae18e · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Qwen3-Omni Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7c42e758-6134-4eae-a7bd-475225cd6e4f · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Qwen3-vl technical report,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6410d80e-c6ff-4ba8-85ad-5d2a89bc9f61 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Qwen2.5-Omni Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ae7501ad-991c-4e6e-8fac-cec4f55cf572 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dcf79bbe-8e66-4445-b164-6e4cb83da00c · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Msr-vtt: A large video descrip- tion dataset for bridging video and language,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04f357ea-a740-4cd0-8c2e-085458660645 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video question answering via gradually refined attention over appearance and motion,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70b39ce8-58f1-439d-9f5e-d2abb3c9de62 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Tgif-qa: Toward spatio-temporal reasoning in visual question answering,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be9e0995-0aa6-46e3-a5a4-d021b08da945 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs LongVU: Spatiotemporal adaptive compression for long video- language understanding,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 677d658e-a194-4cff-8db7-3ed14280f2c0 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b9cb58ef-339d-40ea-bf26-e2da5cad6ca0 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 85920590-edb0-4d63-8a45-ec6f1905523a · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Ego4d: Around the world in 3,000 hours of egocentric video,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f94a8dbe-cb34-452b-bc54-d7321e746e94 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs A visually grounded language model for fetal ultrasound understanding,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8713d225-e2c1-47b9-9fb7-02420a9ec188 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Streamingvlm: Real-time understanding for infinite video streams,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99f19cce-399a-45ea-af49-7a991abf30fc · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 32b32525-956f-4587-8a61-1c3826126fd5 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Adaptive keyframe sampling for long video understanding,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 448d42e9-20d1-4e16-a889-747a316d02eb · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs FrameFusion: Combining similarity and importance for video token reduction on large vision language models,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70f7cc58-b488-4f44-92b4-c764110c54cb · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Moviechat: From dense token to sparse memory for long video understanding,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a8fa6e-0ea2-424e-bad9-8af29e776310 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b2e50ed-d4b3-4c13-8886-42123c4e3bbe · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Memory-enhanced Retrieval Augmentation for Long Video Understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ace06fdb-019e-44c5-aecf-8cefd999604a · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 71c727b2-aebc-4c6f-ae89-7fa506eb5885 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cf307e9e-eb4a-4574-a123-362d95f657cc · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 86725e39-b5db-470b-94c2-6a5be27362f8 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d9bc7633-e4fb-4873-8808-3c15b431dd7b · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Open-o3 video: Grounded video reasoning with explicit spatio-temporal evidence
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 26674536-e164-4883-88af-f7ed786e4eae · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 534f1959-5ab9-4481-86a3-16af66200fc0 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-language understanding: A survey from model architecture, model training, and data per- spectives,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7256a394-d786-4b1a-97b0-e67e3f4961f0 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video understanding with large language models: A survey,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49bafe87-e9b4-4b20-b944-ddd57ffc1cf8 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs A survey on video temporal grounding with multimodal large lan- guage model,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab7bb81b-feb6-4cb1-b6ad-8dd26adc9593 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-lmm post-training: A deep dive into video reasoning with large multimodal models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d4c59d4a-3ef9-4e6d-8198-32ccac65e378 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs A Survey of Reinforcement Learning for Large Reasoning Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 65f00e7b-dba5-4b01-9f86-3ff893e99903 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Memory in the Age of AI Agents
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ed615186-123c-4b51-a335-d212b8a0ef71 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs arXiv preprint arXiv:2505.18227 , year=
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3db0dc68-3a88-4d5a-b518-32cb56fcfb3b · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d6210a4c-4b8d-4fb4-97d3-46ea3dde4d8b · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Lita: Language instructed temporal- localization assistant,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e096ad9f-002b-4963-94d1-a56e9187972d · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Universal video temporal grounding with generative multi-modal large language models,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d94bdf2-c212-44cd-803b-289b88d55d3f · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs arXiv preprint arXiv:2512.14698 , year=
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6774cf53-d140-4d87-9dfa-02fc6a1bdd31 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Towards one-to-many temporal grounding,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d88011c-296d-4909-bf1d-4faf9afd5e13 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2f565c03-1547-4ca0-a376-663912c12253 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Sama: Towards multi-turn referential grounded video chat with large language models,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8188500-5fc2-45dc-bdef-1a93e263223d · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Streaming dense video captioning,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f46ae04d-e8d2-4d1b-bc2f-fa8ddec060a7 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Do you remember? dense video captioning with cross-modal memory retrieval,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e7d640b-f394-4356-885a-76ad781b6cae · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary en- richment and online refinement,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1db51307-ad73-434f-9901-01772a1abe2d · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 00fc8319-181c-4837-ad2f-97b1f9e2570d · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d344e5c1-beb9-4a5f-a9ab-a1c298b9f080 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a3e788fe-5b9e-482f-a1bd-9f06988f7cb3 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Baichuan-omni technical report
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 02327fc4-30ea-44bc-9b36-2beebe1677fc · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Ming-Omni: A Unified Multimodal Model for Perception and Generation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation eb3d44dd-5b67-4797-96c8-e90f5dc44120 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b3a8cf48-e086-43d3-bc8e-60ddc7414b0b · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f4e8a51d-2404-4008-af5e-42bb65278d08 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs OmniCaptioner: One Captioner to Rule Them All
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ec4f2d40-b768-4f6c-99bc-dd8fe6a0e163 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Omnivinci: Enhancing architecture and data for omni-modal understanding llm
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation de4ae604-e237-4b7a-9565-2c18234cc729 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c035664f-04b6-46ee-8fde-057eea9f1ce2 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs DyCoke: Dynamic compression of tokens for fast video large language models,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 009b482a-7cd2-4c0a-a22e-50cb3ec90cd4 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Streamingvlm: Real-time understanding for infinite video streams
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 25a0b264-e534-4b6f-9a11-ec6ff6e28042 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Vtimellm: Empower llm to grasp video moments,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 122cca4f-cfcf-49b6-b085-ecbf659a35cb · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55f91030-75e7-4d02-9ae3-047ddde04ebb · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Distime: Distribution-based time representation for video large language models,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 732cb7e7-3395-41d0-b089-50e7272270c8 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Self-chained image- language model for video localization and question answering,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b2b1842-8324-47dd-9e39-7e9d9215ef89 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 84a2c07b-a05d-4151-8dfd-a752d5421d95 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Timesuite: Improving mllms for long video understanding via grounded tuning,
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99336aa3-e415-44dc-8026-07c0c700cff4 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Scanning only once: An end-to-end framework for fast temporal grounding in long videos,
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7e3701d-ac02-4841-86b0-48089a54eb63 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Trace: Temporal grounding video llm via causal event modeling,
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8c49ad2-98fd-4b20-b646-437a84e9f354 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 43060d12-bce0-4242-89ee-08bd37db8ef6 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bf592336-9039-4686-9448-fc6a9d60a463 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Momentor: Advancing video large language model with fine-grained temporal reasoning,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22782f23-b055-40fe-a3fd-480129dee614 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Luowei Zhou, Chenliang Xu, and Jason J
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cf68d152-da0c-4486-9e6e-bc9893b58257 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 34ae96de-f14d-4c2c-a5ae-ac121e656b05 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1c4b1f02-8d40-4948-9db3-8f3101b27eee · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Videozoomer: Reinforcement-learned temporal focusing for long video reasoning
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d83e18d7-2dd2-4760-ab08-85c1c7832d6b · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Datasets and recipes for video temporal grounding via reinforcement learning,
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cba4df5-6108-4d84-b38a-8c1e04efdca5 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d24a5721-fb42-4ae0-a17f-db1a69665c82 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs arXiv preprint arXiv:2510.12798 (2025)
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9f873f25-c3dc-4730-a153-b7eca279b160 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs arXiv preprint arXiv:2511.21375 (2025)
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bc8ac887-80ac-4566-b640-d4e7ba88adf3 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Universal instance perception as object discovery and retrieval,
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38760e07-9fad-4d9c-9ad7-1b5480c8b3c8 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Multimodal Referring Segmentation: A Survey
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a6671a1c-94cd-4ff9-8395-227a8aacc08a · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Deformable DETR: Deformable Transformers for End-to-End Object Detection
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8f74da62-f36e-4bda-9e73-0b8dfd4e9c90 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Lisa: Reasoning segmentation via large language model,
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9037922-680a-4b5e-9933-fda8ab7a1e73 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding,
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 979b6460-4c29-4ef7-be79-b4af886f6a9f · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs High-Quality Entity Segmentation and Grounding
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 17791f4d-f18e-4477-91de-7bb3273a8790 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs SAM 2: Segment Anything in Images and Videos
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1a18f4a7-0e7f-4b1d-bee8-ae5063d4a4fb · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Unipixel: Unified object referring and segmentation for pixel- level visual reasoning,
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89f988ab-d5e2-40f5-9bf2-850cd4b214bc · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Samtok: Representing any mask with two words.arXiv preprint arXiv:2601.16093, 2026
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d572f914-b166-4af9-b5ea-d87b7587650a · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Collecting highly parallel data for paraphrase evaluation,
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e50f0776-31b5-4c2f-afa7-736e19cab602 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-llava: Learning united visual representation by alignment before projection,
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 325f0750-8865-48cf-b56b-997cb7b9819e · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1c250328-5078-438c-b812-bd495f678941 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video recap: Recursive captioning of hour-long videos,
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74bbc708-9bc2-4185-a15a-3c23914329cf · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e9841d58-d9f5-4f29-93ec-bd548fc09fd3 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Fine-Grained Captioning of Long Videos through Scene Graph Consolidation
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6909c904-fb97-4a95-aee1-5bbac09e37f0 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 577b781c-615a-4add-b747-514c1e9853a1 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs video-SALMONN 2: Caption-enhanced audio-visual large language models
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c8439cfd-3612-4e2a-a1f4-fcd335241947 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a29d4349-76e2-4025-9e38-e08a8d8b8536 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1d92c611-654f-45e6-ac0f-5eeb48ac7f38 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Towards fine-grained human motion video captioning,
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29488250-0ac9-421f-936e-ed56697a307b · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Sharegpt4video: Improving video understanding and generation with better captions,
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92c72d79-9b4a-4015-9f16-f2bf4b205e71 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Panda- 70m: Captioning 70m videos with multiple cross-modality teach- ers,
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0be33b41-f6ea-4f99-add3-b7a71a23750c · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Vript: A video is worth thousands of words,
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 695a5423-fe49-4fbe-8a57-6e6b5c895561 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs IF-VidCap: Can video caption models follow instructions?
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 280879c3-079f-47ae-bd7f-b70571d57359 · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Any- cap project: A unified framework, dataset, and benchmark for controllable omni-modal caption- ing
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation af362dc5-c511-4ee3-b77c-0495ce0eebac · outbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Intentvcnet: Bridging spatio-temporal gaps for intention-oriented controllable video captioning,
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2acb1284-9429-4c5d-be3c-714d5ec81a97 · inbound
LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents Watch, Remember, Reason: Human-View Video Understanding with MLLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d613dd20-b8ac-47c9-abcb-98a6b7bc44a4 · inbound
StreamFlow: Dynamic Memory Flows for Streaming Video Understanding Watch, Remember, Reason: Human-View Video Understanding with MLLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.