Pith. sign in

Paper Citation Record · LEDGER

Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 59 inbound Pith citation observations for arXiv:2408.15542.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.15542 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 59 of 59 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:35:52.439764Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:29:15.346488Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6afa325e-c4aa-4373-887a-6e3feca6923d · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.200813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:896ca993bdf1d333c6913e41aa1e35f59f41f6731797d859001aea2738c9d9d5

Observation 2251fc28-348e-46fc-91e3-1c30075ad71d · inbound

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation cites this paper.

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T16:42:55.831551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:42:55.831551Z digest=sha256:e7cee3e2632d8a3f7362e76da4d34a18081633d194b19a8a50f018ea4b23ff0e

Observation 2416ca54-bd0d-4ed3-9b0c-db82703579c4 · inbound

ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models cites this paper.

ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:13.327097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:13.327097Z digest=sha256:dc23da4911724d09bba3c4449e906977e28926bf1225f354e8d3d71af8ca528c

Observation 0660e0e1-ea37-49ca-9754-7c06418decc0 · inbound

TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability cites this paper.

TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:56.143385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:56.143385Z digest=sha256:858876019077db67ebff5c2538733a9b0bd77a217191bb591093c09bbff484ae

Observation b304a8e9-c627-465c-9a19-19f0c90d267f · inbound

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation cites this paper.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.883956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.883956Z digest=sha256:e51b632883e227c1fdac2443d3736f1517dd87abbe341983537af16b8aa90b53

Observation 90257aaf-f10c-41f9-8290-643d4af0669a · inbound

VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation cites this paper.

VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T04:56:01.317796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:56:01.317796Z digest=sha256:263d45a70b09f58d91f9eb7d25173dffd2e77b2820be546cb7ab424feb0555e3

Observation 25e68ef2-663e-4d2e-931e-53e167d8ace5 · inbound

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information? cites this paper.

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information? Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T23:19:11.558964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:19:11.558964Z digest=sha256:2fdf23964e5d4c9c506c41f396c0ffbf5b6bc89d5ba075cc528be7ce9f67221e

Observation 3b77870b-956d-45b3-9b29-202d1f92d7df · inbound

LinVT: Empower Your Image-level Large Language Model to Understand Videos cites this paper.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.256483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.256483Z digest=sha256:6e6da0579e4957f5abf986fa2cbe983a33c7960919999e11669a65d7a97a2e7a

Observation 5fe818af-c160-4879-946e-57db2d2bb695 · inbound

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM cites this paper.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.836999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.836999Z digest=sha256:ef3a194fc431e4808161f2fefb4da4f458a4495279365329036fbfad3d7596ef

Observation 9e9f8983-2816-491e-a802-47cb02a7a7e5 · inbound

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions cites this paper.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.276119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.276119Z digest=sha256:3aa24b3e784db728bd894f0ebca2870f1ba1db2fa33d5074784ba51404948252

Observation 942f3a43-a41a-4e8b-a493-028820c034e1 · inbound

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models cites this paper.

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:42.508497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:42.508497Z digest=sha256:fe6ab14cda55a95d7c4245d5bcc3025e42dceef59b1c1671065d0d956db719f8

Observation 2b5693f0-fcc0-4efd-a127-70584d444670 · inbound

Apollo: An Exploration of Video Understanding in Large Multimodal Models cites this paper.

Apollo: An Exploration of Video Understanding in Large Multimodal Models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T16:11:10.557687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:11:10.557687Z digest=sha256:2727bdac134b61f21c9d55b8d8b494f177c9bc9c8c7697af81f6d694defee3d9

Observation 1ad0e2d1-bf2a-4d42-8477-199a3c4c9914 · inbound

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding cites this paper.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:51.006992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:51.006992Z digest=sha256:ec3007ecaa0f8fbc889eb489eec9f30a9ffa85d01c6b23fee10c35397e624b23

Observation be6c2b11-9589-42d5-a18e-c8b6131ba895 · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:17.549297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:17.549297Z digest=sha256:bc44ad18879da78341102fa253ae1a193346045ce536c42e0e91ed96c836ebfd

Observation fc2efdfa-89c7-4199-81f7-5ccd0e43f5b6 · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.448546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:b34432bf63b008a1a0dcc6896ca831b95a217688f62ee844d188dd0f4a33c25d

Observation 1b56fba2-4ee0-4dae-8050-b7b41f78140f · inbound

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding cites this paper.

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:56.499473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:56.499473Z digest=sha256:e732921e869b47ccd21596847b4d2202f5db534afc1bf3b6f664aa3954fea94c

Observation 4c3f7916-40be-4917-86d3-699740523c97 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.355335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:4d49ef8bf3502e2f89f3822340a5f3687187366bb2150499cb3d3133069441d9

Observation f9545335-a006-4f6c-be0a-6d137fbca0ff · inbound

ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark cites this paper.

ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:25:47.090873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:25:47.090873Z digest=sha256:65fee4a699c7009cab5d794fbe55258bc8ba1d6ae8e3150f13327d794847be05

Observation 8ddac59e-487a-4567-8482-dfe89c57a8dd · inbound

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design cites this paper.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.028341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.028341Z digest=sha256:d55711bdc99451a481dc2545e0030b77308f2c8d83ebd46188c0a6ab40bba5d4

Observation a75c095b-8f2e-4bbb-815c-15164036c109 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.287485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:551c0b6c227dea6a3cff67787d3a4a196c426ba3fc0377b5398447e20d6c9820

Observation c71c0487-de8e-48a9-ac5f-a66dde5d4b27 · inbound

Temporal Preference Optimization for Long-Form Video Understanding cites this paper.

Temporal Preference Optimization for Long-Form Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.195645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.195645Z digest=sha256:2f9abef96265ee1fef7b0951bc99b3fe0310dfec7468ce2bfb9ca3133ca87710

Observation ff23c886-4446-4908-a7b9-a8bf1dff7b4f · inbound

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding cites this paper.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.852698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.852698Z digest=sha256:d1f392224a47bc04bb848aaa94694eb01e5efe4f5ede76b9f8fc4b9b7f0633f2

Observation 3ab614dd-274e-44cb-9262-13cb022a3a00 · inbound

CoS: Chain-of-Shot Prompting for Long Video Understanding cites this paper.

CoS: Chain-of-Shot Prompting for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.940927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.940927Z digest=sha256:f8310ad627746670801abc36846743dbf2a8c89f4425b814f27b86c14306e902

Observation 7b31e901-5b45-4c98-bc84-9819d5b33790 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:43:00.408504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:41791e453e9b79cba344460c53b5d5f99a916afe7988df768c8c9bafd9cab401

Observation 3cd6b174-5f79-4bb2-b775-89284f17b9df · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:23:51.734057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:b5e8fedf57efd0aa7aab4540e10a45661eebb577870dc725410879a8f6c7f85e

Observation 8335557e-cb75-467b-b018-6d8712a4a5b0 · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.694251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.694251Z digest=sha256:5a75c6069dbffda3c552072b4e83ff2ab61db9afe7887ae9847afb60dd093abc

Observation 81ccf352-fc30-40ad-a176-ed4d766704f4 · inbound

Clapper: Compact Learning and Video Representation in VLMs cites this paper.

Clapper: Compact Learning and Video Representation in VLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.733144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.733144Z digest=sha256:01ccd32ae78adb030216edbaea2432abc00bcafaec4c06f65231dcecdb753994

Observation 06c314bb-4e5b-4358-8610-210596f6fd2f · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:01.765688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:01.765688Z digest=sha256:764c815e62a40df3597ad162b71dde451c657cb35bc63276406554949f091924

Observation a3fcd63b-ceaa-46f9-bd7f-2fcd422481f8 · inbound

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration cites this paper.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.535106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.535106Z digest=sha256:908bde25227e8764d1b24889aa0d12de03b3a6e8d96ada67fabeb88795422db6

Observation 6443ad05-adcd-41d2-b2e6-1dc1fa460b2e · inbound

Reinforcing Video Reasoning with Focused Thinking cites this paper.

Reinforcing Video Reasoning with Focused Thinking Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:13.508706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:13.508706Z digest=sha256:895cb4293826279e872d17c3261758a9ee6432f4067e8bd12bbd412968699c06

Observation 45e15b40-6cc5-4825-8b1a-1dea06352fcd · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.645844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.645844Z digest=sha256:da74c6d7139fa77c72e8500a03d416cfa9a61388c28800e246c90ccdc7f23477

Observation de8b4712-1d7a-4f71-9dcf-3612323dec24 · inbound

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding cites this paper.

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:27.753715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:27.753715Z digest=sha256:b6938ebcc15f150100956261b987dd84de754eeffaddcd2d24fa72ad064945de

Observation 82c0dd4b-7efa-417e-8b24-09987e4f0f69 · inbound

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding cites this paper.

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:03.120334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:03.120334Z digest=sha256:bd78da0b37ed9f0f89bb770a406f8baa1e8d5d1b9660d04ec746bde104aaf89c

Observation c38d776f-7f12-467b-b91e-4bc3b14c4a9a · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:22:37.480720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:6ca56110acddee97f6c9b516e644bd4689f5df9deba7a22066e056bd193ac0c9

Observation c7ba70d3-efd9-41c4-bb52-345f72de1a67 · inbound

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency cites this paper.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:46.174675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:46.174675Z digest=sha256:924510ca3857707db6af0fda692e23b9887494e2a2ea36a823e7eb6367698b19

Observation d2d036f9-03e3-4183-a4e4-e3dade6832b8 · inbound

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning cites this paper.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.373080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.373080Z digest=sha256:6ab910cce4be4d16c8a9ab65a7c2580062c5d8667d3f2bd126f368f85866f06f

Observation 0f33d806-78f9-4ccf-9bd3-34caa259491c · inbound

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs cites this paper.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.351183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.351183Z digest=sha256:8d64836e78d0a786089efac5fef6c60ba8d39155ce7a79d39c23abe06c0ba7e1

Observation 0f4f6773-ee34-41bd-bd60-e24713ee2a90 · inbound

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams cites this paper.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:05.338028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:05.338028Z digest=sha256:30a244f33f0e99c29ade07b0a4ee209dbac25deb273463c789659da3f38b11c5

Observation 79cdbcc8-6c07-48b9-9e72-e300b01626e0 · inbound

LAVA: Language Driven Scalable and Versatile Traffic Video Analytics cites this paper.

LAVA: Language Driven Scalable and Versatile Traffic Video Analytics Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:32.407540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:05:32.407540Z digest=sha256:ce57e25be40312e09ef0c9fed6b7a6a8b33f93f1da0880f49e12a12f6874afe6

Observation adce1126-a8a0-4236-b4ca-617e4f826596 · inbound

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models cites this paper.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.438516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.438516Z digest=sha256:0002eb958596e6f80d128fdc0a46777b64d0fc4000d34493306b79ae7f7b9643

Observation 83510f07-e3d2-4591-9f11-3089933d292e · inbound

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding cites this paper.

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:16.769169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:10:16.769169Z digest=sha256:de61c224e2900e42e1757cd470d957ca2d5b56dbb8796505330ee7dda96e1ec9

Observation 9faf2417-4896-46ee-9883-69e0bdcc3e56 · inbound

CAViAR: Critic-Augmented Video Agentic Reasoning cites this paper.

CAViAR: Critic-Augmented Video Agentic Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.988625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.988625Z digest=sha256:92c13fdd28a2a41cdbccd83af6734eeb7eb7462b5827c354e18ad6cff19e96dd

Observation b5c68ae3-7955-4fa6-b102-f464b8aa660a · inbound

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos cites this paper.

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:51:28.132505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T02:49:12.987772Z digest=sha256:e838242c0ed2b357e0b042a25e3ebc634ed5aab623c614cd48ff1d357913e8bb

Observation 8313bc12-2f58-443d-87b3-565b183979ff · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:53.873910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:a5df01deed5af52a5e9c0eee0ec14e2648de6d2cb7f023e485219ce15b18fd83

Observation d1c7aabf-cc5c-4e9e-adc7-4a15863bba87 · inbound

Small Vision-Language Models are Smart Compressors for Long Video Understanding cites this paper.

Small Vision-Language Models are Smart Compressors for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:15:51.479344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:06b490da06f8b40081d1c87bb74aff3f9aca7342467fd60840775cf44a70bdb6

Observation 82579acb-6fab-499a-978b-dcefacfa5f2e · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.569125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:7bbff0e5c6c7929bd967d45928b941df08a6043552680f9aaaf0cc8ac9d9957e

Observation 3e04ced6-a4e9-4ed8-b2f9-66302f455250 · inbound

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning cites this paper.

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.759603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T06:51:52.861981Z digest=sha256:f680be0c798ac7f435cb4dbb503eec31d546104452038fec3ac5fc51e5d09ff3

Observation 92803c71-da0a-4e87-a538-1536a6e706fd · inbound

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration cites this paper.

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:21:09.522561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T20:11:11.410051Z digest=sha256:c9fc5109f92c6501b774eb35dbf8d983d75be3af3348efcacc43fb71174d01ab

Observation c8418107-8250-4579-88e0-92a470464e44 · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:20:56.479634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T02:30:55.939351Z digest=sha256:23a660ce350c1931fd507b8bfe8f8df1c0cc77e38d2fb55fac21810575b39caa

Observation c1005e51-57e9-4c04-9b1d-862c7d6bfbde · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.700487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:6aa7c0e558c90a9534a3889378733118358432bc174e39b11ee6b534f83a6da2

Observation 88c56c0f-3dab-484a-878e-e22f4a3402a5 · inbound

An Efficient Streaming Video Understanding Framework with Agentic Control cites this paper.

An Efficient Streaming Video Understanding Framework with Agentic Control Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.452976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-20T11:30:22.151045Z digest=sha256:9f6aef233ff023d041d58216b19d3390fec2c4df8ba40025e350f726d4489574

Observation 06199d9c-9c90-4ba0-bca6-6dfab8cdebb0 · inbound

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering cites this paper.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.071338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:c8378aedb62b518e6e0d975db6ad0b53cb2db9a6db925dedb351e2073d716813

Observation 4f5b8170-f9d8-43a5-82cc-fa3598a11b04 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 296

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.192666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:6b359bd325ac3da5681054de04b212a00608ebb6d66ab579d8cfb720a8beda10

Observation ed14d6ca-5e62-4cb0-9eed-d7f32f883594 · inbound

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs cites this paper.

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:58:57.726244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T01:02:37.632461Z digest=sha256:5e998ff63a8edaccdd866daf97dcf6607f7f3438cb53176a82986a9a0999179b

Observation a605ff3f-0bea-437f-a1a5-671a2adbaa57 · inbound

Native Active Perception as Reasoning for Omni-Modal Understanding cites this paper.

Native Active Perception as Reasoning for Omni-Modal Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:29:15.349201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T21:14:11.297384Z digest=sha256:2e4a08badb9cf961343136f2a1c1f7afe559d1bb6fdd887bd2e96e3aa777ee83

Observation 71114075-e508-49b7-8ba5-8ebae0751e93 · inbound

Native Active Perception as Reasoning for Omni-Modal Understanding cites this paper.

Native Active Perception as Reasoning for Omni-Modal Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T10:57:19.810078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:57:19.810078Z digest=sha256:30485c3041bb310fe06aa596b456a4b0e16565644ed2a02ee138e785f61f0593

Observation 01d453a6-8134-402f-ba94-496cc1c508df · inbound

Latent Visual Cache for Video Reasoning cites this paper.

Latent Visual Cache for Video Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:3a98a47c46c22eca76d256185942c03a1ec7fdff5aee936c56abe6e456daf165

Observation 30cccac3-b3ac-40f9-adef-3c640cb6616b · inbound

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization cites this paper.

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:13.271024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:13.271024Z digest=sha256:34708c0e28445dbd14acc978711d4073ae7b4735d2e2a51b75cd2ec52aae4ddc

Observation 650dd8cb-943e-4d56-a45d-42a97fce4584 · inbound

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models cites this paper.

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 223

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:52.439764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:35:52.439764Z digest=sha256:0ce13b6eff584397bc4cf6ed6b121d33fa89b51b652db8318cfe58987d06f4bd