Pith. sign in

Paper Citation Record · LEDGER

Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2408.15542.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.15542 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:37:05.338028Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:29:15.346488Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6afa325e-c4aa-4373-887a-6e3feca6923d · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.200813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:606771f803975554d50b6faa4e6ebe75523ec6aced0c77b90bf76fdd7ee1d4c1

Observation fc2efdfa-89c7-4199-81f7-5ccd0e43f5b6 · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.448546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:ada2d323dae673708c0a30652130b4c012b4a456e8dc4e01e09e7b489a8171be

Observation 4c3f7916-40be-4917-86d3-699740523c97 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.355335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:4153a856cf18cebf14442c04228bac1eb9d94701adfbefa106a319dc7b35cfd6

Observation a75c095b-8f2e-4bbb-815c-15164036c109 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.287485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:861cdbd02b0ba3bcc1155b4af743b750bbef1310a3a0647cdcd2b7f41cdc2604

Observation 7b31e901-5b45-4c98-bc84-9819d5b33790 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:43:00.408504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:fc05d608bf15e5ec80d505449488090b56b80a98a62e9062c6c665fdc8732e0a

Observation 3cd6b174-5f79-4bb2-b775-89284f17b9df · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:23:51.734057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:a3883e1f8b7e90e9b99c05dbd6b4e381f8684074905ff102aa8e41a13be43e04

Observation c38d776f-7f12-467b-b91e-4bc3b14c4a9a · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:22:37.480720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:b5ea4445be3729349102a93d5829e45f2c511f8a792cdb3816bb3dfc4d549b70

Observation 0f4f6773-ee34-41bd-bd60-e24713ee2a90 · inbound

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams cites this paper.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:05.338028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:05.338028Z digest=sha256:4a173455a5e016e520853f94516aac398a06e007a8761ac5e6e3ecde5520d37a

Observation 79cdbcc8-6c07-48b9-9e72-e300b01626e0 · inbound

LAVA: Language Driven Scalable and Versatile Traffic Video Analytics cites this paper.

LAVA: Language Driven Scalable and Versatile Traffic Video Analytics Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:32.407540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:05:32.407540Z digest=sha256:35f73b605607b3ae92bda82e48c59bc744d8bc3c4739d11ab5da8e90a1df337d

Observation adce1126-a8a0-4236-b4ca-617e4f826596 · inbound

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models cites this paper.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.438516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.438516Z digest=sha256:ccb7fc696d2beeb39502232b8ba0d977093a95616193e801de72a890df2ddc42

Observation 83510f07-e3d2-4591-9f11-3089933d292e · inbound

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding cites this paper.

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:16.769169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:10:16.769169Z digest=sha256:23504f24180115a5f164684d26b6bdd2c651cc950b78b8f0a4e5da413ca55b01

Observation 9faf2417-4896-46ee-9883-69e0bdcc3e56 · inbound

CAViAR: Critic-Augmented Video Agentic Reasoning cites this paper.

CAViAR: Critic-Augmented Video Agentic Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.988625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.988625Z digest=sha256:527b069523c021bba0d42b7ca00d8b6ad3e81e444104ea4ff8e73ffc30cd284c

Observation b5c68ae3-7955-4fa6-b102-f464b8aa660a · inbound

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos cites this paper.

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:51:28.132505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:49:12.987772Z digest=sha256:7fe2d707b94367b5eed3a2e596ed84d935d979932b5e6d673311cd9fb3a1c640

Observation 8313bc12-2f58-443d-87b3-565b183979ff · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:53.873910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:0a7070a00bbd4aea72d6bee43c1a28322af0227147aac261f86e244040ea1aa6

Observation d1c7aabf-cc5c-4e9e-adc7-4a15863bba87 · inbound

Small Vision-Language Models are Smart Compressors for Long Video Understanding cites this paper.

Small Vision-Language Models are Smart Compressors for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:15:51.479344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:cad53682520112eb1587e63d5c12f3ddb499da0ff7f9f5cc4187dfe86c39de0f

Observation 82579acb-6fab-499a-978b-dcefacfa5f2e · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.569125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:90d6bbadc63042fc97c2052dcdcad988476561a787e5df44bb75f26efc5bd168

Observation 3e04ced6-a4e9-4ed8-b2f9-66302f455250 · inbound

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning cites this paper.

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.759603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T06:51:52.861981Z digest=sha256:b7e1221d9197e1e2908627b8ea4f9f0d02f73b8f1a10c877903023c9364b6916

Observation 92803c71-da0a-4e87-a538-1536a6e706fd · inbound

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration cites this paper.

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:21:09.522561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:11:11.410051Z digest=sha256:569b37c54ec08f044f8d7fcdea113114d71f5314950ea41b75592de019d12c6f

Observation c8418107-8250-4579-88e0-92a470464e44 · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:20:56.479634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T02:30:55.939351Z digest=sha256:ab2d71f88dee6937b20e1a013fcf2d5f60ddd71c379160e1a05e4b2ed75302cc

Observation c1005e51-57e9-4c04-9b1d-862c7d6bfbde · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.700487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:7a43561e957cdbbedce95b4d356e0d2982d8c2e5ef5aa4b9c5221f40af5144ea

Observation 88c56c0f-3dab-484a-878e-e22f4a3402a5 · inbound

An Efficient Streaming Video Understanding Framework with Agentic Control cites this paper.

An Efficient Streaming Video Understanding Framework with Agentic Control Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.452976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T11:30:22.151045Z digest=sha256:de78220cee7a50b4414eed35f5ebbd8c381db0799eb1c7c081fd239f9548eebb

Observation 06199d9c-9c90-4ba0-bca6-6dfab8cdebb0 · inbound

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering cites this paper.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.071338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:560defd9225ee5f40afe003ea0e7b05b515a586557666d04ca27a6a9b22a34d5

Observation 4f5b8170-f9d8-43a5-82cc-fa3598a11b04 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 296

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.192666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:15ad2b4fc0fac80d053ea74cacc9ae55bc61b3dd6b150fe74cd21bfc1d700f6d

Observation ed14d6ca-5e62-4cb0-9eed-d7f32f883594 · inbound

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs cites this paper.

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:58:57.726244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:02:37.632461Z digest=sha256:f8f9a7e6b1436d6fe7fa784b0155d363a30b54f917a8e442276661f85e4e3961

Observation a605ff3f-0bea-437f-a1a5-671a2adbaa57 · inbound

Native Active Perception as Reasoning for Omni-Modal Understanding cites this paper.

Native Active Perception as Reasoning for Omni-Modal Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:29:15.349201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T21:14:11.297384Z digest=sha256:031b1b856ffab52a890f62bb7cc1bbf350bab88ec7499a3c57ce82a50998d643

Observation 71114075-e508-49b7-8ba5-8ebae0751e93 · inbound

Native Active Perception as Reasoning for Omni-Modal Understanding cites this paper.

Native Active Perception as Reasoning for Omni-Modal Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T10:57:19.810078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:57:19.810078Z digest=sha256:caf2fe7b29bfd2fecc6b6efb628bb3b33ae2fee173e988d5f6b8023cbe9594d0

Observation 01d453a6-8134-402f-ba94-496cc1c508df · inbound

Latent Visual Cache for Video Reasoning cites this paper.

Latent Visual Cache for Video Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:9a43bc18773e4a0e4f6e49728c90f0da15e5ba318a096860b4ae99e1fba51d59

Observation 30cccac3-b3ac-40f9-adef-3c640cb6616b · inbound

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization cites this paper.

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:13.271024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:13.271024Z digest=sha256:fc97a03bcc5ddd0d76b26beabfaccb04e638b8b210c3378b4a1c2ecc46e29724