Pith. sign in

Paper Citation Record · LEDGER

InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2403.15377.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.15377 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:29.351108Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.518508Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0ea28090-2a9b-430f-a743-afc02aa2e0ea · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.244814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:8724b32b2b9a6dad0e77d34d6d5368a2cf72ed83a2630cec81f9f0bb89990eb9

Observation 6482212f-2fea-4ee8-af46-15f76d04a960 · inbound

Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation cites this paper.

Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:40:00.050951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T14:39:59.870039Z digest=sha256:d55fcdc904cc943a553ef946f12bb50e083f6e983e5c3e8ca06c17c53400840d

Observation 67648cd9-4a98-42e9-9df2-3427496945b5 · inbound

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding cites this paper.

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:53:33.707536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:53:33.585035Z digest=sha256:54040070ebec1d84caba7330878a7b1362d21b9b39db4277de70317e3553c800

Observation be600ca6-544c-4333-806e-450eb4add321 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 257

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.143903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:08724d7992b8d71cc747c25603c6b4e70127d718670495d4e10172a4551bbf78

Observation 9b3a2c62-41ad-4638-849f-bff192cadec6 · inbound

DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization cites this paper.

DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:02:41.738013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T06:57:50.897865Z digest=sha256:e48b8f2df1f12b97231025ef65fabb78f8941d579a78d1c13f89db0fe2cc1cea

Observation d4ddbc8d-88bb-4485-b83d-33783b4f2547 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.358964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:e47631e3916b9ef13262522b8c8bb26835128a2b50a7a00514653f2fb7bf9a68

Observation e76c4e79-5a2b-4e07-83b2-a35350833ea4 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.554003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:b5eded5531c7cc550ef792fc31a64a931aac0466eafbd394002f77591a23c38e

Observation fe182bdf-a1eb-452c-ab36-07d2a8315a47 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.742495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:6eb6d26770baf145ecc8a06bca900ef4aff6f9eb3e5312361d7331cfccdd47d2

Observation b85f535a-1e35-4405-ac04-20df1c5043b4 · inbound

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks cites this paper.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.351108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.351108Z digest=sha256:872c9052b31d36d03802e8428566be57e2f8e74dc9a54f3027de042ff5f60e5c

Observation fbc88a1a-335d-41cc-9183-f0aab1822dcd · inbound

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models cites this paper.

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:10.374002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:24:10.374002Z digest=sha256:4d002236dffff103671c41638b403d6b4f334c51884eaf5e7188d8d2a5608aae

Observation 4f82f8a4-75f2-4b0b-bd9f-12cdf4e124f5 · inbound

HCQA-1.5 @ Ego4D EgoSchema Challenge 2025 cites this paper.

HCQA-1.5 @ Ego4D EgoSchema Challenge 2025 InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:47.260706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:47.260706Z digest=sha256:4ab28944f055b866fef4de41cbb46b4be4e63f926a8ad13740d0392572c1b8c6

Observation c51851bd-3b04-4f8f-9f48-14366c20f7cd · inbound

HuMoCon: Concept Discovery for Human Motion Understanding cites this paper.

HuMoCon: Concept Discovery for Human Motion Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.401308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.401308Z digest=sha256:4a287db5af7607e79a3e807e65c079bb85e6452a1975842582c74cf2f8ec71cb

Observation 5c096c6d-ac30-465c-a225-b79619168202 · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.499105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.499105Z digest=sha256:858e9d3cc92b761111fbbf50ced8ff89c2404e91e073fe303a1e41486a7353b7

Observation 34c1518b-6a96-4e58-ad3a-81998148f13c · inbound

VideoMolmo: Spatio-Temporal Grounding Meets Pointing cites this paper.

VideoMolmo: Spatio-Temporal Grounding Meets Pointing InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:30.926631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:30.926631Z digest=sha256:48ff2c361a779e261d9f1f4bb40a0f36f1a4fd67b73cbc4af22a662ce11bb3eb

Observation a206936a-a89b-4446-b0ca-5da22eb7bbec · inbound

An Empirical study on LLM-based Log Retrieval for Software Engineering Metadata Management cites this paper.

An Empirical study on LLM-based Log Retrieval for Software Engineering Metadata Management InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:59.061572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:59.061572Z digest=sha256:14aa15c556426f645a233a9cd60390fe0f36892902c10a0c85f72e514b4f9f74

Observation 30471e61-9c6b-4718-a4bb-197fbf6e0fd0 · inbound

DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification cites this paper.

DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:51:56.851307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:51:56.851307Z digest=sha256:67e30f9c6fd56ebce5b2072fb08d1c4b1e8075db34e8b18b14e4f4e4c05c1d98

Observation 0c6c2b82-1304-4e98-8301-7975bacc3446 · inbound

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization cites this paper.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:00.161920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:00.161920Z digest=sha256:281c054908cf0595cf93ff32839100527bf12c5763f4f696ff8b53dcd1e02669

Observation 26244bf5-0094-4a2f-a175-3efb65d83ac2 · inbound

How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering? cites this paper.

How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering? InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:05.964749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:05.964749Z digest=sha256:a65556fbbecb0ab62234072fbb9496e3041890b5566eeccae140d67067a639cb

Observation f2293dc2-fba5-4555-8a7b-352c33d71e45 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:33.036457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:33.036457Z digest=sha256:8e7fda9c30c79331f34d19c480295da3a4689f3ddb9bda7efc809b50f831c576

Observation e683b406-530d-4883-8e6f-29049fa30f24 · inbound

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents cites this paper.

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:10:15.130757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T14:10:14.929207Z digest=sha256:eb9cea1162790ac514f26dd2fb0b6a735f996987679f6eab476f54ee5ba81aa3

Observation 3f95cab1-e444-4ba3-b563-c49f4618a8e4 · inbound

Whom to Respond To? A Transformer-Based Model for Multi-Party Social Robot Interaction cites this paper.

Whom to Respond To? A Transformer-Based Model for Multi-Party Social Robot Interaction InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:24:58.979468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:24:58.979468Z digest=sha256:85964703e778ed59f70dd14c6a4817c9e8cb1c136e9ab8f1059ca83b830c76ae

Observation 49664a06-07a5-49c3-a551-aeaea6f87f3f · inbound

HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly cites this paper.

HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:04.588143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:04.588143Z digest=sha256:b0a8800f3d8104afea45639044e77d2cae59588dcb4e7c0e59b5a57c57f2d824

Observation 63b3f5bf-aa89-4242-8139-ecca80d15c9e · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:36.886005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:36.886005Z digest=sha256:25cc51f59761afc3fe5b35c2fe25765e8d9001268ddc9d4d2e1c111b65d32aaa

Observation e4d6e6dc-1985-459c-a7f7-9e3438c90669 · inbound

StreamingVLM: Real-Time Understanding for Infinite Video Streams cites this paper.

StreamingVLM: Real-Time Understanding for Infinite Video Streams InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:51:33.436812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T11:51:33.345812Z digest=sha256:4d8bf8a69f998430af1c9005c8602664996bed26f803e9799ec5182baa222076

Observation b4991dee-1341-458f-860d-3b3bdb084ac4 · inbound

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding cites this paper.

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:43:14.848489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T20:40:41.380829Z digest=sha256:f7e47f80a4dbb55e403698c1c8d2ffa719ca612492a53ea7e32243b54cf78569

Observation 9d8ecc9e-0f2c-4a83-b438-04eb4b7f16cc · inbound

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition cites this paper.

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.219911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:35:48.379198Z digest=sha256:8a62b157265e1c55ff9a91a5f1e16ae55f733a4b3d57ee25332cc6ab5adc78be

Observation c27af26f-fadf-4167-a8b4-78c00b2f34a4 · inbound

AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding cites this paper.

AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:47:53.818862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:43:29.123615Z digest=sha256:255ac54d897bfb6a29f81780d29641a29af6ac783283f13ed887924ffffa3531

Observation 3d550601-08e8-4fef-bd0a-d450e29c01e8 · inbound

When Vision Speaks for Sound cites this paper.

When Vision Speaks for Sound InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:13:46.852107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T22:12:52.160596Z digest=sha256:e922f4467c5908253ffba667462f04e445ff7ebf9b046a56c9ef4767698ba89e

Observation 825cb563-f5c1-4de2-8225-143668a3cd13 · inbound

GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval cites this paper.

GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.375723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T19:12:19.056273Z digest=sha256:755ee13786bb2fa9f3dcf0d6ee713ec4ba63a5fe2c9ee1df066d221aeaffb0a0

Observation d949fff9-4f47-4bd9-9828-14cf0d8d1e52 · inbound

VidMsg: A Benchmark for Implicit Message Inference in Short Videos cites this paper.

VidMsg: A Benchmark for Implicit Message Inference in Short Videos InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:30.063976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:25:06.594946Z digest=sha256:455282ffdeafe58474b79ef5cadf60405f7b773adc9972f375e6b0dc1bf26cbb

Observation d76c4a66-6383-49dc-b854-8df5bd1deb6e · inbound

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention cites this paper.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.379175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:a75168ede8973b1acbb661a7fe9e414d0e6e6b95c89bbf1976e2b9c5d93258bd

Observation f705bd2b-15de-4e9e-a1c3-f31dd6f8147c · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.519902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:28247c690aafa12462feb4e05dc6dbf457fed3bd4b551805c291b6a0886fe450

Observation 32b0cf74-2c11-4ecd-993b-6de5dc48818f · inbound

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration cites this paper.

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:05:50.568899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T04:18:02.341742Z digest=sha256:9f0a0202a0fd518e304fba9a30f68a3b24e4a494d9e48dfc9ef4826da38b4b2a

Observation f73a5272-7603-4603-9b33-02aa4aa36cda · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 141

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:c542c9e6e5c05c3cb35f59ba4e15b4dff443f4740f44726e92cec408d65dfd04

Observation a5be6167-3f15-4bf3-8fd8-5b7bbee536d4 · inbound

Reinforcement Learning: From Algorithms To Foundation Models cites this paper.

Reinforcement Learning: From Algorithms To Foundation Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:04.287534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:04.287534Z digest=sha256:5fd9d4334188b05fa5af66e34a6b0b59533c7fb4928d83c5e516ae63c888b043