Pith. sign in

Paper Citation Record · LEDGER

InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2403.15377.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.15377 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:29.351108Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.518508Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0ea28090-2a9b-430f-a743-afc02aa2e0ea · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.244814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:363b7f33953e27177ded48430d52e372d920248a04180cf2a7efa88c4f3a6f95

Observation 6482212f-2fea-4ee8-af46-15f76d04a960 · inbound

Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation cites this paper.

Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:40:00.050951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T14:39:59.870039Z digest=sha256:08c2040992f8a44414414cb6ff37b197ea25ff1b326bf2a6db7726b600fdf6f5

Observation 67648cd9-4a98-42e9-9df2-3427496945b5 · inbound

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding cites this paper.

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:53:33.707536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T13:53:33.585035Z digest=sha256:f9e388067e63aa22d5a9702f6f8ec4cc703ac93b28168e73039c1370840729dd

Observation be600ca6-544c-4333-806e-450eb4add321 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 257

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.143903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:6d9264869cc7520d0a5b39f06e31fd3463092fcc1c5e632455c5751f317255de

Observation 9b3a2c62-41ad-4638-849f-bff192cadec6 · inbound

DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization cites this paper.

DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:02:41.738013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T06:57:50.897865Z digest=sha256:02678cc4df6d3336e10d48f24e0351fcd1060d06ee87ce2f8ccb38c508472a94

Observation d4ddbc8d-88bb-4485-b83d-33783b4f2547 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.358964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:5ab6599439894a26010c14d731cf0d0c905b09b4fe8b44cab71b80ff675aaeab

Observation e76c4e79-5a2b-4e07-83b2-a35350833ea4 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.554003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:ac154f083199e589fd0180aa053eab1fda90a4e9274a7bd75d2c82c39f086c20

Observation fe182bdf-a1eb-452c-ab36-07d2a8315a47 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.742495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:68826f70b340aea9352d4d74f58320201ba6521613f67da6c56718fcc4bdd935

Observation b85f535a-1e35-4405-ac04-20df1c5043b4 · inbound

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks cites this paper.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.351108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.351108Z digest=sha256:872c9052b31d36d03802e8428566be57e2f8e74dc9a54f3027de042ff5f60e5c

Observation fbc88a1a-335d-41cc-9183-f0aab1822dcd · inbound

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models cites this paper.

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:10.374002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:24:10.374002Z digest=sha256:4d002236dffff103671c41638b403d6b4f334c51884eaf5e7188d8d2a5608aae

Observation 4f82f8a4-75f2-4b0b-bd9f-12cdf4e124f5 · inbound

HCQA-1.5 @ Ego4D EgoSchema Challenge 2025 cites this paper.

HCQA-1.5 @ Ego4D EgoSchema Challenge 2025 InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:47.260706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:47.260706Z digest=sha256:4ab28944f055b866fef4de41cbb46b4be4e63f926a8ad13740d0392572c1b8c6

Observation c51851bd-3b04-4f8f-9f48-14366c20f7cd · inbound

HuMoCon: Concept Discovery for Human Motion Understanding cites this paper.

HuMoCon: Concept Discovery for Human Motion Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.401308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.401308Z digest=sha256:4a287db5af7607e79a3e807e65c079bb85e6452a1975842582c74cf2f8ec71cb

Observation 5c096c6d-ac30-465c-a225-b79619168202 · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.499105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.499105Z digest=sha256:858e9d3cc92b761111fbbf50ced8ff89c2404e91e073fe303a1e41486a7353b7

Observation 34c1518b-6a96-4e58-ad3a-81998148f13c · inbound

VideoMolmo: Spatio-Temporal Grounding Meets Pointing cites this paper.

VideoMolmo: Spatio-Temporal Grounding Meets Pointing InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:30.926631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:30.926631Z digest=sha256:48ff2c361a779e261d9f1f4bb40a0f36f1a4fd67b73cbc4af22a662ce11bb3eb

Observation a206936a-a89b-4446-b0ca-5da22eb7bbec · inbound

An Empirical study on LLM-based Log Retrieval for Software Engineering Metadata Management cites this paper.

An Empirical study on LLM-based Log Retrieval for Software Engineering Metadata Management InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:59.061572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:59.061572Z digest=sha256:14aa15c556426f645a233a9cd60390fe0f36892902c10a0c85f72e514b4f9f74

Observation 30471e61-9c6b-4718-a4bb-197fbf6e0fd0 · inbound

DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification cites this paper.

DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:51:56.851307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:51:56.851307Z digest=sha256:67e30f9c6fd56ebce5b2072fb08d1c4b1e8075db34e8b18b14e4f4e4c05c1d98

Observation 0c6c2b82-1304-4e98-8301-7975bacc3446 · inbound

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization cites this paper.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:00.161920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:00.161920Z digest=sha256:281c054908cf0595cf93ff32839100527bf12c5763f4f696ff8b53dcd1e02669

Observation 26244bf5-0094-4a2f-a175-3efb65d83ac2 · inbound

How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering? cites this paper.

How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering? InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:05.964749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:05.964749Z digest=sha256:a65556fbbecb0ab62234072fbb9496e3041890b5566eeccae140d67067a639cb

Observation f2293dc2-fba5-4555-8a7b-352c33d71e45 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:33.036457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:33.036457Z digest=sha256:8e7fda9c30c79331f34d19c480295da3a4689f3ddb9bda7efc809b50f831c576

Observation e683b406-530d-4883-8e6f-29049fa30f24 · inbound

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents cites this paper.

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:10:15.130757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T14:10:14.929207Z digest=sha256:522871ebd18dced32478c03a3bed44357d423d428101d04245465a8a2beb914b

Observation 3f95cab1-e444-4ba3-b563-c49f4618a8e4 · inbound

Whom to Respond To? A Transformer-Based Model for Multi-Party Social Robot Interaction cites this paper.

Whom to Respond To? A Transformer-Based Model for Multi-Party Social Robot Interaction InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:24:58.979468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:24:58.979468Z digest=sha256:85964703e778ed59f70dd14c6a4817c9e8cb1c136e9ab8f1059ca83b830c76ae

Observation 49664a06-07a5-49c3-a551-aeaea6f87f3f · inbound

HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly cites this paper.

HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:04.588143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:04.588143Z digest=sha256:b0a8800f3d8104afea45639044e77d2cae59588dcb4e7c0e59b5a57c57f2d824

Observation 63b3f5bf-aa89-4242-8139-ecca80d15c9e · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:36.886005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:36.886005Z digest=sha256:25cc51f59761afc3fe5b35c2fe25765e8d9001268ddc9d4d2e1c111b65d32aaa

Observation e4d6e6dc-1985-459c-a7f7-9e3438c90669 · inbound

StreamingVLM: Real-Time Understanding for Infinite Video Streams cites this paper.

StreamingVLM: Real-Time Understanding for Infinite Video Streams InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:51:33.436812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T11:51:33.345812Z digest=sha256:20090007b12902b054ceeaa296873c521ecd4f32655d701e68c1a6b3482b0129

Observation b4991dee-1341-458f-860d-3b3bdb084ac4 · inbound

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding cites this paper.

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:43:14.848489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T20:40:41.380829Z digest=sha256:c5d1a65f4b2f2cdf143493e7ac430839c16c11060e07ae94bcebf3cc0ab6a844

Observation 9d8ecc9e-0f2c-4a83-b438-04eb4b7f16cc · inbound

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition cites this paper.

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.219911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:35:48.379198Z digest=sha256:570c6f35948392800aa7c01abd5cc6af0a0d9381a78bb219a448ee4e9d76a401

Observation c27af26f-fadf-4167-a8b4-78c00b2f34a4 · inbound

AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding cites this paper.

AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:47:53.818862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T19:43:29.123615Z digest=sha256:d78fd4eb56a3b396913a96f4af96731f197e83ea46cfc13688c4ef1b16dc0eae

Observation 3d550601-08e8-4fef-bd0a-d450e29c01e8 · inbound

When Vision Speaks for Sound cites this paper.

When Vision Speaks for Sound InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:13:46.852107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:12:52.160596Z digest=sha256:96d8652fb1310529bd118246ee542312ce0c4910fd6c821820c9b4687e8cfee1

Observation 825cb563-f5c1-4de2-8225-143668a3cd13 · inbound

GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval cites this paper.

GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.375723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T19:12:19.056273Z digest=sha256:6504d47e174e0d3bbc9d4535a56daf19509278869c66ea7935d43e26e323c693

Observation d949fff9-4f47-4bd9-9828-14cf0d8d1e52 · inbound

VidMsg: A Benchmark for Implicit Message Inference in Short Videos cites this paper.

VidMsg: A Benchmark for Implicit Message Inference in Short Videos InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:30.063976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T10:25:06.594946Z digest=sha256:9d7093616e92c5e711d3746178d195a792986844db3089e0b3a014925b3048ac

Observation d76c4a66-6383-49dc-b854-8df5bd1deb6e · inbound

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention cites this paper.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.379175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:9ecb8da08a52ecf1acca2b2294d61b06e7b94093d0ee49ece578394eb4380749

Observation f705bd2b-15de-4e9e-a1c3-f31dd6f8147c · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.519902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:80bbd069f12ba4ed98e64dbdaeece2798eee9adb50c882ec92334d7461f4bbe5

Observation 32b0cf74-2c11-4ecd-993b-6de5dc48818f · inbound

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration cites this paper.

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:05:50.568899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T04:18:02.341742Z digest=sha256:a49765967aa06b980ba14d270ee8e9cd46a27a87bcaf67a56ae9c9b1b106f07c

Observation f73a5272-7603-4603-9b33-02aa4aa36cda · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 141

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:c542c9e6e5c05c3cb35f59ba4e15b4dff443f4740f44726e92cec408d65dfd04

Observation a5be6167-3f15-4bf3-8fd8-5b7bbee536d4 · inbound

Reinforcement Learning: From Algorithms To Foundation Models cites this paper.

Reinforcement Learning: From Algorithms To Foundation Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:04.287534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:04.287534Z digest=sha256:5fd9d4334188b05fa5af66e34a6b0b59533c7fb4928d83c5e516ae63c888b043