Pith. sign in

Paper Citation Record · LEDGER

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning

As of 11 August 2026, this Paper Citation Record lists 100 of 300 outbound references and 1 inbound Pith citation observation for arXiv:2606.12195.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.12195 v1

Coverage vector

measured 100 of 300 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T09:48:27.652901Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:04:17.133693Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 300 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved59
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch34

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d15cbd7b-c4eb-4929-acc8-f3e374f16fb9 · outbound

This paper cites V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:02.917527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:176191c9332b0a0b79eda63dc974f273c11d809290b6d4380b8d294047198167

Observation d2c6d8f3-4748-4950-940e-60d27a1e7afd · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:02.914700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:7263a57b63298b9783ff6021d8d1431c401646c305e002b2721601e6a014cca4

Observation f49e3c2d-f3b2-41d6-85e6-2cc0fdb44ec5 · outbound

This paper cites Revisiting Feature Prediction for Learning Visual Representations from Video.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Revisiting Feature Prediction for Learning Visual Representations from Video

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:02.916908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:abfa3d83b3a82c782d37195b8d8903fcc62cacf06d028842b960b26ab8fc413c

Observation e8d72ffd-a59a-4c99-a667-148e9f066f75 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:1de8e90854a773b6c115101b023da94236190814fcd2c8de8ac511ef9b0a368b

Observation 2efbf8a5-55c5-446a-8ffb-303883f2ebf4 · outbound

This paper cites World Models.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning World Models

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:02.911671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:37ca4b1dde8d74a44f44c39fdf8ba2ccf5c4340772fad287b5f335b5d3b913f2

Observation 3520b587-a9e9-4f67-800a-215826e6ea14 · outbound

This paper cites Kwai Keye-VL 1.5 Technical Report.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Kwai Keye-VL 1.5 Technical Report

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:02.945834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:683dd5067a7e5317ac1e80458b20b5797e86c40056f4d9861503e019d954f55c

Observation 4e02f87c-5822-439f-a655-8b0104798374 · outbound

This paper cites an unresolved cited work.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Unresolved cited work

Reference 7

Resolution
parse uncertain
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:c8eff3230563d82d70be7dcf85c7ba2ebd8c1d04153accaf63045c20d510e3f6

Observation a1bc11ee-87d1-4e01-8a90-1ca3840e8f7b · outbound

This paper cites an unresolved cited work.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:0074f2ad62c15ca375ccadf4f05f0c751bc0b886c32be26da7b8dc082159f6b4

Observation 04d415ec-2abf-47fd-b308-b3a8ed199088 · outbound

This paper cites OpenAI GPT-5 System Card.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning OpenAI GPT-5 System Card

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.187254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:edc927e400f1ba82c19101323a68946476f6262a55355e1244e67bb7e1497f2e

Observation 6386566b-b47c-4ad6-b4f5-088ff4416594 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.189549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:395ee0c31e78f0783153a38ba15b37844e01fdbc7af266316be4da420f636707

Observation 1d188e8e-622f-4eaf-9e52-54c304c7be7d · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:fd8e898a7af44a2543b1f35bb905551de74b91ba22b5ee6903ee47a3a0988195

Observation 763142f9-c533-48d9-8897-df2db1b26bee · outbound

This paper cites Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.177534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:51d32a986d55a2688126ef9df9bab876e75748ee02d458a38ef5221a5f542042

Observation 5aaf3362-7b6d-479e-b45c-56a8414758a6 · outbound

This paper cites Advances in neural information processing systems , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Advances in neural information processing systems , volume=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:fafeb5f542e89aa24cbeaef763d0c28975dedce00a483342d1429354fe57eda5

Observation b4b281c9-5d41-46e6-8fbd-94bf7ba8e314 · outbound

This paper cites Advances in neural information processing systems , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Advances in neural information processing systems , volume=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:2e0bb3be0a6db98b93bb775565550af9662347d09fcb56a8d732a2e712a75835

Observation 0b5494e0-da3d-49d0-a7a3-7ace9e33d6ff · outbound

This paper cites The eleventh international conference on learning representations , year=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning The eleventh international conference on learning representations , year=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:068c4b56eeb3902d5af7f7c7b9b7d3e08976fb8d1ae1ecd60bd161573221e73a

Observation 3fb389e0-15c5-415f-acea-10b39932c277 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:c36a9361422051a4649a19a7620c36f205829dd5bf7cc0de964ac2c92247de13

Observation 4766ea8c-42cd-4fb6-b276-b26db55b585d · outbound

This paper cites European Conference on Computer Vision , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning European Conference on Computer Vision , pages=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:e0263151896839d23199e079d7d5ff54a91700836caff3061286ef2d6adee121

Observation 758d6d7b-162d-4725-896d-1b330ecb53be · outbound

This paper cites European Conference on Computer Vision , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning European Conference on Computer Vision , pages=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:cf26dd78260c3f51c38eb1f462b2a16abd34627bce361e61be505acc0678864c

Observation 3079e832-34b4-4472-b8aa-f3743e308d3c · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.180078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:db37e662cc31474c6a5dfd282bcc7ca7807ab0c7232752c06a133a264ea2f5e5

Observation caea3a30-b8d1-4ff7-858f-ba8a3a6f3ef2 · outbound

This paper cites Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems , pages=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:1408513dad997a86aa2cc1ca9b708ae9a5dc5144d284ea44a59264de2b43cfde

Observation fb811108-c52b-4430-ab08-62953bd0d25a · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:560b119f0e7cc01a8f90193244d9ffa1849d1316530e3049c9c8e79ed99848ae

Observation 0a049c9d-5eb0-47ba-9ba8-955ebf952fb2 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:e414a38f81ddf923c434b685fc05a5b9ee6729ec4ad4a93ceea88263c4702b12

Observation 39bd2809-6fa6-447c-a519-6f6e7fb985f1 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.172762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:8a4d5bc1c5534bb4fe729f9794d02c4a066c52962f2335355efff9383bf2c54f

Observation fbd84177-dcc5-43bb-8649-df61f4558246 · outbound

This paper cites arxiv , year=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning arxiv , year=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:e725775df4102686f32fd18cab2a599f1062e891cd0adbebabb3c03b2c5b3876

Observation 6f94c4bd-f0f9-4aad-99bc-854e50711c10 · outbound

This paper cites FineVision: Open Data Is All You Need.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning FineVision: Open Data Is All You Need

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.179788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:74f6d3454c5a6a365af6a60048a4d747e9216ee51e06d1d69e9eca17cc9eda86

Observation 067fdded-660c-4c46-95ad-f2f2b80fb3ae · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:48:03.182059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:d324eb245359d960589a082dc54494001d4a5c57412308efae1d771966022732

Observation d7fdceb9-7612-44e9-a1e4-bb1f62bcff62 · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.182357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:8e6036a8c9e47b61c146dec2045c1e819d1b95caf8a9163e53fcbff1cb7c4f38

Observation 3b93fb68-def3-433a-81bc-c2ecb7273df7 · outbound

This paper cites MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.184975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:15790527fb758a913e4d8a7226da1625f5a8599d6f6225a67ce3c7591bdc6ca1

Observation 5d1d5be9-58c9-4e64-852d-8549a0d27deb · outbound

This paper cites Qwen3.5: Accelerating Productivity with Native Multimodal Agents , url =.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Qwen3.5: Accelerating Productivity with Native Multimodal Agents , url =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:3e062df889b728cff6fdadca9d2a5816f24a9127a302dd3d76b3b009e1d3e1ce

Observation dc3672fd-bfff-41b5-a7b6-7a949378ff18 · outbound

This paper cites an unresolved cited work.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:5934492e5e16e0511c998bf15cba9e0c46aa84aef9c07d74a8b179fb55d9dde0

Observation c1a9e372-b008-4451-819c-f737c0d9c06d · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning GLM-5: from Vibe Coding to Agentic Engineering

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.153084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:9b052e1608014604d9d60c3d3cf1a4e49584118b954c28acc46beef5b42db7c0

Observation 35862d1e-06e6-47f7-8d1c-63d4307702c9 · outbound

This paper cites Huang, A.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Huang, A

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.156179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:22f84108c3f26ef7942ce892af7f4c0fee1981d6d4baf0c835ade3480951f383

Observation 30e83669-fd66-44ad-8758-9ce7dca9ae74 · outbound

This paper cites an unresolved cited work.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:7335491aff02f299179be66da0ad4caf2b33a930eb937ef3de5df9fa8e418287

Observation 095b7fd2-d509-4917-8359-0b9271782241 · outbound

This paper cites 2026 , eprint=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning 2026 , eprint=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:44362f3eb41cc6d2a5b7a477884e445f9b1162571445c2d4aaf7e5fc51da1184

Observation 6e45bfa6-8a70-4f3a-941a-baab0324fb70 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.147721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:eae6b2646076eb46d90a2e1d89a0fe59f0494b3f8cc13a9470b97f0ef93751d6

Observation b23a701e-4e54-4303-b3bc-3886b49a8bc0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.153299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:0ab0e5182d73bafa338ef1719439eaf9848fa115186beb4493c06ab360928be4

Observation 3e276bbf-6ebc-4ead-83e0-f496a9d6499a · outbound

This paper cites Kimi Linear: An Expressive, Efficient Attention Architecture.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Kimi Linear: An Expressive, Efficient Attention Architecture

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.158827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:78177d81edb3d740c3edb18684df1ad6da9cc4a9e7deedeb80441ce8a629d361

Observation d6c40ece-16c7-404e-881e-385796a8be7b · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:db51c01370bd8cb7ff1e91f888285e2fc7f67f16af91a2d6a2c20e5caa087db3

Observation b884b8ad-6975-4862-a1ff-cb73bf9daf4a · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Kimi K2.5: Visual Agentic Intelligence

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:48:03.135272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:5f00cd7a272e2a761f029cbcb6761730e88347abfaa754be6ea54a2ec74661a6

Observation fcf9511a-254f-449a-92d0-9631d348c4be · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Advances in Neural Information Processing Systems , volume=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:af92b85bf377ed6a0dcd93705716cf38eb2f2217228e603ff5c417f6d493e848

Observation 1617a9e2-5470-49b2-8af5-2a165cde6ae6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Advances in Neural Information Processing Systems , volume=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:c8f0eee8cb704ce4494789df004724bb41c669c3eb7a5a23b2392c02fc38871a

Observation e120af6a-65ae-489f-b59e-a8714757e604 · outbound

This paper cites Proceedings of the IEEE/CVF international conference on computer vision , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the IEEE/CVF international conference on computer vision , pages=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:d069137fa735d2b751207b133062d92acc007f04a01be6991cb1c74d9e9c39f7

Observation 9ca5ab0a-376c-43b9-9541-07646a7c21a7 · outbound

This paper cites 2025 , eprint=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning 2025 , eprint=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:1754b20d2a1eab3127dc368f3e7dd0e66284b008a9a74be4f04420fc9e31024c

Observation 851f6889-d71d-422a-af83-29af343c2709 · outbound

This paper cites Step3-vl-10b technical report.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Step3-vl-10b technical report

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.138148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:7fddeba12b5dd863bc2f45615a4669f263509cc4e4de22af808d576e10f7b29f

Observation deebf6cf-ae76-464a-b721-a0c36a43571a · outbound

This paper cites Qwen2.5-VL Technical Report.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Qwen2.5-VL Technical Report

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:48:03.142809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:064b5a098c2300f84d83fbeb8d6840400c3d9d45c68746afa763e8fe659a7308

Observation 64c512f6-3988-429e-a4cf-077e1340ae1e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.161113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:0ea820c72d8f6dc1c0b87541467b6aef63392296e3bafac57bb578f6fae9514f

Observation 54c2117f-fa73-41f8-a638-c48fdf0b39e8 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:48:03.191919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:0b84880fc3cc281f99ad8b67b3589e78eede9592a5dac010d77cbe48540d8e0a

Observation 36156d0c-4d33-48af-87f9-17346d110183 · outbound

This paper cites International conference on machine learning , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning International conference on machine learning , pages=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:924f002569b9ca5db348c512fce3beb86df437cf7c82aee9e534b5ae8a305dcb

Observation 7a9c2c7b-3ea1-424f-be20-76922501b865 · outbound

This paper cites Advances in neural information processing systems , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Advances in neural information processing systems , volume=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:951b3c334c2e50e70409f94d95d0faa9b3fbd23ded64e80bf3c4de6852e485bf

Observation 7b9dc48b-114c-48e4-9079-8d726a9eec72 · outbound

This paper cites Dsi-bench: A benchmark for dynamic spatial intelligence.arXiv preprint arXiv:2510.18873.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Dsi-bench: A benchmark for dynamic spatial intelligence.arXiv preprint arXiv:2510.18873

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:03.205242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:92dd0f73fdccaa46e677a5219989f8ca126d2188756d57ae44a888de10e39836

Observation bffcd23b-6c9e-4041-b333-1e5c27b8c6f9 · outbound

This paper cites MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.160875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:5e6329134ccc8482726453bda89b7e371a9a8efd362429ed4047731fb58bb968

Observation a343a160-20c0-4ee9-a5ac-5d074027ab7a · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:7b30bf133ba42884d2121dd1b9f974adc8716ba94a7774627bb008d6de86e17f

Observation 901738ca-96b5-4233-b361-92f8af6c1835 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Advances in Neural Information Processing Systems , volume=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:03aa23fec8c30458ea006a109ccbad88ad510f79e464b021afa06d08670fa65d

Observation 5a404da5-84a8-43b0-9bd6-65e01f1f573e · outbound

This paper cites International Conference on Learning Representations , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning International Conference on Learning Representations , volume=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:f01797aaf6223b25e59948ca22d6e95d5d7244e48f25965863ff58cf270c5f31

Observation ccd22cc7-6830-4dd6-904a-38d1716991d5 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Advances in Neural Information Processing Systems , volume=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:37a7acae19f706b60b4cd83328dcbbafb32308fec0be3b6b709449a175718ab1

Observation ead62b95-4bfa-439e-9736-5edcc714551e · outbound

This paper cites International Conference on Learning Representations , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning International Conference on Learning Representations , volume=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:e80f74ca23caa2a86f417258c7c8ba647adcee7df1ca7fefa6841cca5422be72

Observation ebd14c95-47e6-4e3e-8484-3ac90d65ce0f · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:8abd1fd806933c382556035338906fb931df4f8642527e90e7f74bd3d2a94302

Observation c748f440-9e64-4282-bcca-6afff8f876e1 · outbound

This paper cites International Conference on Learning Representations , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning International Conference on Learning Representations , volume=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:b06c18b5b38ad1c524b930b472c7ac104ab99804cc18ca31ef16d347d863a0e8

Observation 51c1e0f5-38e3-4ad5-a0f4-6801fc3a1689 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 59

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.184805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:f7a0c75653d717881444dec95729850043560246c3cafa49d8108130f712d325

Observation ce157bbb-ce03-4484-b12e-e3315a6afc37 · outbound

This paper cites Hugging Face repository , howpublished =.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Hugging Face repository , howpublished =

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:76c126fa24692af865e3192f06bb000a6fd3933c7e50493fd9a1ebf3daa77682

Observation 40458348-34ce-4498-8d7f-9e5d31b50f4b · outbound

This paper cites Mmsi-video-bench: A holistic benchmark for video-based spatial intelligence.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Mmsi-video-bench: A holistic benchmark for video-based spatial intelligence

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.214993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:b5ee9e335d7221c5021177c3b6e0e3fab178db21bc6f579fd217a4b28af97107

Observation 9008e37c-b381-4886-91a4-39d4d05a29f9 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:2e14c2243bcb3598a57d56709792d284f77193bd851ee5514b7c9f22b3cfc947

Observation debe2a1b-f960-4630-a621-115eb0da35fe · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 63

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.194402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:8109ad050a2c2b8434ec41d0513a5d08bc67f16b7f7f51c3f2f6034eb9ba9831

Observation 553d6468-3e84-4ac4-8d72-63e0f1f53020 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.196700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:4f61ede759dd58155a5b4b631131af3e5f7123bd8115fde3235c0aef8d85497d

Observation b2186011-4fae-425a-9efa-910710d4bf73 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.106833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:34e9e774c526548833875e7c2feb891aedf9ff8f3bbd0ea7c9a3f7725b9e7eeb

Observation 43aa0c1c-645d-44ec-902f-d248e76a9180 · outbound

This paper cites ICML , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning ICML , pages=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:59225f31a034a3a3cf84d1bf0f273a6fbfb754e1e7ce0e34ad3ee415acdd1a00

Observation 3abe2257-2619-464b-a16f-0d7353333928 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Advances in Neural Information Processing Systems , volume=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:c75a74513930aaa504996e0423498ee416abb7559445b066e8aa2173cbd02b1a

Observation 2c338775-2a18-4995-bcee-08a20d98e88a · outbound

This paper cites Neurocomputing , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Neurocomputing , volume=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:c5b81d5eab017bae4b0e58b9ac09366f2a80b0e7a9e9cfa3f396f0fa2704d15e

Observation 1412aba6-5475-408f-9e70-158db8acbd48 · outbound

This paper cites CoRR , year=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning CoRR , year=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:f68fd37c853f95bb69b7e9e17fe7df38ea959ea4f35a58abfb4c30d375b07526

Observation 00ed264e-e0b8-4b82-97ad-493476fe6b1e · outbound

This paper cites CVPR , year=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning CVPR , year=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:76fb25ebaded7cf92b0eb1c9666ddaa309d25814de7934f7326211ef946be00d

Observation bef0ecdd-1dd3-452a-bf45-091c0466b3b5 · outbound

This paper cites Advances in neural information processing systems , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Advances in neural information processing systems , volume=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:570fd8668b4091cbaa2b2009ebe9c6f6d01d8087a74c804e53c3f2403cae8508

Observation 4d65bcc4-ab84-4259-90ef-d885c961300e · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2020 , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Findings of the Association for Computational Linguistics: EMNLP 2020 , pages=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:a39c13d2384245db30da81bf9c0d1183fbd10f41496d7ab1ba0ff55a8cc639d3

Observation 84b2f344-73e1-4fe2-943d-63aefc74c4b0 · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:ed3776c00e5044115b07f7a050fce22691818af4d92c2b70685069017ac799d1

Observation 6377a3c0-4ddb-4c39-a156-f8a01125d46c · outbound

This paper cites Proceedings of the IEEE conference on computer vision and pattern recognition , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:f935550c0459723187601579f8e058598ba69dde7ea95c650ceef43a5a6fc75c

Observation a9e4387f-ba8b-4e27-9052-eb84d5f43e5c · outbound

This paper cites Temporal Difference Learning for Model Predictive Control.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Temporal Difference Learning for Model Predictive Control

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.124660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:a5eab6a7255ef4f9dc3b585065baa62101ee12ca9d488c0c5baa02cae23643c3

Observation aca501ea-fc5e-48a3-9942-7a3e92d290a5 · outbound

This paper cites International Conference on Learning Representations , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning International Conference on Learning Representations , volume=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:053bbed88ef767b5d1de58f03e8ba6595d5488a3997e7e90c56485e3758bc63e

Observation 6ca6f77b-f490-4541-92e6-5da835a0e0d9 · outbound

This paper cites Journal of machine learning research , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Journal of machine learning research , volume=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:95697ab974b4cadae137ed690443867d8c10866a04aa7653be3cb9cdb351720c

Observation 21853ee7-3b40-4d7b-8b0d-99ad7ddf5964 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Advances in Neural Information Processing Systems , volume=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:45248dad9f12483b7bd8f921fa51a39f73eb96f438abd5b61fc8fae90250da16

Observation 96660d4c-a67d-4273-8969-4162b0a93d01 · outbound

This paper cites International conference on learning representations , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning International conference on learning representations , volume=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:aa99b2448f63e3b5f64fcb72d100e01dfb9c6783eb46cda9eeb4d4b3947b6ee9

Observation ab0c2821-f6ad-4c31-9afd-b2b64f33209d · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 80

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.111707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:94fea6d9807fb03eeef86d5141cc80c9796e6a7eea7fb7da5731237f88680fe0

Observation 90208d52-4020-4672-974c-a79e7f817295 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 81

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.117007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:0ddaa81485bc080e822419d6c53cb13533087aca02f9659a2dcceb223f4cf292

Observation afad968f-9070-406d-b57a-5006cb51a389 · outbound

This paper cites International Conference on Learning Representations , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning International Conference on Learning Representations , volume=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:a1b150039b93bc497963f5165cdc21b21741df79aa769e6be61a09cfc76c36fd

Observation 176ea7c3-d959-4670-8140-e271abc99c6a · outbound

This paper cites Thinking Machines Lab: Connectionism , year =.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Thinking Machines Lab: Connectionism , year =

Reference 83

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:067f2089db7c69f2ace86af9cbd8d5b9a942c6233c5b52e899a98204d9eb7966

Observation 5bc1d740-5c45-46c2-a4f9-fcc350a7a827 · outbound

This paper cites International Conference on Learning Representations , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning International Conference on Learning Representations , volume=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:dbc9b28b53dc942f5ce4b9e3f0f342235b227cfa0093d8c23f148a0ee88b1d8a

Observation d5615fd5-1c14-400d-b15d-07952292b2cf · outbound

This paper cites Advances in neural information processing systems , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Advances in neural information processing systems , volume=

Reference 85

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:e0aa4b5f7f69d64a1240202692bed94702800d03b850e44373f664eae92d342a

Observation e3ee22a3-9185-43b5-966e-bda73f98f7e3 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:88b145c998fa87498b7b2e22b5c15ba93c7a994ab4472da1be099b953900c337

Observation ca224461-22b6-4536-8b01-e8b77178c645 · outbound

This paper cites an unresolved cited work.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Unresolved cited work

Reference 87

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:e78aefd0a875a50903af59eecee8012e818bfffc4b1e39a537b7cc23df4073dd

Observation 800280d7-f652-4a8a-9a3c-b592cf6ed20c · outbound

This paper cites Proceedings of the 36th annual acm symposium on user interface software and technology , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the 36th annual acm symposium on user interface software and technology , pages=

Reference 88

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:019f9f1cb4f826e36ebcfe0b7632d2d841a9b027c4b7712431bf61d924a90014

Observation ed12176d-b8e2-4609-8bdf-c4f543a6209d · outbound

This paper cites Advances in neural information processing systems , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Advances in neural information processing systems , volume=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:2ceda82ed35e58c97c02c8ffe8a8ced7c11e2a46274e285971c08965ca376c68

Observation 421d85fa-12d5-47a3-9ce9-32708813056d · outbound

This paper cites Qwen3 Technical Report.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Qwen3 Technical Report

Reference 90

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.089719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:f7487728932582d2135b4a855d401f1362628c39a33d7a2dc2be560be9608e87

Observation d5d9fe82-50d6-4e2c-9c75-dc811ebdbf4d · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 91

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:02.966022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:081ce2150a035407d9551f31cb3285cdec3a4b85c5417dba1803a06186cfc5f4

Observation e94d2c01-5a30-492b-90be-87e876950a5c · outbound

This paper cites 2024 , publisher=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning 2024 , publisher=

Reference 92

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:0d76c8507186df9ddf3600618b84c1e00b8acf5d8456ca6b5dba6cc980a00a92

Observation 85f5df00-8a9c-4013-951d-2f2df62f2a0e · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Dream to Control: Learning Behaviors by Latent Imagination

Reference 93

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:02.981512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:c978e34d9bc08cb6f36d73bcdebe5d9a6e5516ab911eba6f75d7c6ce784059b7

Observation f87dbc2e-fb79-42a8-ab84-f078ea458654 · outbound

This paper cites Mastering Diverse Domains through World Models.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Mastering Diverse Domains through World Models

Reference 94

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.095524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:40467724ebaaab16a08d42debcc8b214712ff423954164c8f3611b5c2f43d81e

Observation e4ff952d-310c-4ede-bd07-77229975e78c · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Advances in Neural Information Processing Systems , volume=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:6675bf28b6f3bd0de76c6f5789ec567408f16c4f57a15e389c5e355b53d3bb27

Observation 164d0dc0-e567-42a8-8963-4a8e141c1486 · outbound

This paper cites Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:03.109381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:0c1b0faf2789994228a31d2dbc2aa2f93893df2ea93c106f8621c03d84d85afa

Observation 897c40b9-1f1a-40d6-8a8f-cf06f9ff5cfe · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:f448a27955e7c185cdf62fad2237df0b529358677b64da9129d693a29fe55454

Observation 1af5d675-7834-48da-a95a-42ee83a1a70c · outbound

This paper cites LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory

Reference 98

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.122373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:d056e48788cb78aa2194a9a840e6c912e4eccd210f2967a3d8a1f43f645f1b75

Observation b951adc4-3f46-4663-bb76-a4a131a3b5c8 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 99

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:bddfd46cba25089f2ee0ac82fa50f6f7b5d365a353702f179ba6f2f7f12beb9c

Observation d9cf8775-a783-4640-839a-95d477bfba82 · outbound

This paper cites International Conference on Learning Representations , volume=.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning International Conference on Learning Representations , volume=

Reference 100

Resolution
unresolved
no resolver link, observed 2026-06-27T09:48:27.652901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:231e2a3a5d1607a5fbd00b0e651e64831a66cb8b8c1d94890dac27dbea7a811c

Pith citing papers

Observation 02c8c4c3-0f6e-4e29-b8cc-cc9cf28e9563 · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:17.133693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:17.133693Z digest=sha256:0ae5094aff77044ae749c11592266b4a1e08393e30f923394e8eab477b60df4b