Pith. sign in

Paper Citation Record · LEDGER

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding

As of 22 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2507.18552.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18552 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:36:17.742562Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:26:49.679168Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T13:26:50.220608Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dcb6ff57-d4c4-4983-902c-364342c2f332 · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:17.669298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:17.669298Z digest=sha256:156d0a8ef7e16fd87318e459a22d4b5721add7594e263bac3b501819fa72705d

Observation c3035e13-6f62-40a2-9792-0eb942947abb · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:17.940088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T14:36:17.672416Z digest=sha256:017bd6f1440cee229af1ce7d554ff81fbe279c952ce69f91b367b13aa078df07

Observation 207c1d92-d5c0-4ce3-99d3-cb6b3e437068 · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:17.931333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T14:36:17.675479Z digest=sha256:3d8f7d3284bd0c8d70a327a53a41436c88a0c71ad63c89898f48722fd4cd6cc0

Observation 1c1d00f3-a256-47e3-8b6d-0f04d9b0a1db · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:17.922870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T14:36:17.678374Z digest=sha256:d094c0c59a5425bad72216f31a990ed393fb17dad2e553916548504c2f3d4679

Observation 3b5435fb-640a-465e-92ea-ecc13e6519cc · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:17.915323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T14:36:17.681472Z digest=sha256:3dbfe2360e5807c09e9a0168761cb3d51662b2e280588ed17ef030144617656c

Observation 140b7de1-be6c-4bf6-bac1-ff3acd5307e3 · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:17.684196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:17.684196Z digest=sha256:8e9923e66366c53a4f198c068a4532c34621602835ec9909deaaed178318abfe

Observation ccec6382-bb75-44ad-87ad-bca9a277fb83 · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:17.896004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T14:36:17.689452Z digest=sha256:5a546848908d5c4962d832f88657a39162b2b70b72769b2e445e131a823beeb5

Observation fc12e226-4896-44fc-832f-dc33d30f9289 · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:17.888389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T14:36:17.692137Z digest=sha256:730ed30a856dfaf0dfbef3fe5b29dd2859ce12bd79d9166859964dc199e23bba

Observation 82d745f7-4377-4fce-8909-bcd39332254b · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:17.698429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:17.698429Z digest=sha256:85f75d60090dc19d6881869352f6c65e01d73af505289037ee794f02217e5cca

Observation 51b9c1ad-c75e-4ace-8554-3329c6a693c8 · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:17.867816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T14:36:17.700991Z digest=sha256:78db04d43e4790e977a67c0cce074064d85d013a481554f3f93574496ec4866c

Observation 52fa79b7-9c94-4e00-9ebf-fc75278eeb70 · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:17.860421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T14:36:17.703583Z digest=sha256:b1afb4ce8dacbd83cba11de3b6226c1b2f0ceb63fc9330466ecb28dc3fd29781

Observation 8c2a5f34-8d74-4e9e-aee1-f01d811fa393 · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:17.706991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:17.706991Z digest=sha256:bf193eb1d3c2ea176080a6df54413a7e0ce982a5dff199d2675c909a7f37a844

Observation a595082c-c994-44b3-bec5-4db5b5050eca · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:17.852279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T14:36:17.711139Z digest=sha256:e7c918a726cba7a7536815a03a668334d95d58b21c70dd5166692089ed7168bd

Observation 1d3e9038-0a5b-4ea3-b712-bd4efdeffec1 · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:17.845087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T14:36:17.713694Z digest=sha256:d649cf64d477224e70c2e477e7c1be65c6c92a6c4388a73f21a07e4987ea10f5

Observation 4eb03fb5-9f1f-44ee-a19e-f3889548655b · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:17.716107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:17.716107Z digest=sha256:04ac91cedd41c7117f8c294ad51303001a27f1cf293065a7d90491983edd41f0

Observation 48d14979-6788-41b4-a8dd-8520ce911b75 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:17.718956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:17.718956Z digest=sha256:b498e872424424de91046eb99a4148ef4966af7588f53399392f058358666dd7

Observation c72c969d-f1a1-4537-938b-8c0820829b92 · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:17.833880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T14:36:17.721605Z digest=sha256:426809e1d7562381744dbe2cf74faa15c61e3a774301a9a78a61bc66e023b2d6

Observation 3d0b065b-b671-4a08-90de-556ab4257c7a · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:17.724007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:17.724007Z digest=sha256:1184bea92c99934a5471dd4b6eee517e7b6e4d8c2d39f0dc80b9e042332af59c

Observation 2228c1a3-fef8-4a05-9d27-bde4302af7f0 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:17.726835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:17.726835Z digest=sha256:e3c1a020f22f4574d8c474a91a51aca515e84884af4a1a9ea6385d2bcbd0ebbf

Observation d8e201cb-3c24-403c-9741-f9b85f9d0127 · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:17.826001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T14:36:17.729874Z digest=sha256:b0156c4c946583510ab5921baa8cbdbaefd294217b8ef5499b15cbbce3298dec

Observation defe94cb-2268-4054-a089-90505a5b28ea · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:17.818364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T14:36:17.732305Z digest=sha256:07ef264cdb9421b28031d040140978e803b96e6f6b1754adfd1f0aa8d4f987ba

Observation 1093e8d9-469b-4368-931b-5200d6fd8e13 · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:17.735208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:17.735208Z digest=sha256:2d5d1703b35519c5fbb0f918c1a7970bf8a935d341dc1991e784234fab254b7b

Observation 15f36700-c942-4730-940c-ab614da42fed · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:17.806252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T14:36:17.737673Z digest=sha256:32c7f62b54000a017bac1b663c1d59caf9d989f7ad01234a279ca40136fec6f9

Observation 8adcdb16-d28a-47a2-8e91-34efca6867d6 · outbound

This paper cites an unresolved cited work.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:17.740027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:17.740027Z digest=sha256:dbec2f7da86f10492b36f3ad8826bfc874e67365f5911b064c2e0b34b3b7a828

Observation d994a22b-3a6b-4c21-b371-85d8f4a306b1 · outbound

This paper cites GME: Improving Universal Multimodal Retrieval by Multimodal LLMs.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:17.742562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:17.742562Z digest=sha256:03ef9ae82566eb99d1a08ede61e35e0a5c35be19f806ef3f057343e475228b8b

Observation 5253b98f-802d-41e5-869e-c8d0dc454592 · outbound

This paper cites InIEEE International Conference on Computer Vision.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding InIEEE International Conference on Computer Vision

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:36:17.904087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T14:36:17.686968Z digest=sha256:2411aa7356205e289a4bfd2ccf80da30d6ab4c96f5154f86bfb4a7f7e0ff3277

Observation feba9fdd-08f2-42e7-ab01-93652b152ea4 · outbound

This paper cites In IEEE/CVF International Conference on Computer Vision.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding In IEEE/CVF International Conference on Computer Vision

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:36:17.880291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T14:36:17.695722Z digest=sha256:d486d9280be5d65ef4a4fc12752d14538b2ae62543c54645ba3099c19f39a635

Pith citing papers

Observation 6eda65ef-6e64-430e-9cbf-91f7bee17e5c · inbound

Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos cites this paper.

Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T13:26:50.225809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T13:26:49.679168Z digest=sha256:f90044256237c9c5d0de69afd141c2c564503479c3014c0403f3f7eec79f0aaf