Pith. sign in

Paper Citation Record · LEDGER

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation

As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2512.10607.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.10607 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T17:09:23.706315Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aef58132-70af-47e3-981d-a8f1dd0a24c8 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:18.919548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:18.919548Z digest=sha256:379f49b16aed1cfbbfc4d8d857258074c00affb6e18ffe1b6cb2a9fd8a4937bb

Observation 3de8bd12-c4a3-4ae6-81dd-9e3040ac3c31 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:19.085406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:19.085406Z digest=sha256:ac6323fb412c2e0352791f2ecb3d973388336ab3fb48d588bae414a896db12d3

Observation 57223da2-1323-42c2-b54b-7aedaac20fb0 · outbound

This paper cites Object segmentation by long term analysis of point trajectories.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Object segmentation by long term analysis of point trajectories

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:19.314939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:19.314939Z digest=sha256:1d2c48b1e274854ea401635e27ad133c5c9ba6b82d9ca5d098bed81acc734b82

Observation 2ec8f9ad-ab2d-4ca3-88b0-90f5975c6255 · outbound

This paper cites Mevis: A large-scale bench- mark for video segmentation with motion expressions.arXiv preprint, 2024.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Mevis: A large-scale bench- mark for video segmentation with motion expressions.arXiv preprint, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:19.393314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:19.393314Z digest=sha256:e1bc84d2bbcd86a46d2e0fcf6f380bce5191050e2e8f8275ca8b6d35ab8bff41

Observation cb279774-a2a2-4d49-aa4d-0317a17c0c6d · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Mevis: A large-scale benchmark for video segmentation with motion expressions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:19.499474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:19.499474Z digest=sha256:fd7a037e28eda893df5eac129c7e772eb604bff048791a0b8a08932222fbbfea

Observation 6e836d57-eb99-428a-a673-fae38e6452d4 · outbound

This paper cites Tap-vid: A benchmark for tracking any point in a video.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Tap-vid: A benchmark for tracking any point in a video

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:19.601245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:19.601245Z digest=sha256:98d3ee64cc4369acd3579cfd679740719a5d4b7619e1cc7b94ccef524e605848

Observation 2cde1e77-91b2-4942-a4be-392db3c47f80 · outbound

This paper cites Tapir: Tracking any point with per-frame initialization and temporal refinement.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Tapir: Tracking any point with per-frame initialization and temporal refinement

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:19.705194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:19.705194Z digest=sha256:d7245981dc3596e81392641447d3bada585a52ecafc1ffb788fcbadf0f7de01d

Observation 3f6755db-4ee8-4051-a93c-1a896198f5bc · outbound

This paper cites Step- former: Self-supervised step discovery and localization in instructional videos.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Step- former: Self-supervised step discovery and localization in instructional videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:19.771927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:19.771927Z digest=sha256:26b270edb9802c9522b783b677922a84ddb41ea9c60e77fe02f390cc00685520

Observation 02fd8f7d-495d-4c5b-b962-a2541a8c1cd3 · outbound

This paper cites Context-guided spatio-temporal video grounding.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Context-guided spatio-temporal video grounding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:19.858416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:19.858416Z digest=sha256:b69b7e1d05a6303b503cf95b5f27342fd90fe52a9bf3571f93745c57f1bdcf5e

Observation ae883612-6ff8-4ec3-b5a6-fdbacb368aa1 · outbound

This paper cites Harley, Zhaoyuan Fang, and Katerina Fragkiadaki.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Harley, Zhaoyuan Fang, and Katerina Fragkiadaki

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:20.013031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:20.013031Z digest=sha256:6654c9b8c613eed37575c9f8ab847f5b69a83fd1213b40c9ece08f3884708783

Observation 79d85d07-c7bc-4506-8df5-80a4a1ea4892 · outbound

This paper cites Pips++: Improved tracking through occlusions via extended point trajectories.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Pips++: Improved tracking through occlusions via extended point trajectories

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:20.111151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:20.111151Z digest=sha256:e5ca75426639b179c5a9d00fcfd50169cd64927eaa083d309ebc16f790c4f21a

Observation b3f971f9-358c-4fc3-a93d-d3026f16ce28 · outbound

This paper cites A better use of audio-visual cues: Dense video captioning with bi-modal transformer.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation A better use of audio-visual cues: Dense video captioning with bi-modal transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:20.297182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:20.297182Z digest=sha256:f70dde5c345589e51ad631a59d529fcbf9aa87bc666e5492594db04e54155920

Observation 2d83ecfe-48f0-4669-aeb5-e8d9310fb367 · outbound

This paper cites VideoRAG: Retrieval-Augmented Generation over Video Corpus.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:20.433930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:20.433930Z digest=sha256:f525640b98c705a6dddd3a4fdcec89f0ff1c60f854e1bd548f87e36217cab2d9

Observation 63cadd34-5091-46d2-8b09-0636e0f1acb5 · outbound

This paper cites Embracing consistency: A one-stage approach for spatio- temporal video grounding.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Embracing consistency: A one-stage approach for spatio- temporal video grounding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:20.509094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:20.509094Z digest=sha256:9f023a897646c8952577237cd3542803e8bf6125fc24445477ea99c3a5f94e65

Observation 0ca2e823-4436-4236-9394-08565064ccda · outbound

This paper cites Co- tracker: It is better to track together.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Co- tracker: It is better to track together

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:20.649128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:20.649128Z digest=sha256:a9220c7aba0b2d34bbdd1efc35fa0ca3b0abfc9aa98b68dfa41dfb60477c5323

Observation 3e7ef5ab-bb89-4ce2-9340-b5cbe8dfb9ba · outbound

This paper cites Co- tracker3: Simpler and better point tracking by pseudo- labelling real videos.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Co- tracker3: Simpler and better point tracking by pseudo- labelling real videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:20.768356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:20.768356Z digest=sha256:75fdb03759a1d4e68a9c0f7ce391e86a37c5aa56dc7dd35c1f9ec1510a5af540

Observation 2bc5818c-8fd0-417e-a7a2-4a303e1c170d · outbound

This paper cites Cospal: Co-optimizing spatio-temporal context prompting and adapt- ing for weakly supervised video grounding.arXiv preprint,.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Cospal: Co-optimizing spatio-temporal context prompting and adapt- ing for weakly supervised video grounding.arXiv preprint,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:20.857564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:20.857564Z digest=sha256:ac6c0518390405a073481b401e44b655bce37dae73ee20515dd3def447168a03

Observation acf569f3-09ba-4e5c-a3b5-bd2fb4392fd2 · outbound

This paper cites Unsupervised object discovery and track- ing in video collections.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Unsupervised object discovery and track- ing in video collections

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:20.957468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:20.957468Z digest=sha256:ff5019f4ed26d9fcb540accdc710153789144da93109a8f1be49ed6b3b24895c

Observation 6f3e8864-aa41-4109-aa52-be2b78164700 · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neu- rocomputing, 508:293–304, 2022.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neu- rocomputing, 508:293–304, 2022

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:21.061329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:21.061329Z digest=sha256:da5ae8e0ca13d1cfe33fa961f1abee96ce49dda0795fa67a971d0ebbe19bc322

Observation 03c7596a-8e0d-48e1-a2aa-b3c9189148fa · outbound

This paper cites X-clip: End-to-end multi-grained con- trastive learning for video-text retrieval.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation X-clip: End-to-end multi-grained con- trastive learning for video-text retrieval

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:21.162440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:21.162440Z digest=sha256:554c667e68194a3597302e677c3df5dac677f23650279893aa9d5e793f4dc49f

Observation cbcefd51-4e0f-4550-a9b3-952260870a1f · outbound

This paper cites Delta: Dense efficient long-range 3d tracking for any video.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Delta: Dense efficient long-range 3d tracking for any video

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:21.268988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:21.268988Z digest=sha256:7728b1f3062002b2677d79a199bf8d3d715f6d841011fdedacb80c60c4439124

Observation 07ff6088-b3f4-4858-840a-d34b39a37d13 · outbound

This paper cites Unsupervised discovery of actions in in- structional videos.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Unsupervised discovery of actions in in- structional videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:21.333781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:21.333781Z digest=sha256:9be2237a265981d3164b1df0f75fd2779e55d0a6c007e19e655cb22fc893cd9f

Observation cfcbe662-b9b0-4172-8b78-064f432908d2 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Learn- ing transferable visual models from natural language super- vision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:21.415070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:21.415070Z digest=sha256:6afc7b1f377325b31c85323dd2affe992a651d5f28a2a49978de400f6e502ef9

Observation b8830128-56a8-407e-8ec4-7b20d4a3f0b6 · outbound

This paper cites Two-stream con- volutional networks for action recognition in videos.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Two-stream con- volutional networks for action recognition in videos

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:21.447677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:21.447677Z digest=sha256:9cfeaf6061beb84c2ac0b22105af3d6449265d1ac3e34a2f0bd5858fa8230f5b

Observation 3caa1e21-e743-41b5-b7b7-64487c2b8549 · outbound

This paper cites Human-centric spatio-temporal video grounding with visual transformers.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Human-centric spatio-temporal video grounding with visual transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:21.479717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:21.479717Z digest=sha256:ae988ba137ad68b2dca25a4136a96b424175e783b83e0fcf48b59044ca90484e

Observation 72058787-6fce-4a60-be68-312eb18d98d0 · outbound

This paper cites Human-centric spatio-temporal video grounding with visual transformers.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Human-centric spatio-temporal video grounding with visual transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:21.523247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:21.523247Z digest=sha256:198006820f5e25b25d61fafac51a64787e4ed9b8624f2557bfcd889e0df09a83

Observation bb08ca90-92cd-4658-afde-b968f42db816 · outbound

This paper cites Repre- sentation learning with contrastive predictive coding, 2018.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Repre- sentation learning with contrastive predictive coding, 2018

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:21.689307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:21.689307Z digest=sha256:d912256eab93a6b8c059643a6e6793c783383c57b3fa0d0e9ce2dea352050ede

Observation 646b2a06-6c64-4872-a89d-e77122168216 · outbound

This paper cites Action recognition with trajectories.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Action recognition with trajectories

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:21.810364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:21.810364Z digest=sha256:ea7e5d9716e05e62d93edefe9b8816539c3cc2eb837be985276e06f1cb664656

Observation 0984b6a3-f540-4e40-80e9-ffdbb441b3ec · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation ActionCLIP: A New Paradigm for Video Action Recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:21.923280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:21.923280Z digest=sha256:d7b8b0fe9974a03eec8e8a173b169fc3d9ab1eedba7c5f095cefaa6239e1461b

Observation 2d453055-c304-4433-aee9-73bb93a9b9ea · outbound

This paper cites Tracking everything everywhere all at once.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Tracking everything everywhere all at once

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:22.067340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:22.067340Z digest=sha256:58576b5bdc0acf760a87d859bd0c5b51e451754866f2b97b1adcc18f485bbad9

Observation c865365e-747b-4104-9df0-304c5e1e3458 · outbound

This paper cites End-to-end dense video captioning with parallel decoding.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation End-to-end dense video captioning with parallel decoding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:22.234388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:22.234388Z digest=sha256:593d458803631b41cf00b203a31c91484cfb41b315dc49b40f7d20b544822859

Observation 67248719-9360-4fdf-8aea-7097a183c0e3 · outbound

This paper cites Language as queries for referring video object segmen- tation.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Language as queries for referring video object segmen- tation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:22.353539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:22.353539Z digest=sha256:429c2046afb71e77f80eaa4c459708ccc1ec2d3cfc269df97d7f8698babfd1ef

Observation b85600e1-13f3-4eb1-a4bc-5b4987b4f7fb · outbound

This paper cites Spatialtracker: Tracking any 2d pixels in 3d space.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Spatialtracker: Tracking any 2d pixels in 3d space

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:22.453670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:22.453670Z digest=sha256:ed9bd4276f0c2491527ecd82dfd070dc1d507bf62736a7de77e967070dd3e184

Observation 5163c027-4ad1-48c0-b1ad-548dd93d88b1 · outbound

This paper cites SpatialTrackerV2: 3D Point Tracking Made Easy.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation SpatialTrackerV2: 3D Point Tracking Made Easy

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:22.596059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:22.596059Z digest=sha256:ff0cb424b1953aee6d8e5b8660564f4772830a9840f8cc2c519134e0d9d6f0a7

Observation 636d8ca8-ba1e-4cdc-b5af-8e847c87a78a · outbound

This paper cites Videoclip: Contrastive pre-training for zero-shot video-text understanding.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Videoclip: Contrastive pre-training for zero-shot video-text understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:22.725031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:22.725031Z digest=sha256:39fa6a76a73323492ace9f7b5b47db1877aa84aaa7a25cacaedd76ae793c5786

Observation da50673d-f83d-4fe4-851b-4db90ae7451c · outbound

This paper cites Universal instance perception as object discovery and retrieval.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Universal instance perception as object discovery and retrieval

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:22.836577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:22.836577Z digest=sha256:5155a1d23c2901cfe610b74a169a316dea0cf5404bda845c3f95c0c54d2579da

Observation 1b7764ea-dabf-4c23-a7fc-61399e6be65f · outbound

This paper cites Tubedetr: Spatio-temporal video ground- ing with transformers.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Tubedetr: Spatio-temporal video ground- ing with transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:22.954450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:22.954450Z digest=sha256:6496fb1b8ada464cf3fa264cbf6737d9856ee0cba0b6d73bf89a720542b6a1fe

Observation 52dbcf17-2ab9-434c-be6a-afbed23b942a · outbound

This paper cites Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:23.075433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:23.075433Z digest=sha256:b6029b87210dc71fda8193eb43429284da5411dcdee80ae50cd09278fbf58c3f

Observation a810d00c-526e-4f9d-85c1-081f80f0a3a3 · outbound

This paper cites Tapip3d: Tracking any point in persistent 3d geome- try.arXiv preprint arXiv:2504.14717, 2025.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Tapip3d: Tracking any point in persistent 3d geome- try.arXiv preprint arXiv:2504.14717, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:23.165943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:23.165943Z digest=sha256:d9b1e7733f4116a55617635df363fb82c8a5e46b0c34ab6fd9b7f17ddd31073c

Observation 8d26d584-b579-49ae-b32f-8708aba588bc · outbound

This paper cites Where does it exist: Spatio-temporal video grounding for multi-form sentences.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Where does it exist: Spatio-temporal video grounding for multi-form sentences

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:23.290128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:23.290128Z digest=sha256:5e3357ae19f2f4eeb964663b6d9312d7828b49b5b589bfd283b9aeafdc796032

Observation 08531455-bcdd-4cd6-ba43-80920be9c8e3 · outbound

This paper cites Video- text prompting for weakly supervised spatio-temporal video grounding.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Video- text prompting for weakly supervised spatio-temporal video grounding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:23.418978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:23.418978Z digest=sha256:3fdf8d567fd0f02f48a5a4047f7cbe2b8c7751de7bbec863b0f023624d4495bb

Observation a3da8936-48f8-4637-bf56-0c9cca9f15ca · outbound

This paper cites Unsupervised learning from video to detect foreground objects in single images.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Unsupervised learning from video to detect foreground objects in single images

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:23.508626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:23.508626Z digest=sha256:cfdf9797bc6e29d41e0c83f1cb4c46d433201a31491085909946dfb6443c5f00

Observation 8ef7c3d8-a217-4a02-bc1c-cef4b6e0600f · outbound

This paper cites TAPNext: Tracking Any Point (TAP) as Next Token Prediction.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation TAPNext: Tracking Any Point (TAP) as Next Token Prediction

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:23.609008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:23.609008Z digest=sha256:3b0cb339cd921fa8cc7c1e9fdb9cbb67116dccb804f4cabe01e00ed88abedfb1

Observation 0cb55178-5194-4abe-b175-d239f12c1ff6 · outbound

This paper cites Dense video object captioning from disjoint super- vision.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Dense video object captioning from disjoint super- vision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:23.706315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:23.706315Z digest=sha256:c5de97981247c1f84bad61c135611567f99185f29ecc90ebfddafe35c345bb22

Pith citing papers

No inbound Pith citation observations are available.