Pith. sign in

Paper Citation Record · LEDGER

PolySmart @ TRECVid 2024 Video Captioning (VTT)

As of 20 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2412.15509.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15509 v3

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:24:00.624313Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0952f686-a54b-4aee-be77-f2458595612b · outbound

This paper cites Trecvid 2023 - a series of evaluation tracks in video understanding,.

PolySmart @ TRECVid 2024 Video Captioning (VTT) Trecvid 2023 - a series of evaluation tracks in video understanding,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:00.992472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:00.512091Z digest=sha256:d3789bd918b00536e96477f01507c7de5c5eaba3755d661784d5e2a55ce54b01

Observation a57aa07f-a414-4d17-a4e1-a09ee8d94ec3 · outbound

This paper cites Motion driven approaches to shot boundary detection, low-level feature extraction and bbc rushes characterization at TRECVid 2005,.

PolySmart @ TRECVid 2024 Video Captioning (VTT) Motion driven approaches to shot boundary detection, low-level feature extraction and bbc rushes characterization at TRECVid 2005,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:00.969822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:00.521107Z digest=sha256:17f632865e6e4add36725866827423ca86c9d255a9019f5781091ae8b835bef9

Observation d3b1223f-7b9a-4829-9107-087534c65df7 · outbound

This paper cites Beyond semantic search: What you observe may not be what you think,.

PolySmart @ TRECVid 2024 Video Captioning (VTT) Beyond semantic search: What you observe may not be what you think,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:00.953029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:00.528283Z digest=sha256:9acf5cdbcd2811246c0bb3a75357e09c795b0791bfe794bb032ea2590fa82973

Observation e841fe0f-4cc3-4730-ae12-aa09228f5c56 · outbound

This paper cites VIREO at TRECvID 2010: Semantic indexing, known-item search, and content-based copy detection,.

PolySmart @ TRECVid 2024 Video Captioning (VTT) VIREO at TRECvID 2010: Semantic indexing, known-item search, and content-based copy detection,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:00.933356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:00.534726Z digest=sha256:63db4e84687503286aa2f85d8fdd31946a1dd98f558abbf3cc8156ed0081b429

Observation e62ab536-ff54-41f8-9773-150dcd8d4cde · outbound

This paper cites VIREO@TRECVid 2023: Ad-hoc video search,.

PolySmart @ TRECVid 2024 Video Captioning (VTT) VIREO@TRECVid 2023: Ad-hoc video search,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:00.916192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:00.540959Z digest=sha256:bce15d99d5fa5cb4daaa494cc7d1f85ca91bf5f5c03b884d92bf893a27f31f96

Observation 3845033d-b1c6-4023-b917-807fadc6ff2c · outbound

This paper cites VIREO@TRECVid 2022: Ad-hoc video search,.

PolySmart @ TRECVid 2024 Video Captioning (VTT) VIREO@TRECVid 2022: Ad-hoc video search,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:00.896375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:00.546739Z digest=sha256:407d89af656dbc438a41f722b2f67a2927b68031a78992207169c04ce92ee769

Observation 00832043-6535-4a5a-941c-085441af67d1 · outbound

This paper cites VIREO@TRECVid 2021: Ad-hoc video search,.

PolySmart @ TRECVid 2024 Video Captioning (VTT) VIREO@TRECVid 2021: Ad-hoc video search,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:00.878814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:00.554472Z digest=sha256:4093e25d60d240b6483f5bfe5d9bae6205fba661703474f4c7a0cfacd2a3116b

Observation 272d7590-e842-4cf6-a3d7-6ed271c90cc3 · outbound

This paper cites Improving interpretable embeddings for ad-hoc video search with generative captions and multi-word concept bank,.

PolySmart @ TRECVid 2024 Video Captioning (VTT) Improving interpretable embeddings for ad-hoc video search with generative captions and multi-word concept bank,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:00.860709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:00.559275Z digest=sha256:3a0dd8762f1b9c7797f5ee9e6378eb07f77e3501dfe78b6bb1f43d337dd884b9

Observation c1c70163-2f1d-48aa-b75f-5aad923e2c80 · outbound

This paper cites (Un)likelihood training for interpretable embedding,.

PolySmart @ TRECVid 2024 Video Captioning (VTT) (Un)likelihood training for interpretable embedding,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:00.845709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:00.564527Z digest=sha256:d6531c59c31430d30efab110f8cc87e593a3320f19fcb5818dfab87bb69818be

Observation 88062a10-e92e-4b58-be60-098772b21004 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

PolySmart @ TRECVid 2024 Video Captioning (VTT) LLaMA: Open and Efficient Foundation Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:00.569639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:00.569639Z digest=sha256:9322c05585ab828823ad858816fb13e68651e963d372e7edbf42cbea27a6c9dd

Observation e3034e5c-4625-4e19-8925-9b70123f5a56 · outbound

This paper cites Visual instruction tuning,.

PolySmart @ TRECVid 2024 Video Captioning (VTT) Visual instruction tuning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:00.830461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:00.576951Z digest=sha256:98da87f550ed7103dea7982ab7cf28d51e52ac8d2dd5dab200ba9d2e90b10e26

Observation f6a2133a-fd79-47cd-9279-a99ecfbf533e · outbound

This paper cites Interactive video search with multi-modal llm video captioning,.

PolySmart @ TRECVid 2024 Video Captioning (VTT) Interactive video search with multi-modal llm video captioning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:00.813077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:00.584366Z digest=sha256:2e67567f2d68e1de8be8a9136dda08a8674f99e0ace64c025b5fe6f6eafd6ca2

Observation ebb78dfb-1ae9-4a2a-a40d-bfa9bd502f71 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

PolySmart @ TRECVid 2024 Video Captioning (VTT) LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:00.590366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:00.590366Z digest=sha256:cb4890bce4bd4a2a7133995d1fe2ee597f641d0a8553f670281378fe44d570ac

Observation a2da8ecc-af18-494c-b40d-368f0fddc3c0 · outbound

This paper cites Improved baselines with visual instruction tuning,.

PolySmart @ TRECVid 2024 Video Captioning (VTT) Improved baselines with visual instruction tuning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:00.794576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:00.596241Z digest=sha256:eff930e19ce0d3b47631d69dcef04902423166f3aab6d17c606bbe48d2c682be

Observation 6a3ea1f6-7e26-4449-98c2-b39dd02f8ac8 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

PolySmart @ TRECVid 2024 Video Captioning (VTT) MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:00.601226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:00.601226Z digest=sha256:6719c19c1df9f478ba5fd5ba150fb41ae579ad4aa1d9bcdc2a44014913c4cec6

Observation 02abd44f-134c-4100-a72a-0a296d7517ac · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models,.

PolySmart @ TRECVid 2024 Video Captioning (VTT) Llama-vid: An image is worth 2 tokens in large language models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:00.778607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:00.605749Z digest=sha256:733754d59d341aaa3b5942196b2f5c4ca52b093e6dfce475ea9a4015485fd3b0

Observation 4a09f899-50b1-4187-a53e-e6e38c4a242a · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

PolySmart @ TRECVid 2024 Video Captioning (VTT) MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:00.610637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:00.610637Z digest=sha256:c57f293a2730591f16dd0a50e103b85061fb04631849d3ef9c81185668832f62

Observation 056687fe-f654-4718-8e92-133754ef86b0 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

PolySmart @ TRECVid 2024 Video Captioning (VTT) Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:00.761934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:00.615171Z digest=sha256:5affe38019c64eef6e10326d7946e1655b0258c6bc34b495c8625d2f34fe146f

Observation 9245b8f2-a6bb-41e3-a93c-4ef49887c28a · outbound

This paper cites V3c–a research video collection,.

PolySmart @ TRECVid 2024 Video Captioning (VTT) V3c–a research video collection,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:00.747106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:00.619658Z digest=sha256:bb1c9ae6557ef8173d13ee462dd98ada3b872b05fee60d07d6fcce13c720337d

Observation 48398c15-e2ad-44f2-bbd2-3feab8eace6f · outbound

This paper cites Evaluationcampaignsand trecvid,.

PolySmart @ TRECVid 2024 Video Captioning (VTT) Evaluationcampaignsand trecvid,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:00.729784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:00.624313Z digest=sha256:55a92ef1b18afd3c973587a72943ef438e1708dffd95a0e0bd14ffcd0f3ad36f

Pith citing papers

No inbound Pith citation observations are available.