Pith. sign in

Paper Citation Record · LEDGER

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models

As of 21 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2412.11391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11391 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:03:10.621431Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T21:55:17.388570Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T20:06:13.285932Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 530b7ad4-bc9f-49a5-8cfd-dba9fa042326 · outbound

This paper cites Improving cross-modal alignment fo r text- guided image inpainting,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Improving cross-modal alignment fo r text- guided image inpainting,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:03:11.161173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:03:10.492695Z digest=sha256:5dbbd130cefbbf6d8a0a46123fce67582d22dfab394dcbfc3b07c315ad2a72cb

Observation 4b571c40-358a-438a-84ad-99db03fbd50b · outbound

This paper cites Visual semantic role labeling for video understanding,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Visual semantic role labeling for video understanding,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:03:11.146367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:03:10.497875Z digest=sha256:0832bf1b57a77241f04f4a9801428e20e075984fc58db48ed7c65256401f459a

Observation f4e8022b-8435-4489-814f-c3b5f11a9ff7 · outbound

This paper cites Learning transferable visual models from na tural language supervision,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Learning transferable visual models from na tural language supervision,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.501895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.501895Z digest=sha256:befb6c7a35535f32a997d72661460c1943f3a200e13f854914274d3d2e63bf31

Observation e91f26dc-2b2c-4e74-bb65-af1488614639 · outbound

This paper cites CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.506035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.506035Z digest=sha256:981ef715c6d16c14a9e9f9e6fc3652b68abe5bca339554b67558b5b7108583b6

Observation a03d22b1-0641-482d-b3d2-473d1fd811a5 · outbound

This paper cites Video-llava: Learning united visual representation by al ignment before projection,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Video-llava: Learning united visual representation by al ignment before projection,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:03:11.124505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:03:10.510994Z digest=sha256:888e4f4a0286f10ec8547c9e8a843c36eb0401296dd43674bd6ea3e2ce377dfc

Observation 609e8ad4-1268-4d79-a882-7571ce8c09f2 · outbound

This paper cites Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.515704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.515704Z digest=sha256:bd4c12383f856574576fde82ba5dccabeabd4044445292ec842a1bc82f810e78

Observation 5165f5d6-26ab-4088-baea-38512f4dd074 · outbound

This paper cites Visual in-context le arning for large vision-language models,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Visual in-context le arning for large vision-language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:03:11.111691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:03:10.520540Z digest=sha256:09b00885c4025b40f99d683f575efbbf65727ea02511f44a8d468bb9206751bb

Observation 7b85c097-d117-4282-bc78-9a6cf8be69cd · outbound

This paper cites An Introduction to Vision-Language Modeling.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models An Introduction to Vision-Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.525269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.525269Z digest=sha256:8f4de842f46605b20174177dd02b309ff919cb4f87e12e4a481c948ea935daee

Observation a6ba79b8-b211-4aff-9698-fd29df02bdbe · outbound

This paper cites Triple sequence generativ e adversarial nets for unsupervised image captioning,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Triple sequence generativ e adversarial nets for unsupervised image captioning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:03:11.097325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:03:10.530307Z digest=sha256:8fc6f32090bcba020744fa2d2222d1fdae4efab3c987b8332dfb89564d67ab9a

Observation 13e4d032-0556-4830-963c-ac3152413ad6 · outbound

This paper cites Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.534778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.534778Z digest=sha256:7ab8720c9ee0977da088af40cc625d3a8209a2b024cb64fcb7c917cb752e6c06

Observation 2845fadf-6512-4b6d-a673-60fa1db1973d · outbound

This paper cites Style-aware contrastive learning for multi-style image captioning,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Style-aware contrastive learning for multi-style image captioning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.539315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.539315Z digest=sha256:5b313f6a215b6c5487812b77111901dabae7c89a8216093c8a3dca1ba70e9718

Observation 3a129508-079c-4be8-b8be-12d100853811 · outbound

This paper cites Sketch storytelling,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Sketch storytelling,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.544086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.544086Z digest=sha256:2893dc54672a889a458bd25ebf18ddd10dc3853abb3ce55aa40339111ec1f60c

Observation bb20cf11-9596-484d-8e78-98dcd360342c · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.548291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.548291Z digest=sha256:e7b2a02c059d3efdf456dc544ad353b27151f98da04ad1d4c679073fb57c095c

Observation 6280c8fe-0ba1-4b28-a26b-5b53566389d8 · outbound

This paper cites TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.552828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.552828Z digest=sha256:d1d361f51b6f13b87bd4723c07868305590b59cdac4f8ba520f7fa0e253ab55d

Observation cfec1ed5-b8aa-42b3-97f2-4f04d3fbab7c · outbound

This paper cites RelationVLM: Making Large Vision-Language Models Understand Visual Relations.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models RelationVLM: Making Large Vision-Language Models Understand Visual Relations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.557467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.557467Z digest=sha256:7378297a7b3713310b56514e5edae24628330973352cc2ae3b885288a83e7101

Observation df81e3f0-51c2-4e88-8ccf-aa11e3bbb0b7 · outbound

This paper cites Multimodal event transformer for i mage-guided story ending generation,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Multimodal event transformer for i mage-guided story ending generation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:03:11.061766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:03:10.561787Z digest=sha256:89c72e3d93e87c4879ad0696ddad519873c70ba843602b5b070c5ce9d47e838a

Observation 028d54c5-aad0-4a89-a7ed-b0f224fa0b34 · outbound

This paper cites Multimod al large language models: A survey,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Multimod al large language models: A survey,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:03:11.046464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:03:10.565867Z digest=sha256:ecb17acef78e2c50c67049d9b90184a395d5527b2dd0e71663691d4fee5c9cb6

Observation 2955da0a-9076-4f0e-b18b-bdea223b1209 · outbound

This paper cites Thread of Thought Unraveling Chaotic Contexts.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Thread of Thought Unraveling Chaotic Contexts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.570064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.570064Z digest=sha256:029b2e372ff3c5b43fc5e4a93ce36a2bc5d2c7421c6549bcd085a46b5dda6c65

Observation bd4773ee-68fa-4ec1-8340-d4322bef8daa · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.575147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.575147Z digest=sha256:2b3cd5ef4ac67822767a7457a3ec78a41fca16fd5a4aa8b9b76a7457e2c963a7

Observation 83b130a2-be03-470a-953b-9dcad23c942c · outbound

This paper cites Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.581269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.581269Z digest=sha256:63a3dfd9a0ea56373a8fabc189512d4155086770feb87581f1e268b10b2595e4

Observation 51f9acce-d756-4f12-b611-3ee2ed4b045a · outbound

This paper cites TESTA: temporal-spatial token aggregation for long-form video-l anguage understanding,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models TESTA: temporal-spatial token aggregation for long-form video-l anguage understanding,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.591551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.591551Z digest=sha256:1a2d4d539fcad963186b81bbeb84bbb2196093652afec32ea668991347982592

Observation ea147321-d0f4-4ff8-8735-cbc5263c3825 · outbound

This paper cites Electrophysiological responses i n the ventral temporal cortex during reading of numerals and calculation ,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Electrophysiological responses i n the ventral temporal cortex during reading of numerals and calculation ,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:03:11.031554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:03:10.596776Z digest=sha256:e8d836f8d690414d9f1dbd0226913b43d4d03ace4cdb2b6a8dffe531c2d6c5c7

Observation e421dcfb-d0e2-4719-8a6b-74238e67754a · outbound

This paper cites TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.601232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.601232Z digest=sha256:33eb50168cfd6ffa7216d63115b104a330bdd46d18295f4452609920d927ca0b

Observation cb7e6470-1192-4b0f-92d1-2a9f66ea954f · outbound

This paper cites Enhancing video-language representations with structur al spatio- temporal alignment,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Enhancing video-language representations with structur al spatio- temporal alignment,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.605001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.605001Z digest=sha256:8733ae71fb59ce934ac354187edf6c3c9b983484044211f0ffe6d6bf4ef1a843

Observation 6649c21a-592f-445a-b08e-a77b3f96058c · outbound

This paper cites Temporal Sentence Grounding in Videos: A Survey and Future Directions.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Temporal Sentence Grounding in Videos: A Survey and Future Directions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.608963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.608963Z digest=sha256:9ba9c1be790647b191d68d5c008ae3fcc40677dee2774d2b1adc3a0a8f744e8a

Observation ced600e6-8623-4ccf-81d1-c6fb5dabc37c · outbound

This paper cites Towards Effective Time-Aware Language Representation: Exploring Enhanced Temporal Understanding in Language Models.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Towards Effective Time-Aware Language Representation: Exploring Enhanced Temporal Understanding in Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.613345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.613345Z digest=sha256:acd7785e224cd4068fd752962c650750358f7da4ff6b0031ebea61069d9f75a1

Observation c32036d5-1ebe-4b68-b961-09da0f4f71f6 · outbound

This paper cites Temporal Reasoning Transfer from Text to Video.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Temporal Reasoning Transfer from Text to Video

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.621431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.621431Z digest=sha256:8819f6f176d1028316ad9551b6c5d01f22dd6e129f013bef9c98eeb88abee79e

Observation ebd473bb-d232-4b56-9282-6d7b0b9e27b6 · outbound

This paper cites Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T15:03:10.709172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:03:10.586265Z digest=sha256:7176e6660fb22dead9eb5c704975ca0ad34b2664c04838dbf8de6aa0bffa9497

Pith citing papers

Observation 8e505cdf-3815-43b9-b414-1c01695c8cb7 · inbound

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios cites this paper.

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:13.287967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T10:15:15.129358Z digest=sha256:fa0ceb52c8847c8095eeb82ba7d599a31b39467f7a368c15c0514ae81ca07ac7

Observation f830223a-3a6f-4c7d-a475-893a5c70b8d8 · inbound

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation cites this paper.

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-31T21:55:17.388570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T21:55:17.388570Z digest=sha256:74ab7dff19dec19f0c289cf5fc8d6299b5e0dc8a378fd46e1a00d11302ae5044