Pith. sign in

Paper Citation Record · LEDGER

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models

As of 12 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2412.11391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11391 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:03:10.621431Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T21:55:17.388570Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T20:06:13.285932Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 530b7ad4-bc9f-49a5-8cfd-dba9fa042326 · outbound

This paper cites Improving cross-modal alignment fo r text- guided image inpainting,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Improving cross-modal alignment fo r text- guided image inpainting,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:03:11.161173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:03:10.492695Z digest=sha256:d38571c0629471126560b3f495857ab56ec0937634babfa37b5187b6ec36fbb0

Observation 4b571c40-358a-438a-84ad-99db03fbd50b · outbound

This paper cites Visual semantic role labeling for video understanding,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Visual semantic role labeling for video understanding,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:03:11.146367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:03:10.497875Z digest=sha256:3efe4cba54b44da8cf4060bbe54da6ade584761672601c5d90de6c18cda491e8

Observation f4e8022b-8435-4489-814f-c3b5f11a9ff7 · outbound

This paper cites Learning transferable visual models from na tural language supervision,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Learning transferable visual models from na tural language supervision,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.501895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.501895Z digest=sha256:4678fc7e85a2a1f8091a2e14cf741803addf58da8ca1100adf5dbfb0440e2d0d

Observation e91f26dc-2b2c-4e74-bb65-af1488614639 · outbound

This paper cites CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.506035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.506035Z digest=sha256:872795aa608e618a8d613945def6b72713a2266d97d0b4c528459fb93db3d07b

Observation a03d22b1-0641-482d-b3d2-473d1fd811a5 · outbound

This paper cites Video-llava: Learning united visual representation by al ignment before projection,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Video-llava: Learning united visual representation by al ignment before projection,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:03:11.124505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:03:10.510994Z digest=sha256:3413347dd2f374e37f2025489019d7e8d0e0a883c30c72a19ae322acca24b0b2

Observation 609e8ad4-1268-4d79-a882-7571ce8c09f2 · outbound

This paper cites Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.515704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.515704Z digest=sha256:e58c8c79ae03452cc386be726d7c5458f27b541d09ad3176b7d73bbc4135d6de

Observation 5165f5d6-26ab-4088-baea-38512f4dd074 · outbound

This paper cites Visual in-context le arning for large vision-language models,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Visual in-context le arning for large vision-language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:03:11.111691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:03:10.520540Z digest=sha256:f33eed681b134fa5b064934b4ea15e1f553b00afef396ebf8c843a01a0b139cf

Observation 7b85c097-d117-4282-bc78-9a6cf8be69cd · outbound

This paper cites An Introduction to Vision-Language Modeling.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models An Introduction to Vision-Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.525269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.525269Z digest=sha256:fbb56189aa1cd04421a75afa77b06cb9dce71cb39e1e84a0bb7e8aaed7a25edd

Observation a6ba79b8-b211-4aff-9698-fd29df02bdbe · outbound

This paper cites Triple sequence generativ e adversarial nets for unsupervised image captioning,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Triple sequence generativ e adversarial nets for unsupervised image captioning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:03:11.097325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:03:10.530307Z digest=sha256:d0bae145b5e62b5e13db55b2d6fcff42c314164c368c770a80555aeddb202054

Observation 13e4d032-0556-4830-963c-ac3152413ad6 · outbound

This paper cites Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.534778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.534778Z digest=sha256:61381cf9ca7326a005ca402eec1c65f8c96394391ff38d2c5d86065d8b14323f

Observation 2845fadf-6512-4b6d-a673-60fa1db1973d · outbound

This paper cites Style-aware contrastive learning for multi-style image captioning,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Style-aware contrastive learning for multi-style image captioning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.539315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.539315Z digest=sha256:364adf1c1fc2117b6b20d42ed0f6a09e975b6b096af1cc01ea25cb047fd7274d

Observation 3a129508-079c-4be8-b8be-12d100853811 · outbound

This paper cites Sketch storytelling,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Sketch storytelling,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.544086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.544086Z digest=sha256:b4d8a41caec155578e8d572156e3232ca6c24c6eccdb6294d3646163d9477930

Observation bb20cf11-9596-484d-8e78-98dcd360342c · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.548291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.548291Z digest=sha256:a23c9226cfab961b10a99a6eeeceb7f1049e6db47e5fb63b0edc175682b11f4a

Observation 6280c8fe-0ba1-4b28-a26b-5b53566389d8 · outbound

This paper cites TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.552828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.552828Z digest=sha256:679fa85abbb2cc940c3b34d2df2620c460f7e63bead364984aed46cf189e0db0

Observation cfec1ed5-b8aa-42b3-97f2-4f04d3fbab7c · outbound

This paper cites RelationVLM: Making Large Vision-Language Models Understand Visual Relations.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models RelationVLM: Making Large Vision-Language Models Understand Visual Relations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.557467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.557467Z digest=sha256:c7fb41b130a2694c538e3c1e3e2e15c5789fae252e61fbdbcee00aef0a085fbb

Observation df81e3f0-51c2-4e88-8ccf-aa11e3bbb0b7 · outbound

This paper cites Multimodal event transformer for i mage-guided story ending generation,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Multimodal event transformer for i mage-guided story ending generation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:03:11.061766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:03:10.561787Z digest=sha256:f7693ef7cf87f930c9c4e95f9371cc28ec4f1903d9a39d322b659302500919cc

Observation 028d54c5-aad0-4a89-a7ed-b0f224fa0b34 · outbound

This paper cites Multimod al large language models: A survey,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Multimod al large language models: A survey,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:03:11.046464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:03:10.565867Z digest=sha256:35a8a3a7a5f34cc8ecf758d74778a33272b2151d5cf129184551461711f19b16

Observation 2955da0a-9076-4f0e-b18b-bdea223b1209 · outbound

This paper cites Thread of Thought Unraveling Chaotic Contexts.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Thread of Thought Unraveling Chaotic Contexts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.570064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.570064Z digest=sha256:475aaea206bec7efe3429b6e0313fe800b2694310b15437f5e34daf036ec616f

Observation bd4773ee-68fa-4ec1-8340-d4322bef8daa · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.575147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.575147Z digest=sha256:03a59dbc7e7d98bdafb47d63b30ed9636c0735241fb8a2f92e150bf8545d5303

Observation 83b130a2-be03-470a-953b-9dcad23c942c · outbound

This paper cites Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.581269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.581269Z digest=sha256:901add715f6d932558da418c27360aa622f4c71b7c7aff42b06ebad553c963be

Observation 51f9acce-d756-4f12-b611-3ee2ed4b045a · outbound

This paper cites TESTA: temporal-spatial token aggregation for long-form video-l anguage understanding,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models TESTA: temporal-spatial token aggregation for long-form video-l anguage understanding,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.591551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.591551Z digest=sha256:1bde4f3791792553b80486f36e4a23367ee6dc6294917bf125abd6c3b03e419c

Observation ea147321-d0f4-4ff8-8735-cbc5263c3825 · outbound

This paper cites Electrophysiological responses i n the ventral temporal cortex during reading of numerals and calculation ,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Electrophysiological responses i n the ventral temporal cortex during reading of numerals and calculation ,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:03:11.031554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:03:10.596776Z digest=sha256:80d6ac4f381fec7d4b933ad92884312786bfe0f0248a1ff38e96968762badd4f

Observation e421dcfb-d0e2-4719-8a6b-74238e67754a · outbound

This paper cites TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.601232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.601232Z digest=sha256:ee0250ee9bd7e1f523d2486706bd8fd5a50a2b7479c5ee1690e9169306b0fe40

Observation cb7e6470-1192-4b0f-92d1-2a9f66ea954f · outbound

This paper cites Enhancing video-language representations with structur al spatio- temporal alignment,.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Enhancing video-language representations with structur al spatio- temporal alignment,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.605001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.605001Z digest=sha256:e31aff9362c8bf4f5de6afb792b29b4f5598bec7a0803f0398a6e7b6e779d123

Observation 6649c21a-592f-445a-b08e-a77b3f96058c · outbound

This paper cites Temporal Sentence Grounding in Videos: A Survey and Future Directions.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Temporal Sentence Grounding in Videos: A Survey and Future Directions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.608963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.608963Z digest=sha256:36337b605856c4cc1e5ffdb24a3a3232b595f5c86e6b01e9b5d98ca46afaaf46

Observation ced600e6-8623-4ccf-81d1-c6fb5dabc37c · outbound

This paper cites Towards Effective Time-Aware Language Representation: Exploring Enhanced Temporal Understanding in Language Models.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Towards Effective Time-Aware Language Representation: Exploring Enhanced Temporal Understanding in Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.613345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.613345Z digest=sha256:f64206ba18edb8360727b928852a8c7ac786b0e2968fe5e6c2728e10e32139c0

Observation c32036d5-1ebe-4b68-b961-09da0f4f71f6 · outbound

This paper cites Temporal Reasoning Transfer from Text to Video.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Temporal Reasoning Transfer from Text to Video

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.621431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.621431Z digest=sha256:48d9b57bd42968a5f2311217a562452270aa842f02170f10f1da6add34c55efc

Observation ebd473bb-d232-4b56-9282-6d7b0b9e27b6 · outbound

This paper cites Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T15:03:10.709172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:03:10.586265Z digest=sha256:8020adce607cd2b22a298727eba03d8165c2251380a78f50a3d340ef007c1919

Pith citing papers

Observation 8e505cdf-3815-43b9-b414-1c01695c8cb7 · inbound

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios cites this paper.

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:13.287967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T10:15:15.129358Z digest=sha256:b3f26b3b20d2e49f19943e31679e79fae2271ebd7c1dcd7b3024be08ed62f5c0

Observation f830223a-3a6f-4c7d-a475-893a5c70b8d8 · inbound

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation cites this paper.

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-31T21:55:17.388570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T21:55:17.388570Z digest=sha256:d89f5abb40e5502b4e04eacafd96f3fd31605887930c853e8eadf9d9b3d00941