Pith. sign in

Paper Citation Record · LEDGER

Tarsier: Recipes for Training and Evaluating Large Video Description Models

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 44 inbound Pith citation observations for arXiv:2407.00634.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.00634 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:40:23.734088Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:40:00.906584Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7868b652-9b62-4519-befa-8237bb57d600 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.686898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:7e13beccc31862181b94e611c511e82c8c6dd1326c1ae0f11e806a02e2faa825

Observation 0be139c6-460b-4b87-ae5c-c34e807f64df · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:20:32.962559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:b1e16b6c228ba61f809a43a49d17884871719a1885e365dda6fe6a8b37fe5d22

Observation f6822db5-7cbd-4ebd-8023-6aa5e8df054d · inbound

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models cites this paper.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.734088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.734088Z digest=sha256:b554de3f8acb7046113f17e94af2a37f684ec09f8cb82308adfb2feb9ba892ae

Observation e108ea2e-3b48-4b8c-91e9-463ea9edbf50 · inbound

Open-Sora Plan: Open-Source Large Video Generation Model cites this paper.

Open-Sora Plan: Open-Source Large Video Generation Model Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:42:45.194961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T08:38:27.946746Z digest=sha256:de1bbff98f2e79401575b960e1715a4644591cdb53f8048f39debfaea5976af7

Observation 8b4458b4-060b-4a33-bdfa-057b2f6d6052 · inbound

PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation cites this paper.

PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T05:16:59.601418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:16:59.601418Z digest=sha256:02f0b899904c999f01b78a3c59579397a26ce5f68ffb409a8e1f25fdb7992365

Observation 3ee4e1f2-cd67-4798-a8a9-e436db706eac · inbound

Progress-Aware Video Frame Captioning cites this paper.

Progress-Aware Video Frame Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:58.545529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:58.545529Z digest=sha256:df555c4162bc24cacfabccbe619c8e6f8e2e37772584d0e5c3b05e71e8478f7b

Observation 1f9c87a1-a224-42c3-87d2-2b5769e407f0 · inbound

LinVT: Empower Your Image-level Large Language Model to Understand Videos cites this paper.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.477303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.477303Z digest=sha256:1a25e1d1100243b6478deb007c0b37a568820dabee88d18025b20e73728207d5

Observation c0db57b7-51b8-4e78-8b50-f50461a269ac · inbound

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models cites this paper.

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:42.664942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:42.664942Z digest=sha256:452df65107c966511d75a3d5095200d1a21500a3d2d9f82673bd4bf23cb0ef97

Observation 0b4c1064-cd9d-4bfc-afc1-2fc19c77a335 · inbound

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval cites this paper.

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:00.945147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:00.945147Z digest=sha256:bbcec0809412b2e894adc557219b42d53106590f9adf3f8ccd8a5126307cacfd

Observation 29df199b-4fd9-4a7e-9473-d21570d2bad3 · inbound

BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos cites this paper.

BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-09T23:03:34.652174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:03:34.652174Z digest=sha256:60ce1bfeedb24721d39cb51e0930c9eccbd34a79d5b21f9902bb8919ee1ace36

Observation 63f8375a-2a2e-47a9-a9bf-2f2200751a76 · inbound

Goku: Flow Based Video Generative Foundation Models cites this paper.

Goku: Flow Based Video Generative Foundation Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-08T21:07:32.440706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:07:32.440706Z digest=sha256:17aed38e001f1bebe70219c5e19d3a17f29a2c2f4693c16f9aba3464b65946b9

Observation 1ba35c61-2595-4c3e-81cc-798898ead23d · inbound

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning cites this paper.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.869285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:b3753455e384f619d9960900631040b942d31a02c5081d3421888230cebbee49

Observation d31b4ecc-a26f-4da3-b04f-050dde267969 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.594061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c26ffe585b44b322962a0794935806f94f46fc5c0f1080477d3fb56fc4ec9ef8

Observation 4165373e-c41e-4f73-bb44-aa18251c6109 · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:44.083171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:44.083171Z digest=sha256:40e2e6a64cc3205b1b1ab64aef7aa68f50b4b95c1ca6d2b55684bd96c76e013b

Observation e152143e-f3e1-440b-be7d-f896473bbd0b · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:04.111822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:04.111822Z digest=sha256:47d52c44ceba6dea9a0c8b84606f75c72ef62a7bd0dde9658fc3b90c93416bf6

Observation 8f47666e-7b6c-4623-bfbb-7765955299fa · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.490044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:c196eb52deb9a5152b0a54f3e80856222cd38c223c95860620eee87ba308ec9f

Observation c314e2d9-ad47-490c-a7b0-549e292d25e3 · inbound

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking cites this paper.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.885400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:35.885400Z digest=sha256:10e3f084decd0fa962da3f2d454fee17637db760b4a5f33c9506c3800d3f224b

Observation 543f8e17-0ffc-4fd6-aac0-f39935557265 · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:50.622545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:50.622545Z digest=sha256:96fa899a46dfea2ebe8f53f029f8aeb5d3776f1127c0289da449e2b52946dc13

Observation b82aeb2f-ca6e-4fb8-a96a-3af13f1bd4a5 · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:10.194652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:10.194652Z digest=sha256:dd6d5269c3a20a01ffc714714b54a6af43ac30c22e1e215922695480618de7df

Observation 0cb707a5-6917-4e9f-b00d-1c255f096643 · inbound

How Important are Videos for Training Video LLMs? cites this paper.

How Important are Videos for Training Video LLMs? Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.884760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.884760Z digest=sha256:bb331ade382170afa62e26e4b35d0ff1f43262f26528211462158c5987d28e90

Observation c455dc02-2b25-4543-a95a-d0536125090a · inbound

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering cites this paper.

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:33.251744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:53:33.251744Z digest=sha256:fcdf2d46ed0ed20dd11e95d8d37c176347c21aab5a31d9285197e9c992c6810d

Observation b62dbf61-4e24-4aa2-be39-3f4e90809c3f · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.933713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:e4cdb1d9c66fa8c9e484e09dac335505417578a98265512980456890f147d172

Observation 245512fd-c346-4587-9d79-985335bf2157 · inbound

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning cites this paper.

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:02:42.441350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T10:02:20.477517Z digest=sha256:78a9db1a430a399e6f43e8c3b20c08c35286df507c48f25ca591dd4e89a27fab

Observation acb55947-e818-425b-b074-60982699ea2b · inbound

SCP: Spatial Causal Prediction in Video cites this paper.

SCP: Spatial Causal Prediction in Video Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:50:11.211833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T16:47:44.523606Z digest=sha256:1686d47a0ac8265676ddc226266707243a640cf9f4f27738055dc7135ffdf0f6

Observation 31ef9471-cb5c-43dd-924d-6b79638221e1 · inbound

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding cites this paper.

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:43:14.858948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T20:40:41.380829Z digest=sha256:387ec6356a7284f35d43470975f9654c9a76fe2ba10b8bd7a8fd8737fe6d09f4

Observation 3b876ced-1ea4-4dad-ac02-ba66176a193c · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.401243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:d7b712de75e5defee3b69511406f78cc671a93adf430cabaa132301d3f23a34d

Observation 97a2f1b6-2f1e-4d05-bab6-e6c6324a2baa · inbound

Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction cites this paper.

Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:53:28.934400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T01:51:57.018809Z digest=sha256:bb712d855081e3dcf76e965eaad96aa831971be402789aa8fa4b26421b9fe643

Observation 38714149-295f-4abe-9ab5-d23f4216ba24 · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:28:04.465276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:3173a6d15cd966e2cb668e7de7dd48d7eaecea8364ec822ac154a755abea78e8

Observation c55f7ec3-6104-41fc-870c-a9ab35371b3e · inbound

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning cites this paper.

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.434241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T13:13:57.599970Z digest=sha256:456d4851486ebfdcc9dae56fca4ec7db665c3dce29a263465bbb487bcba21db8

Observation 779d8d9e-b6f3-4051-86a8-a7c4d19b3f41 · inbound

UNIVID: Unified Vision-Language Model for Video Moderation cites this paper.

UNIVID: Unified Vision-Language Model for Video Moderation Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:07:09.272703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T22:56:26.674841Z digest=sha256:ad9eac60cdb50dec720c7197b9b0190120075c9321677ab73bce78b45edcb64b

Observation d58a1299-f53b-4cbb-8f78-1f7c9e333eba · inbound

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models cites this paper.

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:09.548729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T22:49:18.420491Z digest=sha256:768456dbf276873291a95d50cd2f012b8b274b9b3fceba5ed64f2b874cbe3087

Observation 6909c904-fb97-4a95-aee1-5bbac09e37f0 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.770571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:af6c1425fabc859bb5de8ebecf546cad3901b124fb01abb39d5c0077145d8515

Observation f129b174-1b38-463b-a386-c40262bf78de · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.312897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:bf142e1bf15407ee7ef03d0f7c96a527fadc20c96d065f5e6c464160908e8809

Observation e33d93dd-136a-4769-8026-647da8b3bed4 · inbound

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA cites this paper.

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:30.176143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T16:55:35.743040Z digest=sha256:fbdddd393eb2e59571520ffdc2fffb8b84e363235e266328d169332d33cb5051

Observation 9b7429ba-3931-4175-861f-4f270f630a92 · inbound

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning cites this paper.

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.024209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T17:21:38.543724Z digest=sha256:7a2d0d53a806769307c0bbc0ece80353a0577601caf49a85b3ce43db10c80191

Observation 2902fe35-5cb6-4ef3-83a3-04289e077e83 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 297

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.121610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:c7f765271c409cd39afd8a644845edeb7d7c514f8c90922a160417eda30c411c

Observation 1efce4af-dece-45af-8191-c58b9bd5045a · inbound

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning cites this paper.

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:40:00.907967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-25T23:29:24.520537Z digest=sha256:480f1713fdd3c8b86d90212b855b3baed731d694874ca6ea55e6731d0e2eb489

Observation 87c325a0-12db-4fe8-b67c-73b20cc29f7e · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.256819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:e079c205476b4dec74f87f7e3863e790d4cdf4ac51228e1e2cde6de5cec96ad9

Observation de2f74fc-0519-4b49-ba64-ce5e88f3af4a · inbound

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos cites this paper.

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:24:21.158963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T07:21:46.783970Z digest=sha256:5f1264cb47e1fce9b67f3cef52b24e636ab5d11a1001137fa94b2e06bfb9a7f3

Observation a0274c15-76c9-48f4-a8cb-a533aa800a29 · inbound

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning cites this paper.

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:38:39.648627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-03T16:37:06.384435Z digest=sha256:07a76676dbdd350ad72d36d069a878524c093e0d2197d78887cc93280c4b997f

Observation 1a17a28f-d18a-4b89-92f5-7f5b45e0daa2 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 167

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:c69d7e0c261ca8952fa1f037bf1b936857cb743e5b80439658fc13690b68b935

Observation 0f4661c7-a4c4-4aeb-8791-76af60bd4357 · inbound

PercepCap: Video Captioner with Structured Spatio-Temporal Perception cites this paper.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.243631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.243631Z digest=sha256:26a62135d597131aaea02dcf1ac611406c15c9144a672713be2db57436cc8a1d

Observation 9643093e-bfed-4251-9083-b15b50c20ec5 · inbound

RefCaptioner: Multi-Reference Image-Grounded Video Captioning cites this paper.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.012837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.012837Z digest=sha256:f4d2706576f72dbadb0fccb827f3fc7c51b07108dd0e79422015263d8db40643

Observation 90db9274-bbf3-46ad-8514-45a615392547 · inbound

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward cites this paper.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.059127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.059127Z digest=sha256:ffe7afdcacf96b6f7a2a8f79bd223df7af463629bc7df077435aea68c28cbc6b