Pith. sign in

Paper Citation Record · LEDGER

Tarsier: Recipes for Training and Evaluating Large Video Description Models

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 44 inbound Pith citation observations for arXiv:2407.00634.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.00634 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:40:23.734088Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:40:00.906584Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7868b652-9b62-4519-befa-8237bb57d600 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.686898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:4d56a2b9397f49f053ebe8890694ca0015a454d60f7b6deea368a11f9690a964

Observation 0be139c6-460b-4b87-ae5c-c34e807f64df · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:20:32.962559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:aff622624e98e1cb2f5691edc83496007cd1e7a9f224c2edb34c1f742d4d0d2d

Observation f6822db5-7cbd-4ebd-8023-6aa5e8df054d · inbound

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models cites this paper.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.734088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.734088Z digest=sha256:81cd2e5f0a7a8f6f84a4f0a7c6757ce87e6e6c441a4ab4ee58d2cd97e8a5fdc0

Observation e108ea2e-3b48-4b8c-91e9-463ea9edbf50 · inbound

Open-Sora Plan: Open-Source Large Video Generation Model cites this paper.

Open-Sora Plan: Open-Source Large Video Generation Model Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:42:45.194961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T08:38:27.946746Z digest=sha256:43358d1dd77039dea5457ab275cf3598e0857dfac25ae096abbc8c2b79d52532

Observation 8b4458b4-060b-4a33-bdfa-057b2f6d6052 · inbound

PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation cites this paper.

PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T05:16:59.601418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:16:59.601418Z digest=sha256:3d5ea6c42035fb679f6e0d8f2d8861278374ac70425a3c85d1b37c2578e9d6c7

Observation 3ee4e1f2-cd67-4798-a8a9-e436db706eac · inbound

Progress-Aware Video Frame Captioning cites this paper.

Progress-Aware Video Frame Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:58.545529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:58.545529Z digest=sha256:5db8128d1a93c848c6cdb904e1818b486492c579e241e222e292f2e3c5a45cb7

Observation 1f9c87a1-a224-42c3-87d2-2b5769e407f0 · inbound

LinVT: Empower Your Image-level Large Language Model to Understand Videos cites this paper.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.477303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.477303Z digest=sha256:b3d553a8df15bd684861be92c6242051e68bda1644cfcc0652abfd0dcbad227e

Observation c0db57b7-51b8-4e78-8b50-f50461a269ac · inbound

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models cites this paper.

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:42.664942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:42.664942Z digest=sha256:429743a40384cb1607234bbedbc92ba08a9104cff9e0733cca8248c2c6924cb3

Observation 0b4c1064-cd9d-4bfc-afc1-2fc19c77a335 · inbound

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval cites this paper.

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:00.945147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:00.945147Z digest=sha256:79e59dc483733ea368c3192cea8af42ea39f3c923b5204c32446d0253efceed5

Observation 29df199b-4fd9-4a7e-9473-d21570d2bad3 · inbound

BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos cites this paper.

BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-09T23:03:34.652174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:03:34.652174Z digest=sha256:7cf60601c2de67dd247f16e83f637aab632f5b7c604d32d3efd6baea2a0f15e4

Observation 63f8375a-2a2e-47a9-a9bf-2f2200751a76 · inbound

Goku: Flow Based Video Generative Foundation Models cites this paper.

Goku: Flow Based Video Generative Foundation Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-08T21:07:32.440706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:07:32.440706Z digest=sha256:1d775dfd50d9e6301193c36f186b01c2b661993abe74da2445bd5e081fff5b14

Observation 1ba35c61-2595-4c3e-81cc-798898ead23d · inbound

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning cites this paper.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.869285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:16a2284c4aaf31dad4973bdee3ba9e8257016a1ad97ec41df788e4a45584bed6

Observation d31b4ecc-a26f-4da3-b04f-050dde267969 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.594061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:98d3e686361266f41a0c0816d4fab94664e55ed29c9798e245702c44f544c8b5

Observation 4165373e-c41e-4f73-bb44-aa18251c6109 · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:44.083171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:44.083171Z digest=sha256:b76d9d8ba342f803dc78931ef12a63249821d898120f719b5de68581f1d381fb

Observation e152143e-f3e1-440b-be7d-f896473bbd0b · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:04.111822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:04.111822Z digest=sha256:aa5e6da8f9468e709e57bede9f332b330692e215c33720b737f5f4ba9f959dc4

Observation 8f47666e-7b6c-4623-bfbb-7765955299fa · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.490044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:f0fda8301703f5df85a3fe8e8803c96a21041c11cd628bc8e03373c2f9c7be12

Observation c314e2d9-ad47-490c-a7b0-549e292d25e3 · inbound

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking cites this paper.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.885400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:35.885400Z digest=sha256:99436df81d52378552ab671117c6c516ebcd1bece38eca988003b1b04ebb37b7

Observation 543f8e17-0ffc-4fd6-aac0-f39935557265 · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:50.622545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:50.622545Z digest=sha256:4effd050e9e6dc2b046d4e709f277119c9bd4e71a41420bf2c7e3d5baffbc079

Observation b82aeb2f-ca6e-4fb8-a96a-3af13f1bd4a5 · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:10.194652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:10.194652Z digest=sha256:37f24c92b20a0130d921949bca5b7ef4b33654ed69278d2510510b626d8de14b

Observation 0cb707a5-6917-4e9f-b00d-1c255f096643 · inbound

How Important are Videos for Training Video LLMs? cites this paper.

How Important are Videos for Training Video LLMs? Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.884760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.884760Z digest=sha256:56c0b72bd39ee742784b23eabcd9e8eff1f9a2d04bd8b4fdf438d564f88339c5

Observation c455dc02-2b25-4543-a95a-d0536125090a · inbound

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering cites this paper.

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:33.251744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:53:33.251744Z digest=sha256:bb425a668bda037d0413a3eb3278f3d1e8a243cb1c8525bec410ac4a9e1f1c9c

Observation b62dbf61-4e24-4aa2-be39-3f4e90809c3f · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.933713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:d2068e5797ab187247e21b84d94e3dac7cbc0028b11f68ecc4edc2ebe22876f5

Observation 245512fd-c346-4587-9d79-985335bf2157 · inbound

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning cites this paper.

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:02:42.441350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T10:02:20.477517Z digest=sha256:d90abe0e38293640c764b700a84c34e630bb06b17f119c8f8d25920ac6dbf900

Observation acb55947-e818-425b-b074-60982699ea2b · inbound

SCP: Spatial Causal Prediction in Video cites this paper.

SCP: Spatial Causal Prediction in Video Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:50:11.211833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T16:47:44.523606Z digest=sha256:03053ae37ae4ef528ab43545151b538f5494de54249de25c3dd7886fde1ec8dc

Observation 31ef9471-cb5c-43dd-924d-6b79638221e1 · inbound

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding cites this paper.

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:43:14.858948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T20:40:41.380829Z digest=sha256:4ca7d69aeb85f4c916fce70d66dba70a175758c61a8a2364b5fa86c4887a0044

Observation 3b876ced-1ea4-4dad-ac02-ba66176a193c · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.401243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:aba57d10076be3868a3e91b4c57cd607f0148e04ad6b35756d47b4c9e3dc671c

Observation 97a2f1b6-2f1e-4d05-bab6-e6c6324a2baa · inbound

Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction cites this paper.

Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:53:28.934400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T01:51:57.018809Z digest=sha256:823a87130fdd66f21a80187b2b5643e48beab03cc0b647e43329ee1d9887695f

Observation 38714149-295f-4abe-9ab5-d23f4216ba24 · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:28:04.465276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:f0e306084f6df169008fe6e5fc89e74d49cdb45a7a4f684b90fba0591862a2f8

Observation c55f7ec3-6104-41fc-870c-a9ab35371b3e · inbound

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning cites this paper.

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.434241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T13:13:57.599970Z digest=sha256:694ceb5f6e6b4eff06c91365d5712e39dbb22b776d53a6fcaff19d5de034ec7d

Observation 779d8d9e-b6f3-4051-86a8-a7c4d19b3f41 · inbound

UNIVID: Unified Vision-Language Model for Video Moderation cites this paper.

UNIVID: Unified Vision-Language Model for Video Moderation Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:07:09.272703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T22:56:26.674841Z digest=sha256:678ff11473e6a8355bc01d6f31436efbec55181902e1ea3700314941f87d2aef

Observation d58a1299-f53b-4cbb-8f78-1f7c9e333eba · inbound

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models cites this paper.

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:09.548729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T22:49:18.420491Z digest=sha256:d3fa1d4858276926e2e5160b70d34730d48687d7ca51cc96f673885a1a16e0b1

Observation 6909c904-fb97-4a95-aee1-5bbac09e37f0 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.770571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:f5db5b1f77a45df26e26b99b54019f22933b118eec0d4a2d29f5e32b112e10fd

Observation f129b174-1b38-463b-a386-c40262bf78de · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.312897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:a38fbf7e291cacf5acf5a23f4ffe87c91c3eefd3357efafb83d3011b3668ef42

Observation e33d93dd-136a-4769-8026-647da8b3bed4 · inbound

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA cites this paper.

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:30.176143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T16:55:35.743040Z digest=sha256:beee702d5fc66f9df02e28dad86e85c83bc2bd3c43ccf2e52ebfa1f0c80af74c

Observation 9b7429ba-3931-4175-861f-4f270f630a92 · inbound

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning cites this paper.

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.024209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T17:21:38.543724Z digest=sha256:ce3d5c814e57461a583ba673090968ca2374540a496bc72d57a589a97d8221db

Observation 2902fe35-5cb6-4ef3-83a3-04289e077e83 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 297

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.121610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:29e6da0921d0ec8f37724aaf899b09ec0e80e8cc610aac581fb05fa61d1b68d9

Observation 1efce4af-dece-45af-8191-c58b9bd5045a · inbound

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning cites this paper.

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:40:00.907967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-25T23:29:24.520537Z digest=sha256:b539e8628108d90bfd1379eadd439a80be63022b3836aee8ed925918d35aad2e

Observation 87c325a0-12db-4fe8-b67c-73b20cc29f7e · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.256819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:7f62cc3ba84b01414cf0bf71bbf40779b6161054fcc80b626776d22a6941ef9c

Observation de2f74fc-0519-4b49-ba64-ce5e88f3af4a · inbound

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos cites this paper.

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:24:21.158963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T07:21:46.783970Z digest=sha256:4199a40f8e8e650c4773816e7ec67bcdaf63e9586157dd21b73ec269861b9f98

Observation a0274c15-76c9-48f4-a8cb-a533aa800a29 · inbound

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning cites this paper.

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:38:39.648627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-03T16:37:06.384435Z digest=sha256:8b84fc00bd9c518b0e7877a10b9faef2eba1eb9ddf13d78bb43cabbcccac5edb

Observation 1a17a28f-d18a-4b89-92f5-7f5b45e0daa2 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 167

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:1cc74118d5e33cb34f652ef05b67145579b78b63bd3a29c9cc01f76284f3c179

Observation 0f4661c7-a4c4-4aeb-8791-76af60bd4357 · inbound

PercepCap: Video Captioner with Structured Spatio-Temporal Perception cites this paper.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.243631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.243631Z digest=sha256:24bd6186368732070c35dbaec8c118f5a6f4c2845c8aa77b20f35a1c4b8361ba

Observation 9643093e-bfed-4251-9083-b15b50c20ec5 · inbound

RefCaptioner: Multi-Reference Image-Grounded Video Captioning cites this paper.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.012837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.012837Z digest=sha256:8119607fc8591bf903e2e19a9c1501d34340401640a56b227ac2de7df05cbea1

Observation 90db9274-bbf3-46ad-8514-45a615392547 · inbound

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward cites this paper.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.059127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.059127Z digest=sha256:8f260a25ad0b891a89c4057cb1066482e2c38000d6a8d36a2a9cef59401b558c