Pith. sign in

Paper Citation Record · LEDGER

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

As of 23 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 11 inbound Pith citation observations for arXiv:2506.01725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01725 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:57.860880Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:58:52.971482Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:40:00.934274Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a15c7f71-a533-4b6f-90c3-a2dca39bc095 · outbound

This paper cites GPT-4 Technical Report.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:54.001986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:54.001986Z digest=sha256:266273d412bb2feda5d2072a38a28b8cbc58d836d8ba4020c7b79940cf972d65

Observation be9052c9-8e93-4861-a699-4f87d8bb6797 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Activitynet: A large-scale video benchmark for human activity understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:08.019749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:08.019749Z digest=sha256:ad8fe5ee1259eb96afec18ac37896cbe87bd1a10ca3ea5ac3703fb84a19bc9ad

Observation b9c0c257-e40d-4a50-9b57-4793d2ddcc92 · outbound

This paper cites Auroracap: Efficient, performant video detailed captioning and a new benchmark.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Auroracap: Efficient, performant video detailed captioning and a new benchmark

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.563025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:40:08.047640Z digest=sha256:9254cd81b6aac312ff6b841441f3cf01b6e2a8fd4cdc6b17bc46aa89bf29d63b

Observation 87d10cb7-bbce-4570-894c-48ea936d37e7 · outbound

This paper cites M3-embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge dis- tillation.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking M3-embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge dis- tillation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.643118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.643118Z digest=sha256:b54a0c599b7fc40cad4f61e506793cd3fadeb4691d9d2a7d437047d48257967f

Observation 2ef9d158-91d2-4043-b920-430cefab504d · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.778162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.778162Z digest=sha256:262d8a7e15e456de89267f1d6a4e824e80b9c988dc67ff367b52c82fc31dedf9

Observation 03135c81-6c8b-4354-b373-5950dd5531e4 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.868692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.868692Z digest=sha256:5c454d258a407bfc8fdf9b3fa50d299463013069e9ae102576f4183563f4ebe0

Observation eb22eed1-3201-432c-9962-4d3ab340b08f · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.908851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.908851Z digest=sha256:af90ea80475ec9f46ab0732bc96bd3ced1e721a67b917d26b0be4ea7fca3eb45

Observation 41784efd-54cd-4147-bb3a-2519b8b8319b · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.939333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.939333Z digest=sha256:271ab6a4f9c8880dfc39f0472a5447382310f70db01ffffdc61eed61790a3c18

Observation 83d9c79d-4ff8-4bf6-a9f4-4936233864b4 · outbound

This paper cites Tall: Temporal activity localization via language query.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tall: Temporal activity localization via language query

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.301041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:40:10.032618Z digest=sha256:974ce134ca1029802eaab017e172fbcc18467dcd0b360331f94679096029b464

Observation 02ca2843-d1a7-4209-ba52-dea0399d67e4 · outbound

This paper cites Scaling laws for reward model overoptimization.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Scaling laws for reward model overoptimization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.223294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:40:11.502809Z digest=sha256:f00d989b5bb34479dd0b2502b6587f2f1a97b21891bb568199956316243a55fc

Observation 7c245ae2-6d72-4855-8bec-7c8f7ad09bc5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.701653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.701653Z digest=sha256:03f06ef59982cd5edbd4be22a2ba76636c0897b41e43f6e6f14f9f73d160bb3f

Observation 5eead944-f033-4430-97d5-61140c8c3ef9 · outbound

This paper cites Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.807797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.807797Z digest=sha256:e45a4775b5f078caf097751f0ec6ec69fd0af4dce2ad4dcf33efd55285612b0d

Observation 68037354-34b0-40c7-954e-572ffe521b97 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.883387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.883387Z digest=sha256:e789468bb38660691e0e954cb4627047be38cbebfc1908ee3a0017a3b8b8fc1e

Observation 5a3b0e9d-49b5-48eb-889e-a442eaca222c · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.956199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.956199Z digest=sha256:dc42ea985006c57c9bbb2ebedfc1e854675d0a1ca612689b3ea3163d8bdfc332

Observation 903b84b7-4c7b-4e9a-ba12-15ed97a07435 · outbound

This paper cites GPT-4o System Card.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking GPT-4o System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.978948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.978948Z digest=sha256:d22a6b32d9b6542ac98a82f23cd6e03c7ae945eabe6bc2a33abd22dac1b9bb6b

Observation 9b7bb67e-93e3-4e1f-9da3-d259de55f54a · outbound

This paper cites OpenAI o1 System Card.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.995527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.995527Z digest=sha256:72a64f23ecd250c07de3907befd94fc0469d2249c56b326c00dc74d3fcb18e95

Observation 8c2b9347-9f47-422d-8df5-d5a7eaa3ad00 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.012820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.012820Z digest=sha256:e4556c54b06a52e03f07b86135875f3916f1c17f23529af99045b4ed2232c104

Observation ff83fa10-ffe1-48b6-9b70-6a8e139e6a8c · outbound

This paper cites A shortest augmenting path algorithm for dense and sparse linear assignment problems.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking A shortest augmenting path algorithm for dense and sparse linear assignment problems

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.184267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:40:12.037643Z digest=sha256:d05026484d449f1974868a83b09288f2fe3e66792d4b62a7ec6cc619d00fdda5

Observation 3210f0e7-20a3-4d2f-9319-37e1bc1b2cad · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.061896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.061896Z digest=sha256:cded2cce26adad42b953f5afb2da64a42c1114dd6d32986a899b1375b0b0514d

Observation 64cd84fb-ff43-4cac-9c69-120eafb54e01 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.084921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:40:12.097371Z digest=sha256:c79c3f189a5c2e928aa28aa3084facbfd27a23fc58132d5df524c171771e6769

Observation 00491f7a-f816-44ec-89d2-873dd68e7f75 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.120821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.120821Z digest=sha256:ea3e655f51e93502892c6d378b0065c5184e89c7ab08dcb4f32695c701ce70c5

Observation e54b9f27-0590-4e6d-b7d4-57ac29bbd53b · outbound

This paper cites Let’s verify step by step.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Let’s verify step by step

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.167318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.167318Z digest=sha256:5d9b7ecdb6bfbff5e410aa4172e3c3a489b5e19a580a8acaba1c9793931652ec

Observation e962be59-6051-4662-b6ca-b4f3a45d3b9e · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.227193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.227193Z digest=sha256:40fc18609c5ccc221bc1e103de4e38560f4da4691132e571153e68dffe424125

Observation 929c53b7-30e9-4924-bd77-7fa1adf09ad0 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.317088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.317088Z digest=sha256:9ea800c66e2870e5f329d4abb0294513607e7c34292551dab706e3ee8e32a82c

Observation 463aad97-165e-47ee-95e4-5c492569d533 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:23.084340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:23.084340Z digest=sha256:d3311d90cf2df42f826e8831c28982b127e17ee2bc423b12dcff841d06880bb1

Observation c097e315-49c3-4fed-a816-21edd16b6721 · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era, 2024.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Introducing gemini 2.0: our new ai model for the agentic era, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.900706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:40:23.398842Z digest=sha256:d3e2982bcd0e2ed3653e9f6734cc805a65dec1b1b962faf8bda4d0776e697778

Observation 69971c8d-693e-4318-ba6c-b3316bc9db5f · outbound

This paper cites Proximal Policy Optimization Algorithms.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Proximal Policy Optimization Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:23.659852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:23.659852Z digest=sha256:1c025cf2ddc2d7f0cefb59397b9ea404c108daf93b90167fbfd84cc1dc4cfdd4

Observation f77a72e1-7acc-4f7a-a6fb-89a92ca927ce · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:24.131641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:24.131641Z digest=sha256:a87041016c834082a8849e9559ec550eba86e336bb96a02f9b02b47a961ebb29

Observation 9e13b280-d600-4202-a8e2-37e22e721598 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:31.437137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:31.437137Z digest=sha256:3f087a7331f3e787fc3cf6e3b54ded7274151b821c56e3c0b2750ff42a701f9c

Observation ecbc943e-6897-4c25-9e4d-093006c3b49f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:32.880469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:32.880469Z digest=sha256:3f0fa3566e1af84e43e48eec835ca579b8597688ce121d0f0e4742d24b5b2caa

Observation 077385f6-de87-4489-81d0-4e9c4156c2e3 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:34.931386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:34.931386Z digest=sha256:0ff25f5c197176863e643820252b4c124329d32ad964a784995df18d9835eea2

Observation ce95458a-a9b4-41d5-8105-9a804e398998 · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.727219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:35.727219Z digest=sha256:f33535b79f7feebbada52df4a3ff2c745760a4d47ce8f9df1b00511c6471e429

Observation c314e2d9-ad47-490c-a7b0-549e292d25e3 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.885400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:35.885400Z digest=sha256:6fe5a681a98d80b37962b6223c3368a4affbb84a149e8687fd3b98ea4b1d45e4

Observation 835dc834-d0d0-4557-bfba-38afbae77af2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.947370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:35.947370Z digest=sha256:5f74c19b79d43a7fb6a5a6f5e831d4df629def567afe30a5513ed03060e7e9be

Observation c1befbf9-e216-4ae7-bb19-6a771e49d146 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:36.023507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:36.023507Z digest=sha256:c48c3d14f713a6f22ad335a65adc02b6a70ed6f18c63d09016b0af62b0f1d75c

Observation 7be61c09-24ad-45a7-95ff-96a0817528a3 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:36.060722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:36.060722Z digest=sha256:29913c97d3572df31d046c5dfe5bbe0c85316fa1011baa4160b90d0310dd8b72

Observation c42b6749-3580-4a62-b6eb-d81860a6ce66 · outbound

This paper cites CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:45.357734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:45.357734Z digest=sha256:9db19b5dcc4cff6934cf220f388f63372c2ecc657393b17a75b764a2f7116cd8

Observation 84cb20a6-1314-4606-881c-7b4d6619932a · outbound

This paper cites Qwen2.5 Technical Report.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Qwen2.5 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:53.234744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:53.234744Z digest=sha256:d5cacbf1d681ff2a11cc6d4a51add4b95c556f459265744d974b6aa02152f8f8

Observation e19a02cd-8847-4a5b-b60b-5160165c8c7f · outbound

This paper cites Vript: A video is worth thousands of words.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Vript: A video is worth thousands of words

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.725983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:40:54.882956Z digest=sha256:cdda5c4784a68d7ad5fe0322aec19c4d7d7090b0d257fa8db0b8a1e72fbd6745

Observation cfe4d595-bfd2-4502-834d-9b58e82f605a · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:55.543904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:55.543904Z digest=sha256:90e4a1872c1246ee16dadeb4f2158ab91e98b9d2f466f546348ff661d2af3703

Observation 11382a79-e212-483d-acee-0e4633d8e18a · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:55.873734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:55.873734Z digest=sha256:ad0bbac765ca2b7040f1dd55e555778b68b22059dce906a87d8f78574ca63495

Observation b38cb08a-7b3d-4787-9fca-3455a433d275 · outbound

This paper cites Modeling context in referring expressions.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Modeling context in referring expressions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:55.915232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:55.915232Z digest=sha256:e4111a5ba35497f1d5b97331343c30aca7b60e1a3e75cbde4bc248f8c8980c0f

Observation 71e9c9c5-ee26-4ec7-af1a-7aa1e7509853 · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:55.923555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:55.923555Z digest=sha256:689e765d324162931c84075707a99c0f55a0de30eda48c8d86f4c7a7e785d218

Observation 02b25c7c-7ad1-44a4-891f-fc297f035177 · outbound

This paper cites 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.644904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:40:55.953394Z digest=sha256:2c1e1817dab9a5b73c190cc9a0925034dfb709ec07188602def572143e79cd03

Observation b58ff38f-0a06-422c-942e-7381f7305b9d · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:56.019147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:56.019147Z digest=sha256:b464b12582e4e0ed183cc5110525e3e54e133e475ba2aa14e4a296f653d8c1c8

Observation f26bff47-1afe-46bf-aed8-1efb8589414e · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Improve Vision Language Model Chain-of-thought Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.064569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.064569Z digest=sha256:717034e1239ed65147129fb3113dc1d4e4d44e8c86414156c36e13cfd99094dd

Observation 19ab9372-e3a4-457d-8ffd-1ef64bb7114f · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.687621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.687621Z digest=sha256:83d3b3179acc56bfa4ec6b80b5824f51c38b237e13945319dc5c0598fe249b2e

Observation bfc7a347-b255-43b6-97df-b3e9738fef0d · outbound

This paper cites R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.733558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.733558Z digest=sha256:2ac5dcd1522735e487af61b48dd0aed8499bd14f5762130be9f813592e9c159b

Observation 568a8494-8f7a-48f4-b86c-3834ef178039 · outbound

This paper cites MMVU: Measuring Expert-Level Multi-Discipline Video Understanding.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.775507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.775507Z digest=sha256:8fbf6e973f5526146acf1fb783505a8f1ae9b0f51f8e2b943f0ab9a8b0a871b2

Observation c016390e-52e4-4925-aa35-151145ec6a97 · outbound

This paper cites SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.810477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.810477Z digest=sha256:0310b0370d3d6d1acb171db0ad66290cec64bdfee3cdd139207adc9d20def9dd

Observation 48094e61-f3c6-4174-be5d-b770a4d99ebe · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.860880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.860880Z digest=sha256:dec171e2a202ab41ef96bc03207f1b97d55018591a274c621e9a1ca42d4a2851

Observation eb5ca7e3-48c9-404b-97c0-cf111daead75 · outbound

This paper cites an unresolved cited work.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:40:59.416096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:40:09.314893Z digest=sha256:e71c2738c1f66119d0fd8d07ab5b89efb8b66ecd90ae7fd1e3e4607d9ba2531d

Pith citing papers

Observation e5b183dd-b5a6-4e41-9796-1db3ffebe5ec · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 292

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:11.589395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:11.589395Z digest=sha256:4c80ce90f796c8fd7fbbbc9ed1bfba4e5aa7b4ecd62cb042b374d15f8802d686

Observation 2c65558d-f0b6-46a8-9c1d-539e9579eef0 · inbound

OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward cites this paper.

OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T16:58:52.971482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:58:52.971482Z digest=sha256:3e6a62a89347d1c4a584313230588d6b9224beae2bead15ebf6fe5f594939ccc

Observation 4ff98ddf-19d9-4d22-a939-46b89e999320 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.561845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:8c0f53edfe163cc2d29fb4947e90aad5b6be29a4c664975754a31e4b688efffd

Observation 36e73833-b17f-49e4-9a20-48e792f14e45 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.484250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:724c872e06f3d48b5a6ff7e3ea77b5be0459b9a02341e258bf503cf3135901f6

Observation ef22b1aa-0882-4abb-b11b-72a052fbcb76 · inbound

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning cites this paper.

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.493845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T13:13:57.599970Z digest=sha256:8af8d57265bacf5eddf276b39f7d10cb2eff46f1e0c0abf4ab373a2f96a358b1

Observation c8439cfd-3612-4e2a-a1f4-fcd335241947 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.755390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:6cf4a1f4253705dbee38c799a2f50cd64e0577ec14c1fbbed13f793c6b859bfb

Observation b6a4f68f-becd-49ee-a3bc-6334f5766267 · inbound

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning cites this paper.

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:40:00.935800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-25T23:29:24.520537Z digest=sha256:8328c586ed8f31031ab21a08bc891a889323ae487955da24b76ac1e421e6be9f

Observation 9066af67-d826-4a5c-b0d1-5ef885e470a5 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 194

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.942635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:1c7e426519ca4a9fd27636510c4237c95dbc162ed98f5700dd77009719e3a858

Observation b32ac67e-4cd7-421b-a1e9-b83d10d18411 · inbound

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships cites this paper.

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:22.196718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:22.196718Z digest=sha256:9dc818f5d2049f6b9bdd3d7c3a8789b4eda6d36c78d6a4027bdba625f3705168

Observation 32461e25-5782-49b6-a010-1cc41b265953 · inbound

PercepCap: Video Captioner with Structured Spatio-Temporal Perception cites this paper.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.207694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.207694Z digest=sha256:a5cf295f0ddbb48b5ebe28bee226e0a21434e61d9fb01202b0346de661ef1c97

Observation 53fcffc3-aa5f-49b5-ac92-20dd6ee28176 · inbound

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward cites this paper.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.235558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.235558Z digest=sha256:19cb4f8d2659703df2027a0d5df8234aa19a7f7473872aa5d98f85840702f094