Pith. sign in

Paper Citation Record · LEDGER

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

As of 13 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2607.28509.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28509 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T05:08:20.197862Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved42
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9e5ef122-d5b2-4e1c-ac12-b90a3088b2d5 · outbound

This paper cites LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.737340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.737340Z digest=sha256:59b6f47de1bd57fc39ecf6a5a5c630cbbc57802d92b297af9dfa7b553fd3fed4

Observation cc44a2e2-fbe3-4938-89cb-ed8987f1fb75 · outbound

This paper cites Qwen3-VL Technical Report.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.750764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.750764Z digest=sha256:3038673af7b1f98f486340151606452944d30eb6ad1785fbfee7c415bd4dff04

Observation 001d8a6d-16d8-4517-a577-0e49064afd8d · outbound

This paper cites an unresolved cited work.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.763797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.763797Z digest=sha256:008bd2a38960a0d92e47d173fa45bcdec33b9bf4bb3165c252f69c393a39fbe1

Observation bbb30ee5-66a8-43c3-8085-14267234ec98 · outbound

This paper cites Multi-subject open-set personalization in video generation.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Multi-subject open-set personalization in video generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.780721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.780721Z digest=sha256:27bd15acc1abc091ff4beb2b390a89e3a6ef7959fa6511a4f9789ba5d5a96fa1

Observation d96553f7-95ff-41d1-97b1-a146976fc099 · outbound

This paper cites VidCapBench: A compre- hensive benchmark of video captioning for controllable text- to-video generation.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning VidCapBench: A compre- hensive benchmark of video captioning for controllable text- to-video generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.791672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.791672Z digest=sha256:31b09f285be8d6c53d301dbe8b9a53c3833bbc412eacb6960b7d93af38007445

Observation 7439c6ac-9db8-49f3-b763-50921b4a1af2 · outbound

This paper cites CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.804511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.804511Z digest=sha256:b9ef890418a13aaa1d519f0a34ee68b78f3b38c47a2c930cd614e6f3a44ce9e9

Observation 7c3ba9da-2809-43e7-ab9e-ef0315868f10 · outbound

This paper cites MA- GREF: Masked guidance for any-reference video generation with subject disentanglement.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning MA- GREF: Masked guidance for any-reference video generation with subject disentanglement

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.817785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.817785Z digest=sha256:06a771ba01b93a6786ead69b3fe2b8e41a9c8d078177c8d0119fbbf3cc46fee5

Observation 6d25759c-5f23-49ae-9276-95bc2c367031 · outbound

This paper cites ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.828792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.828792Z digest=sha256:83762ff1553f7413fcc6ef50e2b84ea5bf02e55474a4fbd4eabc112e6c4215b3

Observation e930ff39-68a4-4b18-a102-22993e62df74 · outbound

This paper cites Video Re- Cap: Recursive captioning of hour-long videos.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Video Re- Cap: Recursive captioning of hour-long videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.840883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.840883Z digest=sha256:0b065c1529b692232b48e04a1c96eefa92f62f6aebbee0fcfa20eabf6d3c73ff

Observation 530a17a1-e33c-4a7a-99f5-0ece07dd6da9 · outbound

This paper cites Movie Weaver: Tuning-free multi-concept video personalization with anchored prompts.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Movie Weaver: Tuning-free multi-concept video personalization with anchored prompts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.850220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.850220Z digest=sha256:9cc8475c8807f78b827d56fe9f1c835f210978cde5c654c4fa57fb8e8eee2d59

Observation 44944426-2480-4b27-9826-53eead274374 · outbound

This paper cites Video-LLaV A: Learning united visual repre- sentation by alignment before projection.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Video-LLaV A: Learning united visual repre- sentation by alignment before projection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.864718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.864718Z digest=sha256:bcadf4ddfcdbf0b7891f0ae2165afef666a9573622ed371825022cc1bb50903e

Observation 649bffd7-b00c-4762-8276-9ebe87696f6c · outbound

This paper cites SwinBERT: End-to-end transformers with sparse attention for video cap- tioning.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning SwinBERT: End-to-end transformers with sparse attention for video cap- tioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.874880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.874880Z digest=sha256:55347673c2c7342a23ef8eef1d14bfdd74559952327953b4a5e23c34806ca346

Observation 381b38d9-d150-4018-8fed-6a2b508ea59d · outbound

This paper cites Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.882679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.882679Z digest=sha256:f8886d61c3fb300553764341244378525855f2e6d279d7d9f829105e742e182a

Observation 936c2fd8-3e98-48ca-8f67-3c4ceecf8b7c · outbound

This paper cites LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.892043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.892043Z digest=sha256:6798e2afc780d215ac0cc103ee3058505a4bbef0a424bf9a859460b6afa2ce78

Observation a3f69209-b7f3-4682-af80-3f0d9950324d · outbound

This paper cites ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.906135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.906135Z digest=sha256:2c2098134840726e8e0f180d9239e4d11c4e5e0eef0e37d4808450a137371e21

Observation a68f4953-7a95-42ca-948f-95555a8fc5cf · outbound

This paper cites UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.913515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.913515Z digest=sha256:80c09fb0935966dd138453674ea0e1ac7cc769c48668b4d3e8b6cadc5ebe2799

Observation 615d3bc5-6a95-4e2d-b61d-a1eebcacc36c · outbound

This paper cites Video-ChatGPT: Towards detailed video un- derstanding via large vision and language models.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Video-ChatGPT: Towards detailed video un- derstanding via large vision and language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.920877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.920877Z digest=sha256:3a8b70ba0dfff282d3a1bb5209504dcad8ca655b798d0c9953fcba8ab9842045

Observation f9e5cb05-37fb-4071-bc4d-d9d543497470 · outbound

This paper cites Mavors: Multi-granularity video representation for multimodal large language model.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Mavors: Multi-granularity video representation for multimodal large language model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.928150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.928150Z digest=sha256:f1a72005235becb4cb8bc9404bba7d5c3c1acbcf4138a462e482372ab50c55e4

Observation bfe5d3fb-b499-4e26-bc43-6dd95317afaf · outbound

This paper cites Mme-videoocr: Evaluating ocr-based capabilities of multimodal llms in video scenarios.Advances in Neural Information Processing Systems, 38, 2026.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Mme-videoocr: Evaluating ocr-based capabilities of multimodal llms in video scenarios.Advances in Neural Information Processing Systems, 38, 2026

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.934415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.934415Z digest=sha256:9a671da567fc35e6f0cc2bee77c22d824517353bfd27a88f5140aa5f4a0dceb7

Observation e622f811-a402-40e9-a67e-af992d2c6438 · outbound

This paper cites MV-S2V: Multi-View Subject-Consistent Video Generation.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning MV-S2V: Multi-View Subject-Consistent Video Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.942411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.942411Z digest=sha256:88b25f5d0280bf7859cf827b3806be639b1aea616fb25c501122dc01381bd48c

Observation c6eb2ebc-b5d9-4832-800d-c3cb37952d86 · outbound

This paper cites VideoBERT: A joint model for video and language representation learning.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning VideoBERT: A joint model for video and language representation learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.962027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.962027Z digest=sha256:a4f93c47e43d7c72a21ec643f2483dcf7b7d7183258ad657bc425719837c567b

Observation 0e950d10-d95c-4277-b5ac-2d157881787b · outbound

This paper cites KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.970664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.970664Z digest=sha256:37316061af9144b13284a44b2f4aad20a90a9dfca647dbc899da1973ec1107c9

Observation b1320625-5011-40a3-83c6-97bb1683d67f · outbound

This paper cites Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.981955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.981955Z digest=sha256:219bc85a7e30fd7f9d0d8384384eb9de480d5ee30748af7963a6c51bc3f8dcd4

Observation 5e088145-2503-4500-af3d-d9708d3888b2 · outbound

This paper cites Qwen3.6-35B-A3B: Agentic coding power, now open to all, 2026.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Qwen3.6-35B-A3B: Agentic coding power, now open to all, 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.991054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.991054Z digest=sha256:4e08ef48674c4a8d26f698d85aaadc9e8bcc01def503cc51cf81a39f5b0b8413

Observation d1c040a0-a24c-46dd-bfd9-f94d9c9b30a2 · outbound

This paper cites Qwen3.6-27B: Flagship-level coding in a 27B dense model, 2026.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Qwen3.6-27B: Flagship-level coding in a 27B dense model, 2026

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.001766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.001766Z digest=sha256:94085b1a1d4caee93a3015a1f30bee1c9bb591af56e71480603a3f815d25e47b

Observation 9643093e-bfed-4251-9083-b15b50c20ec5 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.012837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.012837Z digest=sha256:f4d2706576f72dbadb0fccb827f3fc7c51b07108dd0e79422015263d8db40643

Observation 504f917d-82c7-4a01-a168-bbe21f75d59e · outbound

This paper cites Monet: Reasoning in latent visual space beyond images and language.arXiv preprint arXiv:2511.21395, 2025.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Monet: Reasoning in latent visual space beyond images and language.arXiv preprint arXiv:2511.21395, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.021800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.021800Z digest=sha256:38bca78bc21579acf72398c8f12a2b6c6c1c72512a940ddb091a8377d7463a69

Observation cdb0bcbb-9b50-40f2-9e0a-9b548a1b3769 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.031978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.031978Z digest=sha256:f646765ec5a74d21b5cabe11f5562cf978b53bb36ce2dded06f9b6895657d853

Observation dd8e0b6b-1218-4c10-b85e-65931206b4b5 · outbound

This paper cites VaTeX: A large-scale, high- quality multilingual dataset for video-and-language research.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning VaTeX: A large-scale, high- quality multilingual dataset for video-and-language research

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.045854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.045854Z digest=sha256:39a22f5debd9c4f638a0f7e28310ac37bb37af9d43832bf4e633d3f03c12b2d9

Observation 832921ac-748d-4087-8db4-f2144980eff4 · outbound

This paper cites Intern- Video2: Scaling foundation models for multimodal video understanding.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Intern- Video2: Scaling foundation models for multimodal video understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.070937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.070937Z digest=sha256:438478cdc475dda461422d7675de99259cfbe7120366d18f61e33494757e54ec

Observation e5dae559-0b17-4e16-a728-7434a8665c0f · outbound

This paper cites LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.086437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.086437Z digest=sha256:fcba8efd0a08ec1f11284041c536e319cb7bd8936297f6c4fa62e4c9187b5ebb

Observation a7a12be9-d250-4a05-8de4-ccfca55f846b · outbound

This paper cites MiMo-VL Technical Report.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning MiMo-VL Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.095639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.095639Z digest=sha256:599584b2f35770e0f2d4292c33cbdc4188fa8c45572b4fa03f07a444a50e91bd

Observation 75d6403e-93ba-456f-aa11-c6fd1c55050f · outbound

This paper cites LumosX: Relate any identities with their attributes for personalized video generation.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning LumosX: Relate any identities with their attributes for personalized video generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.106010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.106010Z digest=sha256:2b0c761b554358a99c17126707e8ce513fd939d881da61fa7cd6b57a6ee10bca

Observation ebe59067-2732-403c-a53a-3905d0203d1b · outbound

This paper cites MSR-VTT: A large video description dataset for bridging video and lan- guage.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning MSR-VTT: A large video description dataset for bridging video and lan- guage

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.115431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.115431Z digest=sha256:b9ff08a2174e62241e3ce6e84c9556dfcc1b1ccb5ebcd403434ef6aef6c6fcc2

Observation eff35ee6-1a97-4754-850f-c57dd0aea30f · outbound

This paper cites Progress-aware video frame captioning.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Progress-aware video frame captioning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.131232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.131232Z digest=sha256:3314bb11510fa6cf9365e5116e30cbd1727f7aa905f11bad4c176d87e4d4c044

Observation e7d57d85-9cea-4001-bb59-f7aac00a37b4 · outbound

This paper cites Kwai Keye-VL 1.5 Technical Report.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Kwai Keye-VL 1.5 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.137192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.137192Z digest=sha256:709a27b2e4806046c38f33ff0e1173dd05d3318e83caac334af9bc37faaf8143

Observation 1fdba910-7052-43bc-a0ce-292612471ffe · outbound

This paper cites OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.144127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.144127Z digest=sha256:4f618347e477c913900d145c873827a9fffe1c1f9dbc1b834a4ac1c5314d49e4

Observation 58e3c8b7-1217-441c-92a7-d2575c5889ea · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.154107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.154107Z digest=sha256:de67966072b54dc5ca19802df8474ab74dda459ca3a118e74f994ce7fd982060

Observation 26dda30b-25a9-4516-8989-979dba28bf11 · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.160261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.160261Z digest=sha256:ed51c6e1f790e75576fba09c0deb80738be93db8a0c33ffc70599f76be56a2a1

Observation 9b1d58b1-77ce-4a71-b977-9a7e07eb5853 · outbound

This paper cites VCapsBench: A large-scale fine-grained benchmark for video caption quality evaluation.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning VCapsBench: A large-scale fine-grained benchmark for video caption quality evaluation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.170380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.170380Z digest=sha256:f0420e9ed2bf7393014077159c18615c8377ab684fc4af2e02a3c21f8cb48506

Observation 16ef230b-d319-41b9-8b12-9628295b96e0 · outbound

This paper cites MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.178883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.178883Z digest=sha256:622940c3d5023b15740483a655a4709ace456dece9b38e6e576840e5b8570ef0

Observation 9799e92a-b957-4f5b-84e3-496b094fb6fb · outbound

This paper cites Debiasing multimodal large language models via penal- ization of language priors.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Debiasing multimodal large language models via penal- ization of language priors

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.185599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.185599Z digest=sha256:1f17019d40c29e2b5d781a10901d3973c03af18b16664b462c8a3d4ce814ef55

Observation 6e3b8fe5-73e7-4103-9538-2a8d605865c7 · outbound

This paper cites from <Image_N>.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning from <Image_N>

Reference 43

Resolution
malformed identifier
no resolver link, observed 2026-07-31T05:08:20.197862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.197862Z digest=sha256:f08e2a693ed000d17fa7bcbc7b9907a13c1c33546b1b313c0da15374fed62715

Pith citing papers

No inbound Pith citation observations are available.