Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:45:51.654456Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2506.11155.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:45:51.654456Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1c60edfc-0cbb-4f35-89df-1696c8185719 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search The Verifier we used is Qwen2-VL-72B
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0cb6278-ba63-4b33-bad5-f1403c087f6a · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search video caption
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08bb520d-f4a2-4e37-8030-ace8a2b4984e · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search The defination and examples of each category is described in Figure 10
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43c230d3-60b0-4dda-8539-13f547e19a14 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search As shown in Figure 6, the Nature and Wildlife category has the highest number of videos, reaching 382, while the Arts and Creativity category has the fewest, with 57 videos
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0edb9973-083d-4661-9ce8-bb956e41b22e · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search For clarity, consider these examples: ## Example 1 ### Key Points {1
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aca6c052-7370-41b5-9be1-b61b192335c7 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e43c75bc-cef8-466e-b1e4-d6c27ddc085b · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9128212f-1432-4a17-a791-81d3712fea5c · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Elaborate on the visual and narrative elements of the video in detail
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 09dc07ac-d317-4bee-b07f-d469fb158832 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Reply to me with a precise yet detailed re- sponse
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa5f7445-cd02-4a89-ade0-b30c3b04a451 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12594294-b8ea-45a0-8d77-a4ea5371fbf2 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search If you are not sure about something, do not include it in you response.\n # Task\n Describe the background, characters and the actions in the provided video.\n”
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 09a987e6-3a93-4b05-93d3-bd4e4708cb89 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Please describe the video in detail
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 612c6a12-7820-4ec1-8ee6-5fb423a35600 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ff35304-bb11-4a01-a8a6-9a32d9b4da57 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Action Description Action description focuses on the specific behavior or activity that takes place in the video
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e35b046-e14c-4907-a93c-8cf020039730 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6407169-11de-4a84-b3c8-456171823de0 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Environment DescriptionEnvironment description covers the background and environmental features of the scene in the video
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a7688bf-5e33-47dc-8ebd-8049fa3c40e8 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c4f6e0d-c3a7-4497-a5f7-082058f8b5ac · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Object Description Describes the features, characteristics, or details of inanimate objects or items present in the scene
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36aad59d-79c6-4bef-9b3c-afc83c6f062f · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9a41f06-b236-4622-b32f-2204c25f501c · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Camera Movement Describes the camera angles, movements, framing, or other cinematographic techniques used in the scene
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d106cb3-23e4-4936-81ab-e38ff7579a41 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Overall Description
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2c286b5-8623-4035-8ed6-1564be95773f · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44daad0e-240f-45f7-a52a-fefdb635de9b · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 207efaf8-4853-479b-b148-e0c90cf7d968 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a01bb6c-ea31-42a4-b312-5f45cfd4fdf3 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Judgment: [yes/no] Reason: [Brief explanation]
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e7a52ff5-1705-41f0-bb08-4f714e14fd67 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73134215-1d39-4897-8815-04b2c4f43b6b · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search For clarity, consider these examples: ## Example 1 ### Video Description: The video showcases
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44eec224-6230-442b-a74a-18647b9d08b2 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search entailment
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3a218d6-bbb0-476d-936c-c61e9096d38d · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search contradiction
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation deaf138a-b65f-4a5b-bfd3-379f268f0095 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search neutral
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f04dc7e4-f631-4893-888d-c690782b4841 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2dead321-8060-4d48-b8fe-099d903d3c0e · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc4c0a8c-5d0a-4cda-b690-413ff0c6f458 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6f322406-4ade-4809-85c7-01a3c1c5245d · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 209f6f20-0529-436b-bd45-60bbf283b7a3 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search [No] Speculative
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e58d046-f8ca-40fd-94a3-9631b7ad3b35 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search </thought> tags
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 440f3780-deb9-417e-923a-65bd5d676425 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Overall Description
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae45ca9a-7395-467c-9248-59acb4d784ae · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search ## Example Input and Output: ### Input: Overall Description: {overall_description} Observation: Vehicles Key Point: There is no existence of any vehicles in the video
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba2e7ca2-4f59-4f63-a9b4-b4249dfbf04d · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 943bbb7a-b0a6-4ed1-97dc-7efff2e687fe · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5cd90b8e-b737-4cb4-9a01-69db871d2438 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Precision / Recall / F1 Score
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91c77bd8-51be-499e-a1b7-18721181226a · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language Models
Reference 2013
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8cc72c16-a1b7-44d9-ba64-53bcd81e4d9c · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Reasoning with Language Model is Planning with World Model
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b0c2260-b588-4876-b0f3-809ab39e684f · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search A Survey on Data Augmentation in Large Model Era
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d31a7ef3-db04-4ac8-b77b-04a581e6beb3 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search AugGPT: Leveraging ChatGPT for Text Data Augmentation
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad3cbeba-4fda-4f83-bb7d-6c5ec7d0dd94 · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b957b4d-0f0a-4132-b5a7-232ce234c2fb · outbound
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.