Pith. sign in

Paper Citation Record · LEDGER

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search

As of 7 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2506.11155.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11155 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:45:51.654456Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1c60edfc-0cbb-4f35-89df-1696c8185719 · outbound

This paper cites The Verifier we used is Qwen2-VL-72B.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search The Verifier we used is Qwen2-VL-72B

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.549246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:49.732140Z digest=sha256:7c40cf55c271749c646d3938076c4aeccf9c5c1f1f6b8dddbd74facfada74117

Observation d0cb6278-ba63-4b33-bad5-f1403c087f6a · outbound

This paper cites video caption.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search video caption

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.536479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:49.769808Z digest=sha256:277c1ae0ffe381276f95a96b30beb1be86cd5478137c6fdc41812a36f2216d69

Observation 08bb520d-f4a2-4e37-8030-ace8a2b4984e · outbound

This paper cites The defination and examples of each category is described in Figure 10.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search The defination and examples of each category is described in Figure 10

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.494812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:49.897453Z digest=sha256:bf2d3868d94660c3572ff4ac5cc65a4d6eff961f5a301f58734296cd48056c0c

Observation 43c230d3-60b0-4dda-8539-13f547e19a14 · outbound

This paper cites As shown in Figure 6, the Nature and Wildlife category has the highest number of videos, reaching 382, while the Arts and Creativity category has the fewest, with 57 videos.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search As shown in Figure 6, the Nature and Wildlife category has the highest number of videos, reaching 382, while the Arts and Creativity category has the fewest, with 57 videos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.523671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:49.807483Z digest=sha256:800a84f827d484ae311f8163b756de1e0e2cb248124a9536f962ddffa60d774c

Observation 0edb9973-083d-4661-9ce8-bb956e41b22e · outbound

This paper cites For clarity, consider these examples: ## Example 1 ### Key Points {1.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search For clarity, consider these examples: ## Example 1 ### Key Points {1

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.128104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:51.250419Z digest=sha256:fabb706fefcadb7aa8a8bf8a1a90ceab0a92f6731eda54d5faff00992f0ba847

Observation aca6c052-7370-41b5-9be1-b61b192335c7 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.655855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.655855Z digest=sha256:58120f1546f209852ac9289667d229f3fdaac5846342b1abfe4930320f93aad6

Observation e43c75bc-cef8-466e-b1e4-d6c27ddc085b · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.510059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:49.877197Z digest=sha256:6ca689341a07530666a760b826662a1365dd6574996366e68e670b691d745841

Observation 9128212f-1432-4a17-a791-81d3712fea5c · outbound

This paper cites Elaborate on the visual and narrative elements of the video in detail.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Elaborate on the visual and narrative elements of the video in detail

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.482137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:49.953882Z digest=sha256:bbeea5294dc9d5f0d69c9b61279cf9e51a8ee0b2f3cc6010085de6dda100a736

Observation 09dc07ac-d317-4bee-b07f-d469fb158832 · outbound

This paper cites Reply to me with a precise yet detailed re- sponse.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Reply to me with a precise yet detailed re- sponse

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.468746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:49.986135Z digest=sha256:21b0847868680220231edc6e030c58e6dcab415fe3eae8fb1d8b0d4258c0a995

Observation aa5f7445-cd02-4a89-ade0-b30c3b04a451 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.455747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.040106Z digest=sha256:3da4694d84c6ee41f0d43099f9f4bd352ba7895f4277fdd7a409812f470b26e6

Observation 12594294-b8ea-45a0-8d77-a4ea5371fbf2 · outbound

This paper cites If you are not sure about something, do not include it in you response.\n # Task\n Describe the background, characters and the actions in the provided video.\n”.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search If you are not sure about something, do not include it in you response.\n # Task\n Describe the background, characters and the actions in the provided video.\n”

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.442836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.063737Z digest=sha256:7f5f6c8f732b1868dcdc1fcaeaa11cd7f19a2731f28df2be06654978ca3df554

Observation 09a987e6-3a93-4b05-93d3-bd4e4708cb89 · outbound

This paper cites Please describe the video in detail.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Please describe the video in detail

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.429147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.091040Z digest=sha256:11d32d0ad0b1ccb28fa1425fb1b7364a6b3851884c12b899fa2b37a369ebc3e8

Observation 612c6a12-7820-4ec1-8ee6-5fb423a35600 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.416554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.113506Z digest=sha256:9d14acbe186d24eb8d52eccce1bf6a562020e95713f83f49ca778d8a362fc376

Observation 8ff35304-bb11-4a01-a8a6-9a32d9b4da57 · outbound

This paper cites Action Description Action description focuses on the specific behavior or activity that takes place in the video.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Action Description Action description focuses on the specific behavior or activity that takes place in the video

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.404105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.157246Z digest=sha256:8e764760fc6fbd3dc5f6f96df43b0d804518afbb59b1c9415bc533a01c97c558

Observation 5e35b046-e14c-4907-a93c-8cf020039730 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.391772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.227646Z digest=sha256:9630f1e7fd7bd92fa9654ea75a7d8f9521991f4d03d8bb0c17fca3da35bea65e

Observation f6407169-11de-4a84-b3c8-456171823de0 · outbound

This paper cites Environment DescriptionEnvironment description covers the background and environmental features of the scene in the video.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Environment DescriptionEnvironment description covers the background and environmental features of the scene in the video

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.380039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.272752Z digest=sha256:f9afb4aece92cdcf0b4faca183fd344dea688b36cf5ed5f2f8f9609d846b0bbb

Observation 0a7688bf-5e33-47dc-8ebd-8049fa3c40e8 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.367175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.321552Z digest=sha256:696aafb714e6ef8d9a220e0407cb0ae1d28efa4fedd962bd4113ddeeb8a2b0db

Observation 1c4f6e0d-c3a7-4497-a5f7-082058f8b5ac · outbound

This paper cites Object Description Describes the features, characteristics, or details of inanimate objects or items present in the scene.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Object Description Describes the features, characteristics, or details of inanimate objects or items present in the scene

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.353915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.384862Z digest=sha256:2b3077b8f4f5fb416da729bc82059f2cac1edde9725e9624380f4f703377d180

Observation 36aad59d-79c6-4bef-9b3c-afc83c6f062f · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.340986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.433451Z digest=sha256:842a8e3e58faf2bb6b09bc7e0b3ef17b48e1bb622ea05e62f9e6afe14eb1ce20

Observation e9a41f06-b236-4622-b32f-2204c25f501c · outbound

This paper cites Camera Movement Describes the camera angles, movements, framing, or other cinematographic techniques used in the scene.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Camera Movement Describes the camera angles, movements, framing, or other cinematographic techniques used in the scene

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.327006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.493367Z digest=sha256:e2705a41d09cc604288673432a00d8bb924c4c15cf0b6d5599f0c990ce58b7e0

Observation 6d106cb3-23e4-4936-81ab-e38ff7579a41 · outbound

This paper cites Overall Description.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Overall Description

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.312951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.556622Z digest=sha256:de9e980ef808ff520d491c3455c03b333584416d6db6406f3301671f21e72632

Observation e2c286b5-8623-4035-8ed6-1564be95773f · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.299776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.602919Z digest=sha256:d3d9943d5d7695d54cd9e523fe8543a24bd2d10cf8ad3c7757118e03f899fb95

Observation 44daad0e-240f-45f7-a52a-fefdb635de9b · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.287434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.648467Z digest=sha256:a556f3b23e8d92c8342d6fdd8ef2ada9305bb4364e06022fb53f3d255b8da4b9

Observation 207efaf8-4853-479b-b148-e0c90cf7d968 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.273193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.685906Z digest=sha256:949c6c49c868cdafdc12da8e610bfbe1c30265095938d4afc481a5b1fa22ec15

Observation 5a01bb6c-ea31-42a4-b312-5f45cfd4fdf3 · outbound

This paper cites Judgment: [yes/no] Reason: [Brief explanation].

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Judgment: [yes/no] Reason: [Brief explanation]

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.259294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.709361Z digest=sha256:2af400522424e82374d2372881b9367b2d570243c0c86295493bcfd4a888ee48

Observation e7a52ff5-1705-41f0-bb08-4f714e14fd67 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.246129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.751336Z digest=sha256:c2955d0f22ed344423785427433ccbc7dbe4c4638bf6f59abdf834bc75c97cbd

Observation 73134215-1d39-4897-8815-04b2c4f43b6b · outbound

This paper cites For clarity, consider these examples: ## Example 1 ### Video Description: The video showcases.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search For clarity, consider these examples: ## Example 1 ### Video Description: The video showcases

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.233155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.820854Z digest=sha256:d6224122bd7ddacbb548577b3e42876630244c51ee86b8fed0740dc69a6824af

Observation 44eec224-6230-442b-a74a-18647b9d08b2 · outbound

This paper cites entailment.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search entailment

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.219323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.866840Z digest=sha256:a60eed40393480c9b3957b45bdb7f60b4b258fb544e18d693ce133c964bed950

Observation e3a218d6-bbb0-476d-936c-c61e9096d38d · outbound

This paper cites contradiction.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search contradiction

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.204952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.920062Z digest=sha256:44be990e8d2cf51c7536089d217d73da59571358fe10ef0dd394e9d31afdad29

Observation deaf138a-b65f-4a5b-bfd3-379f268f0095 · outbound

This paper cites neutral.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search neutral

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.191475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:50.988196Z digest=sha256:32b36722e3e261cd841b009df52e15e27cda7af36ed374230a1ab5676e857f63

Observation f04dc7e4-f631-4893-888d-c690782b4841 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.179059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:51.039572Z digest=sha256:73c7e51e7f253af9bca895d3706fb5687ca7115acc1fc95fa73af18738c7cbb3

Observation 2dead321-8060-4d48-b8fe-099d903d3c0e · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.166708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:51.135838Z digest=sha256:dd760e615e6e99bad3762394e84a26dca58c81bc6d775001822c04253696b9e5

Observation bc4c0a8c-5d0a-4cda-b690-413ff0c6f458 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.153967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:51.205648Z digest=sha256:45e2fcea8053c8f80fff0ce91cb49688157e888632f3ef4c4b671622436c583a

Observation 6f322406-4ade-4809-85c7-01a3c1c5245d · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.141384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:51.226275Z digest=sha256:9c06edd4b917b8374e429dcbbb9e35a1706a0bf56c4e7b16a9e913e7154eba10

Observation 209f6f20-0529-436b-bd45-60bbf283b7a3 · outbound

This paper cites [No] Speculative.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search [No] Speculative

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.114620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:51.286328Z digest=sha256:04bbe3a3988ef8a657375a494b9b8116583bc35c5e5bc6afd0282ac69caa5e4c

Observation 4e58d046-f8ca-40fd-94a3-9631b7ad3b35 · outbound

This paper cites </thought> tags.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search </thought> tags

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.101011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:51.329282Z digest=sha256:9cd59db3c31ca7af3ccf4ab456cd01b941c78a491a1b11f46dc9652d44bdc41d

Observation 440f3780-deb9-417e-923a-65bd5d676425 · outbound

This paper cites Overall Description.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Overall Description

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.086183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:51.449655Z digest=sha256:ae69047c52e158ec355d8b6f7e0ca1b65504aa2962b8f28f2faec25ced626ba7

Observation ae45ca9a-7395-467c-9248-59acb4d784ae · outbound

This paper cites ## Example Input and Output: ### Input: Overall Description: {overall_description} Observation: Vehicles Key Point: There is no existence of any vehicles in the video.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search ## Example Input and Output: ### Input: Overall Description: {overall_description} Observation: Vehicles Key Point: There is no existence of any vehicles in the video

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.072730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:51.496957Z digest=sha256:a99d78f3fa7675e15325f4eaea61231e427c8a8d60cedc132a839a5ad40f330c

Observation ba2e7ca2-4f59-4f63-a9b4-b4249dfbf04d · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.059285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:51.561541Z digest=sha256:b344f6ee53654ad3233ad5bab9a4e903902b363a67319d947d88d9bd6da7082a

Observation 943bbb7a-b0a6-4ed1-97dc-7efff2e687fe · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.045858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:51.625504Z digest=sha256:666cb8ed3ee3eeb05c2ee7338d533e160f65044b5848ca80bc98febfaa5c2b18

Observation 5cd90b8e-b737-4cb4-9a01-69db871d2438 · outbound

This paper cites Precision / Recall / F1 Score.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Precision / Recall / F1 Score

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:51.968280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:51.654456Z digest=sha256:c766b093197dfe50b0ab5be8cfc68afcecf2fcfe3d1c4e16aee7628a1fe69d8f

Observation 91c77bd8-51be-499e-a1b7-18721181226a · outbound

This paper cites CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language Models.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language Models

Reference 2013

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:45:51.840389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:45:49.548866Z digest=sha256:148186bef7a378f9f3e256a890df3a3259257b27dcf774995e55d15d42f37c7f

Observation 8cc72c16-a1b7-44d9-ba64-53bcd81e4d9c · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Reasoning with Language Model is Planning with World Model

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.431685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.431685Z digest=sha256:b0f9784e7b7ec0e0675cd12595045a6712423478947ae275614c3e63188f6a4b

Observation 8b0c2260-b588-4876-b0f3-809ab39e684f · outbound

This paper cites A Survey on Data Augmentation in Large Model Era.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search A Survey on Data Augmentation in Large Model Era

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.695413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.695413Z digest=sha256:db376799889d87a2cb5f2b47be5e4b45e1406beff9f71f858529e83395062750

Observation d31a7ef3-db04-4ac8-b77b-04a581e6beb3 · outbound

This paper cites AugGPT: Leveraging ChatGPT for Text Data Augmentation.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search AugGPT: Leveraging ChatGPT for Text Data Augmentation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.338985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.338985Z digest=sha256:aa30312b67487999d4564ebf9cac38c3a7e9cbad6ea5095fd0ba59cf3fa87bbc

Observation ad3cbeba-4fda-4f83-bb7d-6c5ec7d0dd94 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.619619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.619619Z digest=sha256:e420d0ad444091eab4dcc1556a16a05e36c017badb043b30f3932cd174ee328e

Observation 0b957b4d-0f0a-4132-b5a7-232ce234c2fb · outbound

This paper cites ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.470142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.470142Z digest=sha256:aef9d81ce25e253f97ced6fa02512c9fe0d0867e8b01a334fd8f29b7fb78e6f7

Pith citing papers

No inbound Pith citation observations are available.