Pith. sign in

Paper Citation Record · LEDGER

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing

As of 5 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2605.03276.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.03276 v2

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T01:43:34.898639Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact19
  • verified fuzzy27
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36c2242c-c6f4-493b-9fc5-c85c78507edc · outbound

This paper cites Qwen2.5-VL Technical Report.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:13.946366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:79b31f03e21fee296f42d046c49422ad98b76acc5e25a511827a557544a0af48

Observation 046bbe7a-7f19-48b8-99b0-d3d23e3ab290 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:3c6e3021f33696fccb4030f3e2f4751a45d14984ce236b0d922fb9af8ef2cf14

Observation 9766c89d-524a-4dbe-9045-7ceadbe02802 · outbound

This paper cites Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:40:56.172131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:daddacd899ec54436cb592cea8c4fdd44352d1d8edbd56020baa1455af2d587b

Observation 181d5190-dbb2-436b-afbf-31755d7a4053 · outbound

This paper cites Routledge.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Routledge

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.010414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:edb536e43f3d75f243a6699ed2a5aa5e2679ef96dbfcfcdee31c6a51fdc5fdfa

Observation 6a2909b4-03d3-4426-b70c-775ab0f8442e · outbound

This paper cites Motion-grounded video reasoning: Understanding and perceiving motion at pixel level.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Motion-grounded video reasoning: Understanding and perceiving motion at pixel level

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.007343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:5e00af8346f435af2846eec5d7080ba1081460c5fa1b40793f5338a7fb4f1ce3

Observation 49f06259-5d2f-4df0-8174-49fdfd75bff5 · outbound

This paper cites arXiv preprint arXiv:2510.08559 , year=.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing arXiv preprint arXiv:2510.08559 , year=

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:13.959879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:578e75dadcc7f812b96618ba7be39abf2600820ba5f79c786aad78d7beece225

Observation ba6fc461-36a2-49dd-9d33-13637a3d96ff · outbound

This paper cites Gemini models: Gemini 2.5 pro.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Gemini models: Gemini 2.5 pro

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.020696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:6e764d8f18b4062d893392ea0b5e3dfc48a9468f14f55795a869775bac318583

Observation aaf8f905-9c22-4d1f-8a28-cf0f8565531f · outbound

This paper cites 2, 4, 6, 7, 8, 12, 17, 18, 19.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing 2, 4, 6, 7, 8, 12, 17, 18, 19

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.043630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:cfe86966a01ea359bbcbd6246454d72c1f1dc7354f324176b8b5f1ae5fc94a94

Observation 2e1f3bec-4f6d-4347-98b7-d978410731ea · outbound

This paper cites Rout- ledge.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Rout- ledge

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.041203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:d53ca3383d927c7bb3d8fe07da84fa83f078cba6d9b09cdbedff69b9aa447dd9

Observation 56898397-9ff3-4fbc-8501-74f686ab5c1a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:13.969368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:63e7cb86993dc9623a01fdadef80798cfeb9f23042cfc52bd967d83e9a91d15c

Observation ddcbbf63-5dd0-4327-a7c6-54e35000ab50 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.955297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:fd614bfe14593fff06fe9819f0aeb19625efeeb3d24f522bd222ade9ed6f7c98

Observation 1eef172e-fe64-428e-8567-954c569f4741 · outbound

This paper cites Edit3K: Universal Representation Learning for Video Editing Components.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Edit3K: Universal Representation Learning for Video Editing Components

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:14.009776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:932eec0488bd7c69cffd80423e5f15557a7673de9402192dd3b98853fb308365

Observation 2080c02b-a044-42fb-ae7b-427c50401297 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Activitynet: A large-scale video benchmark for human activity understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.003915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:a6045076af0bad57ca175948698c1cbd1e3470f82b19c76e63f9399236db1e2b

Observation ecd96d5f-e3e1-4aee-a47e-1c5169d31208 · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:77599bc6d02c8ba8105b8704ae84bfdfdc7f0d76eeee5b0d515b0dc52697836d

Observation 17c7c706-eea0-4983-9498-7d7dbfba9c59 · outbound

This paper cites B-script: Transcript-based b- roll video editing with recommendations.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing B-script: Transcript-based b- roll video editing with recommendations

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:24.995907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:e6182e9a91e462e3f9700ea80a5b992b1a7782ca292502a05224686c6192b478

Observation a0735e0e-082a-4940-b3d7-7480a72c8923 · outbound

This paper cites Routledge.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Routledge

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.032927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:188366c8c1eb28a4789095baeda1caef41412d2d0df61166f09e4cab49c2d9f2

Observation 1279e9d8-e6c8-4d53-b329-a2fcc9510f0c · outbound

This paper cites The Kinetics Human Action Video Dataset.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing The Kinetics Human Action Video Dataset

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:14.006405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:7d5efbc95906ef6c17627944a9fc25204e5421460daa9777e7e3dae96464b716

Observation 52f57a4c-073a-46a2-8636-465ae3e7ae87 · outbound

This paper cites Veu-bench: Towards comprehensive under- standing of video editing.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Veu-bench: Towards comprehensive under- standing of video editing

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.026642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:34f47ecfd5e3ba6233d04792597785335a24e408451121238919ab2ea84a5be4

Observation d971314e-1153-44c3-9631-166148f9a111 · outbound

This paper cites Omnivideobench: Towards audio-visual understanding evaluation for omni mllms.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Omnivideobench: Towards audio-visual understanding evaluation for omni mllms

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:13.993497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:c60268a505f7bb812430a7e7769f5fe36fc217132983901613348628341d6ec5

Observation c05576d7-0c16-485c-820f-9b631f10e9f6 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:13.996824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:e3724ab9114d539aedf7c3ebf157e8a496e0ee7e8884139f55715c6f427cefcf

Observation a8c54e84-a4df-4cca-897c-47b0584dc843 · outbound

This paper cites From representa- tion to reasoning: Towards both evidence and commonsense reasoning for video question-answering.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing From representa- tion to reasoning: Towards both evidence and commonsense reasoning for video question-answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.029893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:1b13fa0602bfd14017f841f33e92f7a26b0be726d34064e36800944629d4289b

Observation 7288ce7f-15ce-4d77-8d38-0c9ba640ea99 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.074995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:4e53a521db49fd2717fcf61002b4a51b84b99d1cf4c91615cdc6aa6b6c4a991e

Observation e2f2234b-8652-4a36-b7e1-3feb5e406316 · outbound

This paper cites Shotbench: Expert-level cinematic understanding in vision-language models.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Shotbench: Expert-level cinematic understanding in vision-language models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:13.982301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:31ff6a3b253750e1607579abb2be7e88d3f20ed309854ae367f4651c322f4a8a

Observation 54dfbac0-7998-4141-9c1f-9cebbd314ab0 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.035988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:08e115f5251b8d6148c93345331ebe2c63c36d74ce72c5fad91eaf0a36c2e007

Observation 3559ed95-7c5e-44b8-bd63-e7504db47c62 · outbound

This paper cites University of Chicago Press.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing University of Chicago Press

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.038646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:0f0e310c58be39a14bba097c613b79ae11ea98d76c368112a91b9f4cbbd1f252

Observation 32806bd2-610e-4563-ab90-eccaea93760b · outbound

This paper cites Gpt-4o: Openai’s newest multimodal model.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Gpt-4o: Openai’s newest multimodal model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.017286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:76f897faa38801354c8282d0a3bfa0e5d481bb75091f083922679a935824cc97

Observation 1a1d5f99-c8eb-4af1-b18c-a8005a997050 · outbound

This paper cites Perazzi, J.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Perazzi, J

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.023636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:f3504600e35ffca4f988c07d3531b7fa31089f3f1613e3ddc5ed4dac69ade650

Observation eb9960d6-39dc-45a2-9a1d-54c5718145b7 · outbound

This paper cites VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:13.974098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:8100b17bf5543541d1d1e42e3abb3bf49038ed96df9ee46fbf695a57eb822a12

Observation 9f6ddd63-1eef-4b50-ad80-bc0c7dc857cb · outbound

This paper cites Qwen3-vl.https : / / github.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Qwen3-vl.https : / / github

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:24.987974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:dcd6044dda9224269180f44da90adda141a550b9a316e147d1f0e867ee69b7b3

Observation 19345d41-4596-4226-96ae-3d7696e51159 · outbound

This paper cites Sage Publica- tions.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Sage Publica- tions

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.055466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:b7b2ed064baa8a0dcefe03a1532b577efe17cfab490ca56b1b4a8d4151fdb114

Observation e8833042-c6e1-44b2-a39e-8742d9f4b0c8 · outbound

This paper cites Video-mmlu: A massive multi- discipline lecture understanding benchmark.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Video-mmlu: A massive multi- discipline lecture understanding benchmark

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.062041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:b06c17ec9794283f9b93631460bb7666fdc8d4c26753712784dc296053640aaa

Observation 39a509f4-61e7-4cc0-b868-bfae9ddd69a0 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:13.964294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:e4a5f23b29b8fbdd33baf7b7d60e161bfc8833aaa7046966c14e813834cb783c

Observation 3cb3f822-1bb4-48fa-99e6-ac825bf317a5 · outbound

This paper cites Qwen3 technical report.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Qwen3 technical report

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.058853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:fb387be874ee3c2e466b956734176479e7d9d567b07221309df7e7458851a4ee

Observation ee453567-c7e2-470a-b981-c97471120e93 · outbound

This paper cites Research on the application of nonlinear edit- ing technology in vlog production.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Research on the application of nonlinear edit- ing technology in vlog production

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.068441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:12a57723eef92ef151f108872cda9281a339245253131e48c40963df8bdaa18d

Observation 3df129eb-f5ae-4e41-afde-f1e7fb8bc5d7 · outbound

This paper cites CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:13.978283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:8f07065786e876cd5c9018f3239f9b1d0aa7138ae239556e9a166d9020a2d217

Observation e7890ddb-aaee-42f8-8a92-00b72e438295 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.046624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:fe55913ac106923b7b9f53eb74e84b7038d1ba9b4f1caad42f90d7779787c6a2

Observation 323eaa52-a13e-4814-a781-90ad3a54291c · outbound

This paper cites Beyond raw videos: Understanding edited videos with large multimodal model.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Beyond raw videos: Understanding edited videos with large multimodal model

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.071461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:8685dfad2db5335d23fc7fad45754a5476107e8e41b7a746ee669a0a3aaa75b7

Observation 19766d70-a65c-4db4-90d9-718cd387a5f0 · outbound

This paper cites YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:13.949720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:fb564a01e579422f2cefc762f98820592a6f5967d48126228ba885dc6c08a6e6

Observation 1ed5d487-9e1b-4e07-a2fc-e37cd6298617 · outbound

This paper cites Qwen2.5 Technical Report.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Qwen2.5 Technical Report

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:13.985857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:d29243f667736d32f26eac2cb29f24905b0b1debae0348b4933414ac53f6e65b

Observation d890b3e6-1d6d-45fe-b98f-6abcdd6cfe48 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:13.989379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:549a51dce2d1e30c93f6b62fa229f257d71cf47cae63bd6a41baee100d137f6e

Observation 42e9010d-71af-4681-9788-50b528eb2db4 · outbound

This paper cites Long Context Transfer from Language to Vision.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Long Context Transfer from Language to Vision

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:d936bf15579dec2f0d43790689d7f53ef60d5d57dc5d3025aac37e745179b80d

Observation 3014223d-0f95-496e-8e00-a295897c150a · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Towards automatic learning of procedures from web instructional videos

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.013671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:47941aa6605d6feffd30744417ed4d9d10f2ea95b4f5d476db420eeb5a81890c

Observation c6394ad0-b7b2-4769-bb38-c383566883f1 · outbound

This paper cites When does the{CUT TYPE}occur in the video?.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing When does the{CUT TYPE}occur in the video?

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:24.999781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:87dd54a3aa548727b7ec09a55dc2ab44caf46e6ac9bac3538760878ddda633ef

Observation 44a6e817-0593-4731-ad88-62052e2024d2 · outbound

This paper cites high", "medium.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing high", "medium

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.049100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:3f518fe6a1efc1d21d680709ac431722d40820c476ba7e1584b7157be9f15256

Observation d5b16fa2-6ad3-4e0f-8ac8-67c724b9861b · outbound

This paper cites Yes"/"No.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Yes"/"No

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.064993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:011472129cccb0a1d48f8ada2af23d257285fdf974b756953dd35b474d504840

Observation 331a5521-99df-4ee9-be25-1d0b1ec7116e · outbound

This paper cites an unresolved cited work.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-14T06:08:24.992143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:369eda2adf256da3a835395f6063784a1be27d1aa26841ec379804e655576191

Observation 7df02a14-7bf2-48c1-aba5-0d5f5b8af252 · outbound

This paper cites timestamp_start.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing timestamp_start

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:08:25.051897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:cade3b39e0d64a35c657d136e832b6cbea6ef74a9d8aa63701f3b5de0407bd47

Pith citing papers

No inbound Pith citation observations are available.