Pith. sign in

Paper Citation Record · LEDGER

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs

As of 8 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2505.19535.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19535 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:48.995779Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy55
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 49259069-4357-497b-af56-bb497360a7ed · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.979116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:43.651761Z digest=sha256:57c4c428fafdce001bc419d6adc5703f4374c422171ef7add7d3bae60e307075

Observation d726bc0c-741b-47db-9faa-edd19846277f · outbound

This paper cites Tokenflow: Unified image tokenizer for multimodal understanding and generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Tokenflow: Unified image tokenizer for multimodal understanding and generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.676070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:43.742055Z digest=sha256:59e9e5b682ea7150433c78f06fff6bb95c7a4d5699808a025311a2171f54cb99

Observation 2f1b4a46-ee54-44d6-9c4c-13ed27702cc6 · outbound

This paper cites Text2video-zero: Text-to-image diffusion models are zero-shot video generators,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Text2video-zero: Text-to-image diffusion models are zero-shot video generators,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.429716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:43.841821Z digest=sha256:81258cd079124e6c97265d9a60e62f0ad9e975ce098a38ca799e56fa4bf262a3

Observation 60d1a271-6100-4bc2-8999-e9635db4082b · outbound

This paper cites Ccedit: Creative and controllable video editing via diffusion models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Ccedit: Creative and controllable video editing via diffusion models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.111035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:43.901242Z digest=sha256:48547664ee81f6b9bf6073210ccaebb2761e741274fd4aba8eb32a7a2c84943a

Observation fb830f27-c3ed-4215-ad0d-1187587f0dab · outbound

This paper cites Controlvideo: Training-free controllable text-to-video generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Controlvideo: Training-free controllable text-to-video generation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.919016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:44.024176Z digest=sha256:ddfa6d61308d0c8d973a851d25058c8d8f28cf20e886b515dfedd5a606d10b6c

Observation b9a50780-125f-42b2-988f-e6e5d402579e · outbound

This paper cites Fatezero: Fusing attentions for zero-shot text-based video editing,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Fatezero: Fusing attentions for zero-shot text-based video editing,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.713443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:44.195828Z digest=sha256:829f499a9e7e9448724a85ae5ddb39e0278317a0c83cb84c66116c95ef67e7cd

Observation 05c9366b-5047-46ac-8824-75e3d23a5552 · outbound

This paper cites Flatten: optical flow-guided attention for consistent text-to-video editing,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Flatten: optical flow-guided attention for consistent text-to-video editing,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.555158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:44.339750Z digest=sha256:1a60b2797de0239d3ce534276a98118ed5d42f3b56d39cf90eeb206e9df85eb3

Observation 2941ad73-01ea-4f6a-b494-4538eea50770 · outbound

This paper cites Fresco: Spatial-temporal correspondence for zero-shot video translation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Fresco: Spatial-temporal correspondence for zero-shot video translation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.301203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:44.443157Z digest=sha256:ea9bf8ef863ae643501561b116275e2d3b724bd051f2e5cdbb9e6c60225066e4

Observation 47e9861c-8a2d-4354-b017-20dbd203fab1 · outbound

This paper cites Pix2video: Video editing using image diffusion,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Pix2video: Video editing using image diffusion,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.154855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:44.574580Z digest=sha256:f516fd6169836ae9f1a2bb8d09e3c9106b19af615175066a248559399a3aec7e

Observation 77254bda-3784-410c-9557-33efcdda1cfb · outbound

This paper cites Rave: Randomized noise shuffling for fast and consistent video editing with diffusion models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Rave: Randomized noise shuffling for fast and consistent video editing with diffusion models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.972919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:44.666645Z digest=sha256:bfa421585a09d815615077a657feb0ca66aa57e8cacd91bb3964f3ff9857c7f5

Observation 6d9acb26-1789-4300-a8ff-33f813aa1edb · outbound

This paper cites Slicedit: Zero-shot video editing with text-to-image diffusion models using spatio-temporal slices,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Slicedit: Zero-shot video editing with text-to-image diffusion models using spatio-temporal slices,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.824036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:44.788333Z digest=sha256:43a53bffb1741430e777bf9e98113afc14e79d628bbf9b2f879ef790756d8b7e

Observation ad892360-db24-4767-b16a-1ac2c692ca94 · outbound

This paper cites Zero-shot video editing using off-the-shelf image diffusion models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Zero-shot video editing using off-the-shelf image diffusion models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.720375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:44.869181Z digest=sha256:f1715ffc6597ecce7dafeb08f7b72a47b74a636eefe094da9000bf4b4c8c2812

Observation 76e4353c-1fb6-48f1-ab9d-c5f18c8268da · outbound

This paper cites Video quality assessment: A comprehensive survey,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Video quality assessment: A comprehensive survey,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.573558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:44.964072Z digest=sha256:275314160a281bfd5eed6468eda42a61f346b0842e7abc74902e0bbc192a4735

Observation e4796152-584e-4f9a-81d4-47c9342fe8d8 · outbound

This paper cites Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.401381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:45.062144Z digest=sha256:5916b0e6530d4f26bbdf4fa7dd9aaa1328ebdaf4c22abb5a62b6c9c5196cef69

Observation a5d76651-a57b-4067-ab29-cab838afc2b4 · outbound

This paper cites Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.275262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:45.140757Z digest=sha256:90a7ce6e7717d5e3d9979b65c027fcb367682e78e97a4cdece8b1c1215211be6

Observation 141736e6-6844-4070-9a96-8377261db06e · outbound

This paper cites A deep learning based no-reference quality assessment model for ugc videos,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs A deep learning based no-reference quality assessment model for ugc videos,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.115372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:45.223546Z digest=sha256:43d1667951b333f97dd7ab150b33dcd6cd4539d6ac74f30b05a3253563609eef

Observation 075bcaea-ff09-4a68-b72e-52c2829fc70e · outbound

This paper cites Quality assessment of in-the-wild videos,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Quality assessment of in-the-wild videos,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.923996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:45.303727Z digest=sha256:c90cdaf044632fc3314753c839b5267817273eadedf8c3814b5ede4e6b979d77

Observation db5585b8-a3c6-4ad3-8baa-48519105547e · outbound

This paper cites Ugc-vqa: Benchmarking blind video quality assessment for user generated content,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Ugc-vqa: Benchmarking blind video quality assessment for user generated content,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.755136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:45.396067Z digest=sha256:d7073d50b0e563b03d1a620be7cdd56a29ee55197f0eb0ca47a3539c6a8c4a5d

Observation 015e49cc-725b-4faa-83a4-9c08c135c34f · outbound

This paper cites Ve-bench: Subjective-aligned benchmark suite for text-driven video editing quality assessment,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Ve-bench: Subjective-aligned benchmark suite for text-driven video editing quality assessment,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.658082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:45.493803Z digest=sha256:c7b4044fe778ec7c1eae907f83f35d61e93c758b3e4cc51e7755a13211070aba

Observation 67dc21db-bbe0-4bd6-a807-dea04770e5b2 · outbound

This paper cites Subjective-aligned dataset and metric for text-to-video quality assessment,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Subjective-aligned dataset and metric for text-to-video quality assessment,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.566921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:45.569424Z digest=sha256:d36bdb29b061c48b3f8e49cae65c67e8c5f8b9af31f45cc9e7ecd0235a46cb17

Observation 382ed968-9186-4e37-a265-2431fb21154b · outbound

This paper cites Aigv-assessor: Benchmarking and evaluating the perceptual quality of text-to-video generation with lmm,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Aigv-assessor: Benchmarking and evaluating the perceptual quality of text-to-video generation with lmm,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.397791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:45.660369Z digest=sha256:bb673aa8c8b620d4009a01960eabc90b04d407dfc4ad36ce5b605a3d4c9c8936

Observation 37e63f66-13fb-4635-8466-3d8fe8f257ed · outbound

This paper cites Cvpr 2023 text guided video editing competition,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Cvpr 2023 text guided video editing competition,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.263392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:45.715156Z digest=sha256:26f5f41d6d5ad642677549eb089b1131897cf56d9d1f6b86a4184f01b5732a2c

Observation 8567a7a2-5497-4f38-bc97-57584fa67d93 · outbound

This paper cites Harnessing the power of llms in practice: A survey on chatgpt and beyond,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Harnessing the power of llms in practice: A survey on chatgpt and beyond,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.127769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:45.772617Z digest=sha256:83b01b5065932ede90277419b1dde0ca20039a9b1736aadd3c1606451cd364ac

Observation 909a4cb3-8658-48fb-9a50-7f876484b7a9 · outbound

This paper cites Señorita- 2m: A high-quality instruction-based dataset for general video editing by video specialists,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Señorita- 2m: A high-quality instruction-based dataset for general video editing by video specialists,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.957500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:45.864925Z digest=sha256:b37d43bd1a46699fec7a1a24004502210e7352374398f08aef3aa1ec81350051

Observation fc602e93-00f8-4104-9789-58f37f95a2e1 · outbound

This paper cites The 2017 davis challenge on video object segmentation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs The 2017 davis challenge on video object segmentation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.811565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:45.952544Z digest=sha256:f68d5860d43a2c089b24cb755f625d8995ac5fdbb30744c07b1637345d529a6d

Observation b3a43984-edbb-4e33-8d52-60aa4f3a2ae7 · outbound

This paper cites The kinetics human action video dataset,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs The kinetics human action video dataset,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.691934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:46.050651Z digest=sha256:84cfa26bda97e2279b3d2d99e45a97bf9daf05014ea5f3f7f28db5c23c295d89

Observation a159a87b-60dd-4c58-8742-d0082cd328ed · outbound

This paper cites JimengAI.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs JimengAI

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.573138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:46.117830Z digest=sha256:fb00efe776ce0630ae5c197fb15e8e8e7f3c57650bb5348b73fbb9a042720bef

Observation e6e4e421-f9e6-476f-9d65-a954de5c28f3 · outbound

This paper cites Methodology for the subjective assessment of the quality of television pictures,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Methodology for the subjective assessment of the quality of television pictures,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.428466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:46.186786Z digest=sha256:02a20466602efb5207c3e395d5f6ce997212970aa660c51233fe3a4cfa59149f

Observation c9cdea49-6a70-47ab-9afd-3a6c8b0baa73 · outbound

This paper cites Blind image quality assessment based on high order statistics aggregation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Blind image quality assessment based on high order statistics aggregation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.235971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:46.255534Z digest=sha256:cacb79c9c3a29f9a9231ac0c13e4d17f6a77ea71a32f6ea1075796d9a721bea9

Observation 0fb5fcef-ca42-4cee-ad5b-f95dd1083574 · outbound

This paper cites Learning without human scores for blind image quality assessment,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Learning without human scores for blind image quality assessment,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.125353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:46.336487Z digest=sha256:cdfc66daba83681d2a3623a6140a2fe0f0cad760f3278b2c2db34be5383a3866

Observation 58668ad9-9f78-4c1b-b000-54e70d6d5511 · outbound

This paper cites Making a “completely blind.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Making a “completely blind

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.055030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:46.415983Z digest=sha256:ad9122e42e53ba5362bda88dd70670864f5fcde52ed0baa007566c406cd8eb30

Observation 147b5037-7061-4135-a851-e193d80a0556 · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Clipscore: A reference-free evaluation metric for image captioning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.999272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:46.508413Z digest=sha256:5e96316652b82d9f35d33e2634cb019f315e564525bdea8f49ecb43d5e0a9d89

Observation 0082d370-eec2-4389-a745-f24588fd8b08 · outbound

This paper cites Evaluating Text-to-Visual Generation with Image-to-Text Generation.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:46.589608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:46.589608Z digest=sha256:6e30e730918ed0b6b84164107f466199af7f1a054d32ba9c6b753462efa82958

Observation fa6aed1c-e178-4085-a73d-c7cc251bcf1c · outbound

This paper cites LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:46.647524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:46.647524Z digest=sha256:7ae524ddf0c2ae56872d23bc2afd0c0351d9ce5596affde898b27473ab358579

Observation 9720f9e0-4f87-4be1-ad52-edb8df975867 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:46.707838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:46.707838Z digest=sha256:881c609b62d1f916c08f7a8e7d14bc199ab2e84426bec17934d45df595cb1b2f

Observation b6ab865e-ded1-44f6-b9ce-1167a8e66290 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.957123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:46.796154Z digest=sha256:d8623e053011c1ad569edacde8c4cd27693045d03a1876cb0a97018e80a50d34

Observation 514cfcbf-2390-4329-9e73-c59ea492ffa2 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.808845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:46.889880Z digest=sha256:518c02b0807a214b349ea50f8862d450edc3876987e00d4e530625fdc4602493

Observation 2a2ac725-4e2f-4ca5-9db3-93e2c7a3cef6 · outbound

This paper cites MLP-Mixer: An all-MLP Architecture for Vision.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs MLP-Mixer: An all-MLP Architecture for Vision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:46.972275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:46.972275Z digest=sha256:dddcf71753dcb3f205e4a37f133f9b93874eaadcedf192d40c4ada5646842f1c

Observation 9fc44e78-5762-44f8-b03f-022e60c60f95 · outbound

This paper cites How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:47.047930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:47.047930Z digest=sha256:6ad8a54c2d573a322a3afcdab2bbbe8f7e2235f40b549a5a754cb464cdade77f

Observation 8febaff3-2ab2-4aa5-b1d0-617cdeb1442c · outbound

This paper cites When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:47.098749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:47.098749Z digest=sha256:4719baa3df409156651f20f13c772b19915d62156cec6b62110713e079b9ce36

Observation 0db7e7a8-44cb-495a-88e1-b7ee9eb8c641 · outbound

This paper cites Surrogate gap minimization improves sharpness-aware training,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Surrogate gap minimization improves sharpness-aware training,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.597132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:47.146349Z digest=sha256:6b24db9419780df536cdc0fa4c169abe17a18e252e80e5ca64781dc4b733dcde

Observation 761438d2-410b-48c9-a6b6-59cb3f6a55f7 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Lit: Zero-shot transfer with locked-image text tuning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.438080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:47.196903Z digest=sha256:b6b1831519d196cd7ad0750189d6502ed13ca5f6610212e7d91b41dd00b14618

Observation 1664f839-5ebe-4a7c-bb8f-b579c62ce79f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:47.249382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:47.249382Z digest=sha256:39a66d0ce89d40d7af1b44de9cf05344b8c1fccec976a74c056fe7d8592cc5d2

Observation 8adb38d0-4237-4e12-b5bd-e0eb6b3fe445 · outbound

This paper cites Mlp-net: Multilayer perceptron fusion network for infrared small target detection,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Mlp-net: Multilayer perceptron fusion network for infrared small target detection,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.213820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:47.303914Z digest=sha256:364f6315800de6176e83c7a0550b794150c42b3b3eded12132b4087d509bf2e3

Observation 65a60f14-dd82-401e-a048-46d7db18602a · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Lora: Low-rank adaptation of large language models,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:52.987257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:47.364459Z digest=sha256:c46800690236868b84746e88bf4a5445774bb9b9b8e431b0747647e4f4467499

Observation d53a406b-4189-488a-99ed-9272ada4c374 · outbound

This paper cites Imagereward: learning and evaluating human preferences for text-to-image generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Imagereward: learning and evaluating human preferences for text-to-image generation,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:52.781780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:47.419599Z digest=sha256:45108729e53079f7031b1d8a12a20685b9c41a73483ced0706c626bc9142f6aa

Observation c95245d3-7d6c-402a-aa39-aec719e0a493 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:52.561473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:47.497508Z digest=sha256:9c9b3f24e1f95561d093fcf530632481cd09ffedc4ce751f6c195fe9d25b46af

Observation 4f49277c-9c6d-4030-a140-a3ee6b1f4dde · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Pick-a-pic: An open dataset of user preferences for text-to-image generation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:52.397957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:47.556046Z digest=sha256:3ff8533056f2e4bf6a754724d734c2feb97d0cf70b5fba94c7580e935221cf89

Observation 11517637-b9b8-4fb5-a164-d3be77ecd053 · outbound

This paper cites Building cnn-based models for image aesthetic score prediction using an ensemble,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Building cnn-based models for image aesthetic score prediction using an ensemble,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:52.174546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:47.618742Z digest=sha256:ace4bf4cd3221cd2f667f0f8a268bac1f33f47a4cc026ef0a3c8b80705c34d50

Observation f8b2128c-3ff8-4a5c-b537-99a95b3bab61 · outbound

This paper cites Llava-next: A strong zero-shot video understanding model,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Llava-next: A strong zero-shot video understanding model,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.951951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:47.671411Z digest=sha256:516999e8abe64601b0dde5fea7ac89e5171175f3bbe2c3e10180fdc4e16cbf54

Observation 92317989-0820-4ead-b432-384fdfa782e0 · outbound

This paper cites Internvideo: General video foundation models via generative and discriminative learning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Internvideo: General video foundation models via generative and discriminative learning,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.692478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:47.723918Z digest=sha256:b1c12888a3be268386e0d1703d181ad91a2a2aff5552b4d5d945adda5a144f3a

Observation b97954b8-5d3e-4053-a390-0a9036c01158 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:47.793673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:47.793673Z digest=sha256:5dcc04c8fdd5ab7858b9118c5fced350e6eb2ebe77ac8e6713a52bdb3c6f3c4c

Observation 16b6cae0-883d-4759-b36e-d73e8eec2a68 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.511413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:47.842241Z digest=sha256:403cfe64b3c2623b3f61f6be73421ce819d6b5937bc1ec8c29ce10e649ee2141

Observation dbc1d83f-1c7a-43b1-995f-c66bac60f979 · outbound

This paper cites mplug-owl3: Towards long image-sequence understanding in multi-modal large language models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs mplug-owl3: Towards long image-sequence understanding in multi-modal large language models,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.344286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:47.905970Z digest=sha256:f2e2754c91d29749df32aff694e96a68f7591424a62d9349611b0d5d15ef17c5

Observation caaec614-1d36-46d3-b353-25615a657f3e · outbound

This paper cites No-reference image quality assessment in the spatial domain,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs No-reference image quality assessment in the spatial domain,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:47.959442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:47.959442Z digest=sha256:1cbfc122be6e0a4d0dc7b1acd6ab54a154ce4e31048d33f37220af2470106428

Observation e78e9d04-ab79-454a-9fff-8fdf3e641069 · outbound

This paper cites Blind image quality estimation via distortion aggravation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Blind image quality estimation via distortion aggravation,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.198918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:48.029727Z digest=sha256:217ab78849d108e740e2179cdb4b7f0cf37ca52be73890f48d5127b3d8519f6b

Observation 79cbeca4-a14e-441e-be03-e04a34a93f5b · outbound

This paper cites Blind quality assessment based on pseudo- reference image,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Blind quality assessment based on pseudo- reference image,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.004024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:48.090491Z digest=sha256:244622d0b75169e8d462dbbe5f72a15ba669bba0d481bd663a6b0236718dadff

Observation 403b8b0a-6ae7-426a-a442-ceabac5f51bc · outbound

This paper cites Visual instruction tuning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Visual instruction tuning,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.138760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.138760Z digest=sha256:6f0f8d9503163e7c556901c133923305bf14f5ce75da86fb54778fe3963427d8

Observation 1f8005a3-abc0-497d-a663-8d9b958afd0b · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.190049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.190049Z digest=sha256:c5b80895a01bc59b74360834b90a1584967eecfe7447a035257d0baff9123462

Observation f883cf17-f2ed-408c-9b71-becbad8dd081 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.245414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.245414Z digest=sha256:e3dadef2b8a2a225d60f552991ede961a0fa60b8b2b63b1a7b8bb6f99c9f9f0b

Observation a9df28ab-111c-4a68-821e-a7aaf4172c19 · outbound

This paper cites Improved baselines with visual instruction tuning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Improved baselines with visual instruction tuning,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.349429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.349429Z digest=sha256:6e53f1a7e94e2ff7bdd340603f7531999493f2fb1689ea8b943dc0381fd5442c

Observation 865709a4-8de1-4b48-be07-011908404bcd · outbound

This paper cites Llava-next: Stronger llms supercharge multimodal capabilities in the wild,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Llava-next: Stronger llms supercharge multimodal capabilities in the wild,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:50.758104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:48.412809Z digest=sha256:d2d49e44e3e22a8d3381310cc3ca985e1fe49348b22b7dd1a7cdb5c842345229

Observation 8b7fa63b-bfa7-40b9-993b-d750b86428ca · outbound

This paper cites Llava-next: What else influences visual instruction tuning beyond data?,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Llava-next: What else influences visual instruction tuning beyond data?,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:50.534207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:48.473759Z digest=sha256:556e44bf1518ef08de4188b0a06429bbc6d104270fc2eae25c41e01db79140e7

Observation 18370840-77bd-42b7-9442-253997febc9a · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.537796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.537796Z digest=sha256:7ff90aefb6fa2c1452dea01369ef3e3aaf63d9554ba057efef20c85aea85876d

Observation 2a484005-937b-47b3-912e-9e815524a744 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.595108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.595108Z digest=sha256:74262ff6257c576ddb933c5e20abba36ce09227353e2cf9e36cb877cf99a805f

Observation 3e219ea3-4a27-42f1-9288-8bbe7e7f46c3 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Learning transferable visual models from natural language supervision,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:50.332408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:48.673221Z digest=sha256:47edbe83ab3862d957ae20244f61480c964947f265c78ad3850da541c012837c

Observation dba5ab64-a687-4600-bbb6-142a5c5a19b2 · outbound

This paper cites Unmasked teacher: Towards training- efficient video foundation models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Unmasked teacher: Towards training- efficient video foundation models,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:50.106302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:48.732449Z digest=sha256:9d65876173a6dbeb03b0500ec018d7703c83a055c1480b5a111a21f338d522e9

Observation 7e7a514f-68bf-41ce-aea6-7b24568ecc26 · outbound

This paper cites Stablevqa: A deep no-reference quality assessment model for video stability,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Stablevqa: A deep no-reference quality assessment model for video stability,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:49.870172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:48.810810Z digest=sha256:7918abf1b93986d7ec5db34bc053614062644f1da996eca1653244d6978bc661

Observation f8d41227-00e3-4ca1-8fdf-781525ee55a7 · outbound

This paper cites Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.881635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.881635Z digest=sha256:aacc45072ee2668fd62642340de5303a257c806312407de6b4f246b6aafe36f6

Observation b589e717-2e61-441e-bf33-a1400fd02fa0 · outbound

This paper cites Excellent.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Excellent

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:49.655959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:48.932160Z digest=sha256:33c4396975968a380e2f95c121e59e9b674c9adf5525af7d8f34c38ccb68a0b0

Observation 4d17d73d-ae3a-467f-bceb-a0c55182ad0f · outbound

This paper cites Excellent.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Excellent

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:49.409764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:48.995779Z digest=sha256:ff853715b48d85e609887296d1bde18a1146d73672698edddd02d7d1d9b157bc

Pith citing papers

No inbound Pith citation observations are available.