Pith. sign in

Paper Citation Record · LEDGER

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation

As of 16 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 4 inbound Pith citation observations for arXiv:2505.12098.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12098 v1

Coverage vector

measured 100 of 111 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:46:46.420145Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:16:47.490821Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:56:59.272327Z

Reference resolution

100 of 111 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved70
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 61a35551-7f2a-4690-942f-dad9a23d3306 · outbound

This paper cites A survey of ai text-to-image and ai text-to-video generators,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation A survey of ai text-to-image and ai text-to-video generators,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:45.959388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:45.959388Z digest=sha256:e5fba124249c06d8fe00cf2ca5b5155cfa2ca644de90bfd41358f907ec872097

Observation 5b8f57f8-2501-4f08-9e6e-c24c02ac98f9 · outbound

This paper cites A survey on video diffusion models,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation A survey on video diffusion models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:45.963621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:45.963621Z digest=sha256:78d9f12b5ed42282d00319947fce8a8badafa9fb0fec739001639fe031dfaed8

Observation 2e44c3a8-09d5-46dd-9c68-fe981fa4d090 · outbound

This paper cites Evaluation of text-to-video generation models: A dynamics perspective,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Evaluation of text-to-video generation models: A dynamics perspective,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:45.967286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:45.967286Z digest=sha256:5ee2fe9075d8a3ce7ace969fb14e2b78b877f64d26384bc3c31780545e47d973

Observation 15dd866e-086f-4f47-83ab-5347401615be · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:45.970921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:45.970921Z digest=sha256:ceb7b4aab23499fd05fa64af2bcc3fb5ca5440d0ddf4af9bf67a9b54fc6e5e42

Observation 15736eb3-5dd8-4477-91e5-21bddb86a63b · outbound

This paper cites mplug-owl3: Towards long image-sequence understanding in multi-modal large language models,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation mplug-owl3: Towards long image-sequence understanding in multi-modal large language models,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:45.974661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:45.974661Z digest=sha256:d5081ff7c7faaaa6319d73d0a004e85e1cd765c863a9c8e166ae79bab7c6f82c

Observation 60c06238-d86d-4ead-a420-149f21fbe5ea · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:45.978400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:45.978400Z digest=sha256:694034a6ecec8c053deb72926348d3075eb8888f2f68585cb0f24d4e7f365384

Observation ca2b3726-1037-499f-9fc2-c8bb07261c1a · outbound

This paper cites Aigv-assessor: Benchmarking and evaluating the perceptual quality of text-to-video generation with lmm,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Aigv-assessor: Benchmarking and evaluating the perceptual quality of text-to-video generation with lmm,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:45.982353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:45.982353Z digest=sha256:3e8c30a495dd483d747b67f1822b80e99a0a292f2e550c58d13eb11085b6994c

Observation 99c91dc4-bbf9-4826-bda9-d1e73bff9f72 · outbound

This paper cites Evaluating and improving compositional text-to-visual generation,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Evaluating and improving compositional text-to-visual generation,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:45.986410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:45.986410Z digest=sha256:84a95c74b1e431142f3a23707374bdfbd90ed01b0d92a86b97109af2474955dd

Observation 58aecaa8-fdb7-4a39-826b-71ceb23667e1 · outbound

This paper cites VBench: Comprehensive benchmark suite for video generative models,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation VBench: Comprehensive benchmark suite for video generative models,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:45.990501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:45.990501Z digest=sha256:9fbf38a3ea0f5a6ef1e82803a79a5c96221a22f40f22a9c04fea88d29467487a

Observation b5ad89dd-09f6-486f-ae60-2f0bf0f1592f · outbound

This paper cites VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:45.994146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:45.994146Z digest=sha256:801e7df7e8100f83bd419116cbdc26c62a36c8afe5616b2f2ae84f2ad6d65470

Observation 9793d22f-e635-431c-aa52-446efa8aa278 · outbound

This paper cites Eval- crafter: Benchmarking and evaluating large video generation models,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Eval- crafter: Benchmarking and evaluating large video generation models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:45.997995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:45.997995Z digest=sha256:66de5d422d39b23604ea3acc0572c38eb28ab5a5d8d9331954fe0b99aa6ab883

Observation a74f312d-709c-426c-86ff-64f71f99afc7 · outbound

This paper cites Fetv: A benchmark for fine- grained evaluation of open-domain text-to-video generation,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Fetv: A benchmark for fine- grained evaluation of open-domain text-to-video generation,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.001480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.001480Z digest=sha256:26b4ba56df315361f8c129de98feab1391bee615a511360c161541f24f2e0291

Observation efe373d1-982b-4f13-be01-6203fd93c494 · outbound

This paper cites Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.004863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.004863Z digest=sha256:8712cc381d29920b5894f740339ac563d24041dc976320b810f18007829097ea

Observation 5ca14115-61a2-4657-a9ba-fb77ed92e633 · outbound

This paper cites Measuring the Quality of Text-to-Video Model Outputs: Metrics and Dataset.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Measuring the Quality of Text-to-Video Model Outputs: Metrics and Dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.008655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.008655Z digest=sha256:3fd81713195d71110ef5c0cc0529c84a7272ca8a8a92c6a9cf0c84b34a3c4083

Observation 47c68798-8b00-467d-a50b-6009b943bfeb · outbound

This paper cites Benchmarking Multi-dimensional AIGC Video Quality Assessment: A Dataset and Unified Model.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Benchmarking Multi-dimensional AIGC Video Quality Assessment: A Dataset and Unified Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.012113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.012113Z digest=sha256:02f70e479acb74f27cc7ea7cbfecdfd39924821d3fac6e595131f4a37e47aae5

Observation c867d89c-cfd2-446f-b1b3-820443e436ff · outbound

This paper cites Subjective-aligned dataset and metric for text-to-video quality assessment,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Subjective-aligned dataset and metric for text-to-video quality assessment,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.015859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.015859Z digest=sha256:7a9e4e6dc11c3f3b12e0327c9677b15f255b0ed92c348c00d6dee57679a4747f

Observation c5dddfb8-e313-44b3-80c5-4a77023e37df · outbound

This paper cites Gaia: Rethinking action quality assessment for ai-generated videos,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Gaia: Rethinking action quality assessment for ai-generated videos,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.019327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.019327Z digest=sha256:ddf76d83e03d3f786ba6e12d9a35b1c2d98d4adc9826361d020b17c8c308b2c3

Observation 5c02f201-e318-4cea-95e2-fbbe8a392221 · outbound

This paper cites LMME3DHF: Benchmarking and Evaluating Multimodal 3D Human Face Generation with LMMs.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation LMME3DHF: Benchmarking and Evaluating Multimodal 3D Human Face Generation with LMMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.022858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.022858Z digest=sha256:da5cd795fdda173ff5de8377e104deb119cc0720f737c7fd0e8aff700d4aa3a9

Observation 9be9ceb2-d771-4ee6-8171-43822aac8f3a · outbound

This paper cites Quality assessment of in-the-wild videos,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Quality assessment of in-the-wild videos,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.026958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.026958Z digest=sha256:7d5433222a792048b1b86f0195a109af416d83f16a1ad2eeac0689e9e41f98dc

Observation c41296fd-8525-42d1-ae45-26c8bb5b14cf · outbound

This paper cites A deep learning based no-reference quality assessment model for ugc videos,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation A deep learning based no-reference quality assessment model for ugc videos,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.030276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.030276Z digest=sha256:b1a6ad9d2ee7640786aa8007b56821af342b1f08cfe4ccb198df862535f8d63a

Observation 4d7c3b6b-9ed3-4db4-af6b-4d82902141dc · outbound

This paper cites Fast-vqa: Efficient end-to- end video quality assessment with fragment sampling,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Fast-vqa: Efficient end-to- end video quality assessment with fragment sampling,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.033489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.033489Z digest=sha256:13eda25b2489b7dd742bf7070c384652224d0962f1c18ea4f10efd8e5d79e3cd

Observation 491b2055-99c1-4323-b4df-01a08c10689f · outbound

This paper cites Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.036660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.036660Z digest=sha256:d9d6ba9d281848ccfb2399f44101ac3b3e5ee7ca256351b7e5fbb87033f6f001

Observation 705c6c3a-6ce4-4d84-9a0d-d502e1533332 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to- image alignment,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Geneval: An object-focused framework for evaluating text-to- image alignment,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.040135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.040135Z digest=sha256:3a6782ff8f102396912f7133674a28f71f6b73f9c422a64bb9a953881ead6958

Observation fefb36c9-4029-4bff-8b82-3b649d8dc024 · outbound

This paper cites Methodology for the subjective assessment of the quality of television pictures,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Methodology for the subjective assessment of the quality of television pictures,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.043320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.043320Z digest=sha256:4a5aaebda56b9bbd6406090befa4b23331262276f120415a1265a8e5a390137b

Observation 642c64cc-0541-436c-b14f-fde4909a6721 · outbound

This paper cites Improved training of wasserstein gans,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Improved training of wasserstein gans,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.046897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.046897Z digest=sha256:50add1c593909244384987edda16d8dc5392d3b89ddf942b61e1868c4ce68d7f

Observation bb565e60-b0e0-4449-8dbb-dc2d3829b8d4 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.050421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.050421Z digest=sha256:c94661aa6dad49c5f65a78255107ee741ee108c55096e366189be3931800a6f9

Observation f85bd524-f68f-455c-84d2-dc9821871fbc · outbound

This paper cites Visual instruction tuning,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Visual instruction tuning,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.054000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.054000Z digest=sha256:1766c2175f1e00ca3380844d59c05cfaa8270acf020a10aac69cc7505d2ee344

Observation e42132e7-9319-4fac-a383-c50f264635ce · outbound

This paper cites Making a “completely blind.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Making a “completely blind

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.057427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.057427Z digest=sha256:912d101f24992e1236cae92f97dd743a4cd779ac2977dc38b0445a633e3ad4ad

Observation a6482f19-a5df-4152-8113-e86a493021df · outbound

This paper cites Aigcoiqa2024: Perceptual quality assessment of ai generated omnidirectional images,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Aigcoiqa2024: Perceptual quality assessment of ai generated omnidirectional images,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.060690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.060690Z digest=sha256:847832632c43f62a7bfaea6d19c8ff7dbeed3d95b863770275ef876b13f2e0ed

Observation 4ac8e209-4e15-4fe0-8852-1362bf47cf05 · outbound

This paper cites Finevq: Fine- grained user generated content video quality assessment,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Finevq: Fine- grained user generated content video quality assessment,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.063857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.063857Z digest=sha256:ddbafd5fdc2bf1e765ad2920c0bbaab42f7b0474d32e2a8520b7e54c42b09bd1

Observation 5fe07f86-cec6-48dc-920f-13b7e973806b · outbound

This paper cites Quality Assessment for AI Generated Images with Instruction Tuning.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Quality Assessment for AI Generated Images with Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.066915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.066915Z digest=sha256:50d8e1eb01b4526d88208d60d2e495da92da6e65d07135a35d6a5128ef18f517

Observation efec826d-5ebc-450d-b188-a0d626f9f1cc · outbound

This paper cites HarmonyIQA: Pioneering Benchmark and Model for Image Harmonization Quality Assessment.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation HarmonyIQA: Pioneering Benchmark and Model for Image Harmonization Quality Assessment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.070528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.070528Z digest=sha256:18177d74d323c8641405bf19639a335e49c55942f373cedead521e39f46c344f

Observation 74bee7d7-868f-4ce4-80fb-9050a8314d73 · outbound

This paper cites Blindly assess quality of in-the-wild videos via quality-aware pre-training and motion perception,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Blindly assess quality of in-the-wild videos via quality-aware pre-training and motion perception,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.074336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.074336Z digest=sha256:a18e932ce76c404a51b318a6fe23b8a6507f877e3cecadd1e99d35bf07d85fa3

Observation c5c48b8f-e08e-414f-994a-5413c6cf487c · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Pick-a-pic: An open dataset of user preferences for text-to-image generation,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.077655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.077655Z digest=sha256:1fe843d44d286af62df779160538e1c8e002a284b24c3ab2483bfb2c9a01b94c

Observation 558d7112-d39b-42ab-b3ca-99b2ec0977f7 · outbound

This paper cites Learning without human scores for blind image quality assessment,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Learning without human scores for blind image quality assessment,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.080975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.080975Z digest=sha256:2a8131d0ea1018d46e0c3f163e70545dd0a5b95c190ce0358360e649fb1529f5

Observation 94afa634-d82f-41bf-b106-4bb9f72c6dcc · outbound

This paper cites No-reference image quality assessment in the spatial domain,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation No-reference image quality assessment in the spatial domain,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.084249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.084249Z digest=sha256:d41900bb07a171d191ca7b2b8bf18ec8094a35a0b7ed3468deb41282cafeb007

Observation 8a64aec0-a3db-44a4-b562-8a18a208db06 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.087382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.087382Z digest=sha256:2d17c1ac711fa3a110144a3ec8bb677b8f4742bea95a24e2bd9e346133cd9337

Observation f6c507e3-dc62-4164-a420-9c01f9638d0d · outbound

This paper cites Jimeng ai.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Jimeng ai

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.090759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.090759Z digest=sha256:01aac26dfc1158ff2a576b1d197b38d46d355acd51f016529886533391b1a156

Observation 2d753434-345b-4f4f-9da8-3d7d708f9c0c · outbound

This paper cites Vidu ai.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Vidu ai

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.094124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.094124Z digest=sha256:c1e0cf16647e0ccc6085e577599978b2517c2f5772db3dd9e70f8c757593f129

Observation d64f5add-52ef-4c8e-b85d-d840da771d9f · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Lora: Low-rank adaptation of large language models,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.097888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.097888Z digest=sha256:1a1e0469234431dddb388a2d1bedbfe2bc8ead7551f4b2b747413b10252786b5

Observation 50b0051f-8ca3-4046-a3a8-b833e1ecee52 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.101913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.101913Z digest=sha256:efd02cb4cdcad5b15533c330b1a43c16d597d26993f57713fa1367b7cef41abe

Observation e2945d9b-871a-4074-8e6d-e0bf0d337c35 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Swin transformer: Hierarchical vision transformer using shifted windows,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.105604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.105604Z digest=sha256:4a6a603fd3f6e6895e8bc9030222ae0499888843dba16ee6e6ca29d2dda397f2

Observation 456be38c-fd37-4916-8909-1122959912a0 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.108773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.108773Z digest=sha256:c0231ca316310d3a9f4d0608c7fde056e093b3b3851fb1c5d19d15c0d5192068

Observation 94b9038b-a310-496d-8777-1b5c45d280a9 · outbound

This paper cites Blind image quality estimation via distortion aggravation,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Blind image quality estimation via distortion aggravation,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.112393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.112393Z digest=sha256:b820497f7d885100a173582767f550f2c68ed8f3cbe54b7f66eee3f250912abd

Observation 27a26964-fa7c-4fb5-869c-05b65e8847a7 · outbound

This paper cites Blind quality assessment based on pseudo- reference image,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Blind quality assessment based on pseudo- reference image,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.339662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.115639Z digest=sha256:34baa292a3544e76dffa336b7b0290c6dcb3f24a6511394ebe89bbb80b14c41a

Observation d1f400b7-5a34-49de-aaec-d07c14a38e67 · outbound

This paper cites Blind image quality assessment based on high order statistics aggregation,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Blind image quality assessment based on high order statistics aggregation,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.330776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.118767Z digest=sha256:ae471885914af24d1b0808db59c49b01b96aa27eb8640e42e1fb35617aa4feb4

Observation 13363b51-a1d3-46b6-acf4-f736835c2e97 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.122132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.122132Z digest=sha256:9c7a110ed5cebf9ab1d1b29bcf6f208616b164ec72288efca1b3c5baa32359bf

Observation d046d088-d194-4c38-aa56-1731522c3a33 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.321573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.125407Z digest=sha256:260460262a85b0270b47cba64179462ddc7954f7ce49cfed21661e1f0dbb3337

Observation f31a70b2-0af7-48ff-995b-babaaf0dd569 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Laion-5b: An open large-scale dataset for training next generation image-text models,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.312234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.129169Z digest=sha256:d958d269bc39236f8c0067a02e9c9031cdef37ceea3e919282fa31d0e1eb13b4

Observation 4c3d2a88-1a51-4370-afd6-d445633d07e2 · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text-to-image generation,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Imagereward: Learning and evaluating human preferences for text-to-image generation,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.303461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.132480Z digest=sha256:1aa3e850a631c79bdf9feac77c94925dd4ad892beb47c3065828dd16c15660cd

Observation 0a09fa6c-6547-4cdd-94fb-06c22389c9b4 · outbound

This paper cites Human preference score: Better aligning text-to-image models with human preference,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Human preference score: Better aligning text-to-image models with human preference,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.294412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.135603Z digest=sha256:01b1e9636e51d885fcbf206e8bab9b340ba5d05a9b3ebbb5af8b88e8ee010935

Observation 6c5980de-0401-4a87-bc22-9002791225ac · outbound

This paper cites EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.138692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.138692Z digest=sha256:15b2af3c780ff798a5e3e28d09957f588070b9b13b46d53232ad2cfc00ab86a8

Observation 05ecb803-3790-4fc2-8c0d-8daf8f1fbfef · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.142211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.142211Z digest=sha256:a0b0950c2e5c25725adddc6bf1121b0ce77ec1c8209dec32bb96d11269da911c

Observation 1d4bc4fc-0be7-4b65-a30a-ecc2c620d7b3 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.145646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.145646Z digest=sha256:3ad82acd77b2292331151973db958d770ac0170ee6a673243f79bb611ba14440

Observation 7c114032-4fad-4103-8b82-879326445598 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.149232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.149232Z digest=sha256:f79ef7b868fd5f8bdbfc2d7d052900b301ff89f0e5e3653d1fa68542083ffe1e

Observation 9c1e9db0-baf2-4b5d-876e-623b9dd10707 · outbound

This paper cites Qwen2.5-VL Technical Report.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Qwen2.5-VL Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.266138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.266138Z digest=sha256:0a95f8f3bf0aee654191937c24c61f1ac371788a578b99ba7ede713f6c74adef

Observation cb615db2-9e93-4f70-85ec-ce592d6f28c0 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Llama 3.2: Revolutionizing edge ai and vision with open, customizable models,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.285190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.269857Z digest=sha256:51feec49cc55e24576557bc12828b58fa07da3182bd4fc76798990c342d9f463

Observation 970e3531-6178-436b-9a17-85be54aa8254 · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation CogAgent: A Visual Language Model for GUI Agents

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.273001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.273001Z digest=sha256:deb9473d4cfa380188c939b609e8965405fde7823be5051512d2975020ce0bb1

Observation efe47176-b4ff-43eb-87c1-804b72632c54 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.276526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.276526Z digest=sha256:ea42ca76bc10de688433853476d0a4f4571dc5dc87d5063a070ab7162a138d5d

Observation e651e428-6d02-48fa-a6a5-4c574a25ea3e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation LLaVA-OneVision: Easy Visual Task Transfer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.280449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.280449Z digest=sha256:0a16e847ac1f022360c277e4240544a8de2c0433f74371ebcb9f1fcb87c607bf

Observation ff836913-1b37-4ce2-aa6f-0e068e386471 · outbound

This paper cites Gemini1.5-pro.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Gemini1.5-pro

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.275500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.284031Z digest=sha256:1e42b0b341c2a2c1e485cc7155c180a26cf4b063bcf75773020767d771332667

Observation d482f168-e05a-4ca4-aeca-9b013a45b568 · outbound

This paper cites Claude3.5.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Claude3.5

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.266572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.287156Z digest=sha256:e13ff80deeeea6bf70d4554e0253cf83416f0b63331577da388f3f6a66432f8b

Observation 9144f4d9-4333-4048-91b2-af05fb14161e · outbound

This paper cites Grok2 vision.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Grok2 vision

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.258149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.290577Z digest=sha256:ec1843b7a80d4ee940b8bee22d2eb7fe7d56c615ff5c4132a17d9c3b17178546

Observation 9243cd1b-c85a-4e4f-a461-f9e47bb5fae0 · outbound

This paper cites Chatgpt-4o.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Chatgpt-4o

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.249632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.293994Z digest=sha256:d45cccb54d8625625e65f69bef3dfbbed5ab78f1b554bc3b560aca460e6f4370

Observation df98fcd3-02db-4a78-86a4-45cbea9d647f · outbound

This paper cites Pixverse: Ai video creation platform.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Pixverse: Ai video creation platform

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.241020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.297587Z digest=sha256:9308878d991033f2aa12e94f37337d468db297a0e0d498dea7f12396de852b47

Observation 553f17d8-6a13-4630-b229-d8620dc11c56 · outbound

This paper cites Wanxiang.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Wanxiang

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.231965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.300873Z digest=sha256:6ed42c1a29627181291c4a554580c269b8d74a6b35fd8ba314f02202fa28d448

Observation ba761b63-d2da-4fe2-bde6-7130fe9a28f7 · outbound

This paper cites Hailuo ai.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Hailuo ai

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.222778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.304066Z digest=sha256:887e2f14ad51a449936d1bd82f825ebf45f764189232d359b4f45945c8d4d872

Observation 512faa2f-f5b9-477a-8f56-0af0f4310f3b · outbound

This paper cites Team, “Sora.”https://openai.com/research/video-generation-models-as-world-simulators ,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Team, “Sora.”https://openai.com/research/video-generation-models-as-world-simulators ,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.214627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.307108Z digest=sha256:41b71e2158294bf74e16d1666963a24c81fa3aa39a96bc8a186133bfb3944510

Observation e11c36e8-7e6e-4681-8621-271b32ffcb65 · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.313528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.313528Z digest=sha256:92cbaa703d5241573c6816a5da884943b59a58abc816235aeb17b79c8f5a9fbe

Observation e37b01b3-705c-4b39-8c5c-4b5570bfb9dd · outbound

This paper cites Introducing gen-3 alpha: A new frontier for video generation.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Introducing gen-3 alpha: A new frontier for video generation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.196511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.317127Z digest=sha256:b07896e6439c18e4c43f708027b628351b7f7122b609635d1156449ae82d272e

Observation d3e6c0fd-1c2c-4398-a063-509f1150adbe · outbound

This paper cites Kling ai.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Kling ai

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.187346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.320122Z digest=sha256:bce28244cb006dc902c0f912b3ac1f34c283b7c8549f7c894db3a4826b3c90d6

Observation ece6ebe9-88e5-4fa7-b2cf-ab5b365c8dbd · outbound

This paper cites Team, “Gemo.” https://www.genmo.ai, 2024.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Team, “Gemo.” https://www.genmo.ai, 2024

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.178265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.322949Z digest=sha256:5a3c8acd7e21467b5531d19eb2b7c4b63def166ed081b72064d84c27c6cdeed7

Observation 1827f4e8-bbc8-4fa2-9245-2eccf06fba5d · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.326100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.326100Z digest=sha256:e9465409c79911bb5ab6f350cafe8437f0adb7a2095c783a27d8dcc411ff3cb9

Observation df667d26-fcdb-4311-96eb-eb9b0164927f · outbound

This paper cites Team, “Xunfei.” https://typemovie.art/, 2024.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Team, “Xunfei.” https://typemovie.art/, 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.169833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.329264Z digest=sha256:bc541d50dd650a0e26a00ef8ad363122beb0750c09194a31399487f1e25a3bf2

Observation 11fcc47c-0da7-4c0d-b2fc-199df9ef65b3 · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Pyramidal flow matching for efficient video generative modeling,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.332082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.332082Z digest=sha256:2a09b953dbbd7f921fb30cce4919723827292bc9753ff16563aa183ae5c2f627

Observation b19dd4e6-a8d4-42f3-a353-c4f399307d56 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.335101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.335101Z digest=sha256:3df279004b4e9e92d91679b446e11302fdf0337a6da92949b50d3fd019f02c04

Observation d787cc19-6e4b-48af-8967-89d90ccde02e · outbound

This paper cites Allegro: Open the Black Box of Commercial-Level Video Generation Model.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Allegro: Open the Black Box of Commercial-Level Video Generation Model

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.338428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.338428Z digest=sha256:6c9f153e93fbe2fe01ce4c21d93cb9156243deed854c036278fc7457a5796ba8

Observation 8cb38cda-cf41-44bc-b751-6bf9f3ae61a2 · outbound

This paper cites VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.342097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.342097Z digest=sha256:80b2eed2a529271e86a61eda8465272145785b60c54529912c9375bb59f17ed4

Observation 08df4b1d-f2ad-402c-8e47-dd0efceea38b · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.345771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.345771Z digest=sha256:f22539c8a9b1d42581d58ea9a86155f98179a6fe05ac0b35022c6e78214ed555

Observation 172303cf-eef3-4237-acd8-a9e7b98addaf · outbound

This paper cites Easyanimate: A high-performance long video generation method based on transformer architecture,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Easyanimate: A high-performance long video generation method based on transformer architecture,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.349363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.349363Z digest=sha256:548c34864b1bf3703f24acfcff9e29cb6b58c3ae5ad109f3191a92ac340c8c21

Observation c15cdad3-2c9c-4dbf-9b58-b6c45adcc72b · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.352495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.352495Z digest=sha256:76a5e524ef60282bce7cf818cd4764f502ea3af25c3e98cc3d112ff1b4d4a472

Observation 8ba89d31-8912-4380-a8e3-fe8a242c480d · outbound

This paper cites Hotshot-XL.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Hotshot-XL

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.160943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.356953Z digest=sha256:b5a8301a5773b6719d64161b0280965587221c638605a244949b6bfb177b1cd4

Observation 64d0ee34-a58c-4238-9783-39a4d8531532 · outbound

This paper cites Latte: Latent diffusion transformer for video generation,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Latte: Latent diffusion transformer for video generation,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.151841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.360187Z digest=sha256:06b40ec9b539ffae2ecb3fa3e9fe99fa85cf09a56c39f213efbfeb392399d8a4

Observation 69fd7429-2994-43d3-842e-b96fd21aec47 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.363343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.363343Z digest=sha256:692b3f8f26370a6274322b05de8b4c24aad3f8d757c53f01b2f10accc0fec14d

Observation 7a4fc38e-3dd6-4096-acf5-f0942e7a6183 · outbound

This paper cites Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.366831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.366831Z digest=sha256:870e0581434a20f1b948d2c97872d37a32833a640f2cd07014d5374df6bf5fee

Observation 87c5e10e-66cb-4f8a-8f30-d3e460b0ffb3 · outbound

This paper cites Autoregressive Video Generation without Vector Quantization.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Autoregressive Video Generation without Vector Quantization

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.370251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.370251Z digest=sha256:11be59a1b71c029890f92bebac81ff1dc8a76a4344910841c7a37eae2350af09

Observation b662b758-6f11-440e-947c-bb0e0e3ccbff · outbound

This paper cites ModelScope Text-to-Video Technical Report.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation ModelScope Text-to-Video Technical Report

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.373929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.373929Z digest=sha256:df778ce6a9d542962d47a0849ed6fad589992a76bdfb45810bfaad2cf9101ded

Observation cf06afa6-8297-4d57-9058-49757e28e44d · outbound

This paper cites Tune-a- video: One-shot tuning of image diffusion models for text-to-video generation,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Tune-a- video: One-shot tuning of image diffusion models for text-to-video generation,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.142748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.377241Z digest=sha256:62405ff8b452068a2248c732d58d444a6bc205ff0319467b4d9ddecb48234fd9

Observation ba8d69a9-4038-4117-a7cb-d31f50b5300a · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation LTX-Video: Realtime Video Latent Diffusion

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.380500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.380500Z digest=sha256:9f2fa70e093d61f0f63c49f4b7597678422fc2f3203280523edaa661c5d040ca

Observation 176ff004-7ab7-4b43-bcd6-2f60616b7757 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.384096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.384096Z digest=sha256:7f26686947adc597830d1ae54c1dc42dda0766453f3417bcbf6bd9078c357023

Observation 296fd56a-f4ba-4111-a16c-f4911da5b5f2 · outbound

This paper cites Zeroscope.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Zeroscope

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.133230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.387913Z digest=sha256:b5bc912f44e2999b223e6c776ef2c317a34974dedb6cbd55f2043734188743cc

Observation 21f0a260-cfb8-41ff-9399-a8c1eee5c20c · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.391478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.391478Z digest=sha256:a8638865e51b603f3ad2dc7bb82861e543a45cf1b119aed6a476e44f4a7e235b

Observation 0a5f7f94-2b68-479b-971f-84ac4b686245 · outbound

This paper cites Responsible research with crowds: pay crowdworkers at least minimum wage,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Responsible research with crowds: pay crowdworkers at least minimum wage,

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.123730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.394999Z digest=sha256:39b6e6de96b296c0f40223ffe7e45516a52e4d736df4b9b92afd14fc3445b278

Observation f7c16bbd-9349-474b-8510-06ffa80dae80 · outbound

This paper cites Toward verifiable and reproducible human evaluation for text-to-image generation,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Toward verifiable and reproducible human evaluation for text-to-image generation,

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.114645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.398663Z digest=sha256:e856ce8bad8a331ae5e86204c5e3ef0e41ead187f79e0591f7334e1947f94a32

Observation af42f055-f406-45d5-b940-ce19f78f00ea · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Internvid: A large-scale video-text dataset for multimodal understanding and generation,

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.105861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.402384Z digest=sha256:928e45b6a6d508c7bde53410f29aa99d538a555810a31903d3eb10c68d171967

Observation d3ff9fb9-4325-4710-9bb2-44708319fb72 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Msr-vtt: A large video description dataset for bridging video and language,

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.097227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.406406Z digest=sha256:83c94b91a961b9285d99865f4e574f3f8ed18aa7963fc58a617e1bb4405bc4dd

Observation ed33fdc9-709d-48e4-aa0f-03205f982a49 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Frozen in time: A joint video and image encoder for end-to-end retrieval,

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.088314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.410344Z digest=sha256:5e701ce51b5bac78a71bea6e051767539b28f7346f8d4c1391504ff74fc933db

Observation 3cc1b1da-32d8-4335-bbf1-b9f1b8fc9346 · outbound

This paper cites Tgif: A new dataset and benchmark on animated gif description,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Tgif: A new dataset and benchmark on animated gif description,

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.079523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.414671Z digest=sha256:aba9d16baa7e461e181d1d9707594d9d8058e5e6f41dff324f96a567cbef6c5d

Observation 96c01551-6739-4012-8700-4e3227a3def4 · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open- world compositional text-to-image generation,.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation T2i-compbench: A comprehensive benchmark for open- world compositional text-to-image generation,

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:47.069121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:46:46.417659Z digest=sha256:ff9ac9a215cf5c2e4e3308ba9b5fd9d61dc253526079c9aaf91be94234737829

Observation 6b1a57a9-e7b9-4f94-9a8e-eff47ad23972 · outbound

This paper cites LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.420145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.420145Z digest=sha256:e31695a4e8e50fcb4c80aedd2eebb66311d921e36dbd542cd11b78efdca7d595

Pith citing papers

Observation ea652983-a7d5-42ec-beb5-76a4e9f08b94 · inbound

DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models cites this paper.

DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:47.490821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:47.490821Z digest=sha256:211b8c9be4bb3b52bca277e42ad6c9a4a35e2dc415a40f8551127517fe399cbe

Observation f03c7216-4506-48bb-a9b4-1fe2e7189b63 · inbound

PhyGround: Benchmarking Physical Reasoning in Generative World Models cites this paper.

PhyGround: Benchmarking Physical Reasoning in Generative World Models LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:21:28.543375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T05:17:30.010064Z digest=sha256:68825389d11a2b194605553fdd825fabe4681218398beaafda752d39ee4523d3

Observation 71868fec-ce45-41ca-801b-15d7a07607cc · inbound

LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models cites this paper.

LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:56:59.273921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-02T13:49:53.016586Z digest=sha256:840d3785229dc657184a1d4ec68dab9633d5025e185fc2b055f03b4e8c902d68

Observation 7de1b9a1-3abf-4470-b2b2-10d662cad8e8 · inbound

Natural Language Camera Movement Understanding cites this paper.

Natural Language Camera Movement Understanding LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T05:16:19.750976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:16:19.750976Z digest=sha256:ccfecfd2b54d2b409275a16eca85a3fe926c56a48ef69c3a4be2091e13e22f19