Pith. sign in

Paper Citation Record · LEDGER

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

As of 19 August 2026, this Paper Citation Record lists 100 of 130 outbound references and 22 inbound Pith citation observations for arXiv:2505.20292.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20292 v4

Coverage vector

measured 100 of 130 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:35.677086Z

measured 122 of 122 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:42:42.185508Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.694521Z

Reference resolution

100 of 130 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved99
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f622355a-fc11-46d9-8882-071363473815 · outbound

This paper cites GPT-4 Technical Report.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:27.582566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:27.582566Z digest=sha256:50b0c5af55cdfbbb322217315adb5fad3e698c0bea2b3373ac66e280e4e13a66

Observation 674e39e3-f51d-4ec1-b250-274525a49cf0 · outbound

This paper cites Detecting ai-generated images using vision transformers: A robust approach for safeguarding visual media integrity.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Detecting ai-generated images using vision transformers: A robust approach for safeguarding visual media integrity

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:27.677003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:27.677003Z digest=sha256:6afa1e3285f3ad2512f56c5d224b32fd849d2ea551fda47e9e7c2fe3b88a6dd8

Observation b31a7f79-6b63-43d9-bd3b-b1ec684f5ca5 · outbound

This paper cites Qwen2.5-VL Technical Report.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:27.789237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:27.789237Z digest=sha256:39b8b039d7b3abc7d55a2b92f274eb58ae7ac14909c598042c589520fa172ca6

Observation d4d2a682-d6fc-4019-872e-3d68462c74af · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:27.939740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:27.939740Z digest=sha256:a201b987d29a32c5390f18ed86824684a6eab128ae4bd0baf383052dbbf39b41

Observation e211f830-c275-4c09-812e-117d517cad8a · outbound

This paper cites Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.048350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.048350Z digest=sha256:20c1364213507a1dc5b7b7295a653f71ed3fb069f844a4f8b36488c29d136b4c

Observation c83fa217-46b4-4b85-a1b9-10cffa1791f9 · outbound

This paper cites an unresolved cited work.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.178541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.178541Z digest=sha256:c8b879a65a3eff16c8fae9dcf9fb97d4784a93772b33c1164e3a37e25e373116

Observation eba92d87-f4a5-49c3-8088-45ad15c42153 · outbound

This paper cites DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.322361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.322361Z digest=sha256:08ab529f0ce82117bc895ba4b05d8024fc263d7c401c47fd171f94efcb734f0a

Observation 2b2e282a-64fc-4df7-8bcc-652b8e8d3538 · outbound

This paper cites MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.549879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.549879Z digest=sha256:58a380623a6d9d920e1724eb5e25791afc7ebf5085ff8c1aa6697e122df501da

Observation 5041d6bd-3208-4592-b253-afb119d6efc5 · outbound

This paper cites PhotoVerse: Tuning-Free Image Customization with Text-to-Image Diffusion Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation PhotoVerse: Tuning-Free Image Customization with Text-to-Image Diffusion Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.605625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.605625Z digest=sha256:4f44b2bfdaad9f068c7a2777b6532c7d3b173bf3ffa96f48202f1fd9d244ead0

Observation a83ac1d4-5098-4588-8898-770acbd2eee6 · outbound

This paper cites OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.683440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.683440Z digest=sha256:d6ba3d4a28374cff0ba2254d06023f3c47d0ac54b90da47914392671ff79794d

Observation 9a4bf5ea-8a1e-4037-9f18-93bd8d8d0591 · outbound

This paper cites Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.739467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.739467Z digest=sha256:6ca7c7645183fcd7791f774e8cdf2e7d849ced0c495c7db5cbde2c2ccd006d01

Observation 08a4a1b2-29ee-4e32-b9c5-a14526531a4f · outbound

This paper cites Multi-subject Open-set Personalization in Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Multi-subject Open-set Personalization in Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.849584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.849584Z digest=sha256:a86a1dd6b4595315e1c614bae324d0d9fbddc962b065308d6b8e354a83ff32c0

Observation 32a1c487-0f0f-45d1-b6cb-035ad5f57224 · outbound

This paper cites UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.964637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.964637Z digest=sha256:0f74fd71db249cb9e647ccba2ceaca228a5bd3dbefa8a5199a45ab1daca41463

Observation 9b8bc10a-163f-4858-ab5a-41425d925ea7 · outbound

This paper cites Yolo- world: Real-time open-vocabulary object detection.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Yolo- world: Real-time open-vocabulary object detection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.040098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.040098Z digest=sha256:9f3807318180829100d87bddefde4c57ddef383085d872d2fb6811fe996ee6d8

Observation a6948c17-7919-4f42-9b3a-095e838483d6 · outbound

This paper cites improved-aesthetic-predictor.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation improved-aesthetic-predictor

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.159161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.159161Z digest=sha256:c98e09ca7f62a8f055dcaccd9aa585a4b9b03aceb280b72f1ff96ad46fe9f60f

Observation 8ea5b247-88ce-431e-86f4-3e07a6732a56 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Arcface: Additive angular margin loss for deep face recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.222747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.222747Z digest=sha256:d7c8d074272d2906005bf1b4782e7dc9d6d151accc3f127b1af7a8fdfeb03cad

Observation 7344fd96-2a37-4a56-9bff-dc864b12df8a · outbound

This paper cites CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.343261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.343261Z digest=sha256:acf66465529e0ca9112331bec15593755ba2c64611ab47990f2bfb0d4264ad07

Observation 5f763547-6771-4408-af08-5709abd1f81d · outbound

This paper cites Worldscore: A unified evaluation benchmark for world generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Worldscore: A unified evaluation benchmark for world generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.388166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.388166Z digest=sha256:6c9308879b1653563d2518b4d561b17fa7af86eb7072c9582d5693f9cca12b4a

Observation 78244fcb-96a4-4777-9e21-b732843a7d56 · outbound

This paper cites Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.452382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.452382Z digest=sha256:bbbfd1ab1a52bc0ea27a616bd26d280eee95075b69395d64793eedb877c99985

Observation 87b78b0f-fefc-47bf-a394-e5f5bcd6680b · outbound

This paper cites SkyReels-A2: Compose Anything in Video Diffusion Transformers.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.507056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.507056Z digest=sha256:a8d9c1f714ffc87849fcb0bc0caf05201ac75837588ad3d188ce287328e0472b

Observation 4038376a-8f8a-411c-83bd-87e66b361c56 · outbound

This paper cites Ingredients: Blending Custom Photos with Video Diffusion Transformers.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Ingredients: Blending Custom Photos with Video Diffusion Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.585694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.585694Z digest=sha256:e8427426339b7033311db01e691328d8ea8c2644b3c6676e19131e314b902bd1

Observation 7e1227c4-7833-4794-8ce6-f823e278d27e · outbound

This paper cites AE-NeRF: Augmenting Event-Based Neural Radiance Fields for Non-ideal Conditions and Larger Scene.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation AE-NeRF: Augmenting Event-Based Neural Radiance Fields for Non-ideal Conditions and Larger Scene

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.672932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.672932Z digest=sha256:aa618ec436ee59bb57fcc701c4f489b9d66d2f7736fcdb192f32bec97d491315

Observation 0bcb0d86-c873-465d-9046-9c71c1b4ca33 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.775085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.775085Z digest=sha256:ffeeb72da89263ecfa391b9b135f5574b3367281c632830daeaf1d3394f3c21a

Observation c460b892-36fd-4c44-9c27-2e96a2b65ab3 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.829842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.829842Z digest=sha256:3566561762dac1b8dcec3f694f87d1f319d40016a1f833d32fc317b529673c81

Observation cef193aa-9fac-4a73-9318-e231dbffa46f · outbound

This paper cites PuLID: Pure and Lightning ID Customization via Contrastive Alignment.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation PuLID: Pure and Lightning ID Customization via Contrastive Alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.869968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.869968Z digest=sha256:7c515a4c67ceeb24de53d63d703c6e015677791b97419e7cd19ac78ffe3dda1f

Observation b9ca7979-c9c2-4e8a-8da3-5d1cf9e33c7f · outbound

This paper cites UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.959239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.959239Z digest=sha256:36bc8dbda7a3ede735e6d4fcb5ef6641e759932435887e499286e7f207477be3

Observation 2a96e897-15fa-4f12-9e15-faf72dbb4d0b · outbound

This paper cites VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.049247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.049247Z digest=sha256:c0d1cd939e9fbb4982edce167105d60af866e8a73d2d9a7270aeac2d62eb69e5

Observation 7b23639f-4312-4f64-a82f-555e6e29f70c · outbound

This paper cites ID-Animator: Zero-Shot Identity-Preserving Human Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation ID-Animator: Zero-Shot Identity-Preserving Human Video Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.179985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.179985Z digest=sha256:474527dfe1ff715fe486c120546ebc510c680c77157a5a410238617e8ecd2bc7

Observation 79771ac8-067f-45f3-b446-1b6261545c69 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation LoRA: Low-Rank Adaptation of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.262848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.262848Z digest=sha256:675a8f42cc4cbfb622135c11a3a073563724dd5270650c476a89687037600a2d

Observation b14917ee-acdd-488e-a78f-8696687ca2b8 · outbound

This paper cites Animate anyone: Consistent and controllable image-to-video synthesis for character animation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Animate anyone: Consistent and controllable image-to-video synthesis for character animation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.325203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.325203Z digest=sha256:06b58d8a5f759604d1de64097e2f17e0b656b8e3cb74d11e059076cd1893156b

Observation 3f861033-8931-486c-b7df-ae3fe9cb8860 · outbound

This paper cites Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.416393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.416393Z digest=sha256:307b18ddbaf95c9db86db75c559666b64bb575b002f7b92329be62b07634865c

Observation 63e9d5a2-120e-4bf2-8687-2cb8e04ecf70 · outbound

This paper cites HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.485928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.485928Z digest=sha256:608edac5579d7d7855f3b003088aa637918f931e79b9e0bf32d729740155b36e

Observation d2626035-6181-445e-b482-e4885d7f8b11 · outbound

This paper cites Curricularface: adaptive curriculum learning loss for deep face recognition.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Curricularface: adaptive curriculum learning loss for deep face recognition

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.553705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.553705Z digest=sha256:519c0f145af9e7e33a8dfed3684cf73bde609272943243a9ec0aeebbe0b1ced5

Observation ed245bbb-b229-4733-8b8c-c399a21ac318 · outbound

This paper cites ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.608834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.608834Z digest=sha256:fae78a96ad1a8d47634f454571e3f0e889d58fe7c68791951b45922361033b4f

Observation 65e3c255-602b-4982-bb4b-57471c45dd2c · outbound

This paper cites VBench: Comprehensive Benchmark Suite for Video Generative Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation VBench: Comprehensive Benchmark Suite for Video Generative Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.690138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.690138Z digest=sha256:212337a287e5d673cdb1a4f74e789bb04b02c03c15191556b3806c669b6aa2c6

Observation 288b44fb-26ab-4086-8dd5-595bdb0df7fe · outbound

This paper cites VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.771224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.771224Z digest=sha256:3e721f421bd66650bbd4948079cd455cd777bbf6ede1ad2764904338580ae6f9

Observation 0baf8290-2faa-4be9-bc22-e5438a041216 · outbound

This paper cites InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.860050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.860050Z digest=sha256:ec8ada9c8886ec819358f63cce14532159f0ac39c5d199104f403ceb0c9c4c48

Observation 7d0df9c7-c0c8-473b-a301-0bf36b3a1813 · outbound

This paper cites Videobooth: Diffusion-based video generation with image prompts.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Videobooth: Diffusion-based video generation with image prompts

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.937726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.937726Z digest=sha256:635a2e5bde5a1c3643962162e25e4aba06fdd13a96ff2023a4d7ee46d782788b

Observation 94ca3e3a-605d-4741-ace1-aa094a4bcc71 · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation VACE: All-in-One Video Creation and Editing

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.975711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.975711Z digest=sha256:01d7e6e6714af2fb8093fd4917de80be8a49b5bfe23ef8d9600b13bb019d6d1b

Observation d40dcba6-b12e-407f-8fee-4f445c5d4eda · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.009125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.009125Z digest=sha256:e4c994fb608f8941dda677bfce13071c5e8f9cdc534faea565ffdd9e75130e64

Observation a9f50fd2-8410-4887-94a1-fddec84a7c47 · outbound

This paper cites Subjective-Aligned Dataset and Metric for Text-to-Video Quality Assessment.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Subjective-Aligned Dataset and Metric for Text-to-Video Quality Assessment

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.052527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.052527Z digest=sha256:9e96f155fb794716bbd02556c98dacbb779b86e3ada002163d659133c803e539

Observation 32bfaf9c-8b5c-4e53-b2cc-80377defbc32 · outbound

This paper cites an unresolved cited work.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Unresolved cited work

Reference 43

Resolution
parse uncertain
no resolver link, observed 2026-08-07T13:59:31.099118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.099118Z digest=sha256:fc9c6eb7e91f30dd7c2e361cd592abc07fd07bb06ff0a986739facb35dd1e09f

Observation 75367925-d66c-4aca-915c-9a298cb79360 · outbound

This paper cites Pika-2.0 lab discord server.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Pika-2.0 lab discord server

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.162597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.162597Z digest=sha256:657ece2f6451ccbc0d6ca798540fbe2576119d7d2c752ddae80752ed6883abd5

Observation 6233e01f-9820-43fc-bfbf-9768baece134 · outbound

This paper cites OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.215680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.215680Z digest=sha256:c41591724bc56fcc0a6f96c2d6e1e10956ba50c6f651beeea7c8a3960a1e9f02

Observation c8fa7d41-df9b-48aa-b5ff-6f05e963ecee · outbound

This paper cites Improving Synthetic Image Detection Towards Generalization: An Image Transformation Perspective.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Improving Synthetic Image Detection Towards Generalization: An Image Transformation Perspective

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.251492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.251492Z digest=sha256:1cbd5dcd709f2e26181d370ea03ee2c5ffcd33d4781b8231edf5b49ba093a9cc

Observation a965af12-e251-4d03-a94f-103a3ec054b3 · outbound

This paper cites Photomaker: Customizing realistic human photos via stacked id embedding.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Photomaker: Customizing realistic human photos via stacked id embedding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.325770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.325770Z digest=sha256:e301376c45cff8c1c0248e174cee49fb349630e809a2c4ff5bd52f94142a4249

Observation f70ac60e-5534-4725-a8dd-f8827e1945cd · outbound

This paper cites WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.446588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.446588Z digest=sha256:352607242a057176311941baab329861611c9b18bf9b76570989ac2387250482

Observation cd2fdf07-8d2a-4136-9e97-f530416fa399 · outbound

This paper cites Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.555321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.555321Z digest=sha256:11d35b3e84f4a7be24255fc463669a3881fd42f80dce1e910eb34e52f7833891

Observation afbf14cf-615b-453b-a715-eba9e447d2ff · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Open-Sora Plan: Open-Source Large Video Generation Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.643439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.643439Z digest=sha256:c708903c0be88f0c073f29415ef6e348bd7d9541a55fd9e1b0caff5d071e3b90

Observation 646f6a09-7bce-4ea8-92fc-7faecb7daab8 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Video-llava: Learning united visual representation by alignment before projection

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.756572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.756572Z digest=sha256:1eb895cc4369d21779f8f132786a753e1a1c56ed2e86d89789c4e12a55469e8f

Observation 35763072-1ad3-4e7f-b05d-d4b2202915bb · outbound

This paper cites Stiv: Scalable text and image condi- tioned video generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Stiv: Scalable text and image condi- tioned video generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.827870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.827870Z digest=sha256:de57b2bb1f329778bd7cfe6021e33a3f957ff6e1e543116c49e646186f8afe38

Observation af6623f3-3f69-48f3-a4b0-9ef802ea327f · outbound

This paper cites DeepSeek-V3 Technical Report.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation DeepSeek-V3 Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.890570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.890570Z digest=sha256:8a5dd0d84a8f34270a7801eac929f76aa9362a79cb05577e51a7205c86caad10

Observation 0f02fd0d-b56c-4bce-8c2d-c8eaff2fd744 · outbound

This paper cites Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.958999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.958999Z digest=sha256:89f9a74fc2a32985d158e3f99c917fc5371c66e1a63a080d6ec908976031c05d

Observation 8b4584ab-6c87-4bf6-a062-9600001f2d39 · outbound

This paper cites Phantom: Subject-consistent video generation via cross-modal alignment.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.071589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.071589Z digest=sha256:c581dc67848ac90defe7d821b338c0b79d4952079dde4d3468c7239418384297

Observation 22e8865e-7ae2-4d68-8efe-989ccfa59e3b · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.143783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.143783Z digest=sha256:5faf14768a6b62a47696380612485c1d2b5e270bc79ce81a5117e922e6d222f8

Observation 13596194-ebe1-4953-8804-bda4c428597d · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Evalcrafter: Benchmarking and evaluating large video generation models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.195883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.195883Z digest=sha256:edd82482cb52055bb91377904760db0322783d14c0742283c219ce39b6f375b9

Observation f69d7047-7c83-44c7-b1eb-3f415ea0b1a7 · outbound

This paper cites Fetv: A benchmark for fine-grained evaluation of open-domain text-to-video genera- tion.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Fetv: A benchmark for fine-grained evaluation of open-domain text-to-video genera- tion

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.278537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.278537Z digest=sha256:603ed6187750021c30dcbb42a71261aadc64026069ec4af8edc53c23f69c2148

Observation e992ba3e-adde-423e-a52a-b972d2a2508f · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.327532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.327532Z digest=sha256:04214063032782be3e2e40d5a23c2d8e24a626dfdc7a1cb73bb869eeeb65aa02

Observation 985742aa-2d65-40e4-9bbc-f7c3a5bba31b · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Latte: Latent Diffusion Transformer for Video Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.388980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.388980Z digest=sha256:7706a8b5e5efa4be4188b8c2150b3aa1c2a4d6bbaf625bc70d865007d2dc093f

Observation 79dc3038-4365-4037-a938-17f2ca0da233 · outbound

This paper cites Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.433119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.433119Z digest=sha256:736cba213a10b5cb81e70e964c3ff3db03e94744372213ec449363523b293bce

Observation ddf5a28d-2ff5-4c6f-b57f-cd36ecb9604c · outbound

This paper cites Follow your pose: Pose-guided text-to-video generation using pose-free videos.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Follow your pose: Pose-guided text-to-video generation using pose-free videos

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.518451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.518451Z digest=sha256:16f5860e2a883e98b457281dac76659cac934a94c6f0b4ab6d13a491fcb5919b

Observation 59b261e3-a9e1-42ac-98d1-9da35f05c4fe · outbound

This paper cites Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.625729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.625729Z digest=sha256:368a8d07994f578b77531b75914ee0e5211e56c24c415c579108a75879272379

Observation 085f1cfa-a5dc-4b84-9ce5-37a9cab5b220 · outbound

This paper cites Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.721019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.721019Z digest=sha256:fb20fb2dbffdac432a3cbbe53ffbe1ed7c163e0336f8cdacab435f17d67f8d2a

Observation c40ac82b-7a53-4239-94a5-0a2437246661 · outbound

This paper cites Magic-Me: Identity-Specific Video Customized Diffusion.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Magic-Me: Identity-Specific Video Customized Diffusion

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.790081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.790081Z digest=sha256:40d9f8a5383a0998bf0f91d83f180eccc0094b61ca33088599b5da774bf1acab

Observation ac1eba25-0a75-4758-af1e-bdaaab115766 · outbound

This paper cites Multi-task image classifier.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Multi-task image classifier

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.885473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.885473Z digest=sha256:fd341615bf5b2bd51ada95b4df3074c6389708bbad9af084e8e4f9c6bbece8d3

Observation 37a1e504-5fc6-457a-8d8c-cc8d9be935db · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.952192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.952192Z digest=sha256:074986802716abe704a40f85d2ca32c832c2b468e881a048d9f540e01248bb86

Observation 0cedfbba-bee5-45cd-9865-a6bc31c75e8a · outbound

This paper cites Nyuad ai generated images detector.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Nyuad ai generated images detector

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.006913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.006913Z digest=sha256:35689268841ead256c712a0e82ff80434bf8892b567c8ba010633091a78df4e0

Observation 1ff826d8-9563-40d7-af7f-753049c77cbf · outbound

This paper cites DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.042872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.042872Z digest=sha256:20a016aeb4f3c5109dad4032240d3f40fab52017553d93b9d37a8585df23d1d5

Observation 9f51e9bb-64fb-4fe8-bbe2-17849e2b1c54 · outbound

This paper cites Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.158196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.158196Z digest=sha256:8316a9cf00e1d233af8574383065dec77f5c62d52c7be4562d7dcf4d6e4ed42f

Observation 629342a2-e5a4-4261-bff7-42711df93a9c · outbound

This paper cites DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.247451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.247451Z digest=sha256:229ba61530e8ab754911ac440986680b7f7dd2e3ae89dcd8fd759ca0ea745d5f

Observation 4ed2f581-5ca5-4541-af1a-e79b4c95c62f · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.328258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.328258Z digest=sha256:54681c21182c08b070550f218b3fc73da3172870bf03fb77773b93f3762bb6a8

Observation 39695eb0-95dd-43df-a774-b98b2a9a1723 · outbound

This paper cites Learning transferable visual models from natural language supervision.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Learning transferable visual models from natural language supervision

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.385897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.385897Z digest=sha256:7407e8e95cf3c49251fcffac2e65e766e7c52aa990fecec3406a5ef017546b2f

Observation 80e7fd13-6dc1-40d0-8d2f-1a716c9df315 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.449820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.449820Z digest=sha256:e8e471c0acbf5ff5e89770f40fabfffbd498dedaef875dfc2fd44e9ccd21cfce

Observation 33f3b505-4ad7-4a06-b3fd-bcf031b21a24 · outbound

This paper cites Zero-shot text-to-image generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Zero-shot text-to-image generation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.541902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.541902Z digest=sha256:2eaf30dbfd1c9958b5cc3995079ce93ccf4c272dd368aa0bac5c65493a016d53

Observation 1281ad9c-c073-4665-95da-e2ca2f3f52e2 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation SAM 2: Segment Anything in Images and Videos

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.625427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.625427Z digest=sha256:500c8ee832b49d0add5fb763edd47d243d8be1c2344570441e7eef9b4308aca9

Observation fc7c93b7-3ee7-48ad-8d46-69a0dcb28789 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.744844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.744844Z digest=sha256:d60765742efb22e4d099e7c45bf022d70b3e98973b6e89da29c777ae10394324

Observation a2501434-08d2-465e-b527-0a1c29caccfe · outbound

This paper cites Magi-1: Autoregressive video generation at scale, 2025.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Magi-1: Autoregressive video generation at scale, 2025

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.801156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.801156Z digest=sha256:18e014e40b7454fb7a8100a4dc1f2b12c35c6ef0362f067f01494fb997499636

Observation 88a9b364-495f-4981-a93e-fb0067626203 · outbound

This paper cites Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.869824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.869824Z digest=sha256:d774a16c6103a83f6f9cb6a70e9d9da71c76205c085a875bedeff1f9c8fb41dc

Observation a22d91a2-71e3-43d3-a74f-9fde8ace6b8c · outbound

This paper cites Dragdiffusion: Harnessing diffusion models for interactive point-based image editing.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Dragdiffusion: Harnessing diffusion models for interactive point-based image editing

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.939322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.939322Z digest=sha256:bc66e313f9243271a8141a4d4acbd8b8eba7b1ff537186072deb0d7ef2363766

Observation c6fdb347-8410-49d7-a3f4-87eebfbe09c0 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.017512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.017512Z digest=sha256:49e245781f2b4f3ce0d509d6450c0106e0f65d3b39ce6e90196ba5a2c2b5ac2e

Observation d490f859-6ccc-4e52-afd1-bf1382323b92 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.089977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.089977Z digest=sha256:9743baf99c0ffb7f2f97179ec3e874a9732493dfd949c9cd758169651a1e5b43

Observation b35739bc-9d62-4f04-bc2a-0f1c5a8ea958 · outbound

This paper cites Animate-X: Universal Character Image Animation with Enhanced Motion Representation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Animate-X: Universal Character Image Animation with Enhanced Motion Representation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.211475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.211475Z digest=sha256:cf7207b2144925372f3df84c6200d47d78b060c0c30205005c2198449d88c4ed

Observation cc4fa633-46fd-49cd-8158-098ff232e83d · outbound

This paper cites Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.324476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.324476Z digest=sha256:f05151c59a7c1ef704cd1ef4ed277c4f14186b3d3500488e6a232dee461ae866

Observation 71052253-bbdd-4a14-ae89-66771b695258 · outbound

This paper cites Gemma 3 Technical Report.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Gemma 3 Technical Report

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.446500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.446500Z digest=sha256:71c0d90180221cff2d2c145eb251344b0e1c76ba69dce312b1bf2d6584ba7629

Observation c79d273d-30ad-4186-bafe-ebb6ac539c9f · outbound

This paper cites an unresolved cited work.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.550294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.550294Z digest=sha256:b3daafeaeefd6a7c4d8754484c975a14f8197f94daf13feb6f6eef50a3f06680

Observation 3e55a3bf-075d-4375-a7ec-717765318ab1 · outbound

This paper cites an unresolved cited work.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Unresolved cited work

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.646831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.646831Z digest=sha256:1cea73fe2f9c431b2bc939ba01903b71bea56d4af824e315219ec4e67c3910c8

Observation 6f59a7a3-5f8d-48b7-b9e1-e0984804100c · outbound

This paper cites Paddleocr.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Paddleocr

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.766190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.766190Z digest=sha256:5c5ddae55b4c8537b38862cea481b4426e14cfa64d67bc43cf571742f6da9402

Observation fa802251-b8ba-4cdf-a54d-ce8af139e7e7 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.883459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.883459Z digest=sha256:8be7ab0e131034d8ff347ac8136764fe0b7f0c220981982527400eb6a987e1f4

Observation 6e3c765a-189b-4dfd-b82d-eaca53f409bd · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.953854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.953854Z digest=sha256:13fd956f206922752a5d588a999a3d773836eceb84ace665c37d0fb2faa38e55

Observation 42a426d6-43e7-472f-a8c0-8647655e4834 · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.043662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.043662Z digest=sha256:12e1ba5b4af55f9615a3f8b6a36af17d24ab7a51be5e88bcae63915713eae8d7

Observation 956c6d0d-814c-4d06-b946-94e4a5fb6d59 · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.115504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.115504Z digest=sha256:9cc341ce2b2c627cbd204d47fa2d02a442e1554d52f5f8744d970d980edc533f

Observation 4fd5de33-cbb1-4744-9e07-f6a94bc67674 · outbound

This paper cites VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.179016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.179016Z digest=sha256:71152ffee3260074794f9f539bd1b041b4c89bd03c96cf9da00d7a5ab0dd1c29

Observation c5788791-343a-4a42-b610-91b72e515481 · outbound

This paper cites Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.266101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.266101Z digest=sha256:c654e369c4d252d1164bcb11e6c6a469d0e80eb4c48554f4fb0f2b5fd953cfc4

Observation 8daffc84-86b1-4a83-8ec4-10a2186913ec · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.298215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.298215Z digest=sha256:59e8c3e520e2787a58270457f50530f84a2dbcdc22ff7b039968f102b7355107

Observation 27bd54d9-7a9b-4961-a550-0366d6b1b154 · outbound

This paper cites Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.359517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.359517Z digest=sha256:375d8647d86e7dbbc381c222e3db636dc6c0ad7f18eb01e259a5bdadefeaed99

Observation 2e7d5a2d-e248-40f9-abe4-65937c05c91b · outbound

This paper cites EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.395109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.395109Z digest=sha256:b0de9d4305c609bd8946988f1d97cb468a70c52010cc4470102e7b8f355778a0

Observation 7119b8e0-fdfa-4aae-aab3-d8766f96a204 · outbound

This paper cites Dreamvideo: Composing your dream videos with customized subject and motion.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Dreamvideo: Composing your dream videos with customized subject and motion

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.430457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.430457Z digest=sha256:a01b87fc5d46bcd1d8ca972f77f525793a4bb668cff1fe4c14bdcd9ac89cc50f

Observation e1b967fb-9f0c-44cd-b544-e8908a293dc7 · outbound

This paper cites Exploring video quality assessment on user generated contents from aesthetic and technical perspectives.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Exploring video quality assessment on user generated contents from aesthetic and technical perspectives

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.489028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.489028Z digest=sha256:a13ce6ec1f43de1a40e7c50bda22014667c70a157d8a30836bd0a90b0d51c507

Observation c48c9115-7c42-4700-92cb-3b8a15840695 · outbound

This paper cites Towards A Better Metric for Text-to-Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Towards A Better Metric for Text-to-Video Generation

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.615058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.615058Z digest=sha256:eb73b190263a5e9b1b8ca6ee8550838147441e756029563233947cc95ec9cd6c

Observation 0b5ae721-ba87-40ff-bab8-71e9c71c1738 · outbound

This paper cites MotionBooth: Motion-Aware Customized Text-to-Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation MotionBooth: Motion-Aware Customized Text-to-Video Generation

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.677086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.677086Z digest=sha256:0bd144f956f70a39a9a59047fc2d34498f43ad65125a01466adb600abe4e70a2

Pith citing papers

Observation 3aca3f54-c5c7-42cd-9089-7b36a2819a38 · inbound

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute cites this paper.

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:51:54.359831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T17:51:41.947939Z digest=sha256:b6c070179f4fef29c0ee29cfe002d1e048c50d8f3e178fad3de868b11e20250e

Observation c203b860-2132-44a4-b42f-76f32dfc4acd · inbound

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation cites this paper.

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:34:27.064326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T17:34:26.951644Z digest=sha256:cbe0dbe34acf63adbb8928cd62e03419555e644d9c10c06da42fe98b54a703ab

Observation 1f147232-5df2-4fc5-99ef-ccfb5336fef7 · inbound

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement cites this paper.

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T12:42:42.185508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:42:42.185508Z digest=sha256:299fd0c95b4c2a71206d850a99ca88453be3cdb6e479a2b23f611d6d842b64a0

Observation df606aab-91b1-478b-acf8-815f5d58b22b · inbound

Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation cites this paper.

Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T12:23:05.876881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:23:05.876881Z digest=sha256:4d435b842fdadc951766b094039fbdc81dd92b485a267cee2f99589aba3c3f78

Observation 47eeac02-5dbc-4b30-b9c4-37208fa942b8 · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.839751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:7b57368774b194b467802e2e9bbeab3a36dc5baaba811707c3706d1ebd0fd77e

Observation ca89bad4-82c4-42aa-8181-608e8e4bff9b · inbound

LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing cites this paper.

LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:21:55.095118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T07:21:41.483427Z digest=sha256:2346544e712cf1a45dd70f43931e7dc11284bd04cd8134b10863e2ac280cb12b

Observation 146ce315-82b4-4402-80e4-5c6faf90d9aa · inbound

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation cites this paper.

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:11:19.210744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T06:28:42.129881Z digest=sha256:583dfdf1f0b715082d03b59e8bb4078abc6bb7a0661e1317f45aa5c72bc0c0e4

Observation 3c22f376-4c71-479b-ae69-5a2cacaa96a2 · inbound

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation cites this paper.

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:51:15.130701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T00:50:10.509727Z digest=sha256:c1bf2f02de9b3f9a025c384687b35d4487a06ce4007ef0ab53a57d8e93bea6b9

Observation 3b2377c7-e313-49fd-a656-7d491fbca3fb · inbound

EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation cites this paper.

EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:44.040337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T03:16:03.450742Z digest=sha256:139b3eaed8bf4cd5bb7832290273796700b2ff6c2c530e8e83bfcc3262cc5501

Observation 6884ece6-c3fc-4215-9ccc-eaebec221846 · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.506965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T20:11:33.277349Z digest=sha256:c1a470aeda047c8d188a6474536da75cc541e02cea2db773521f5ea79ad14578

Observation d0708f02-1251-4c41-98d5-c8e04e91a0cc · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:25:00.604747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T19:23:07.044659Z digest=sha256:10b9ec4f7cddade9c21c61c85a9cd4d7a701cdfdcb518fb0fcaf8b3452b859d9

Observation 3f51a193-c34a-4d5e-9ebd-0abde07ca1f5 · inbound

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation cites this paper.

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.404522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T14:08:30.802619Z digest=sha256:14329edfbcc59194e3ff0e0e623d2002504532d0eb7ce89af5432ec180cc954d

Observation ec6a5880-b1f9-4524-8903-030cab407046 · inbound

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation cites this paper.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:18:03.111547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T05:17:11.690484Z digest=sha256:57ca7a9863b70f8fde4e3aaf634fed289c6869bed68a89fe5dd5811e216b063e

Observation 5717e354-4435-4328-84c4-ea7fe842b65e · inbound

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation cites this paper.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.012646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:cb3fff3aa6a078ffdcee8333fb5cdcb759d1c53c8ca9bb1bb2db3f311f84e304

Observation cd359f9b-2198-495b-99ac-3411e362971e · inbound

Bernini: Latent Semantic Planning for Video Diffusion cites this paper.

Bernini: Latent Semantic Planning for Video Diffusion OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:41:10.560161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T06:39:47.124605Z digest=sha256:4e565f1a47269c37f1a6dd81234b2e59ea9fc4d82bdbb03441b120984c0640c5

Observation 4b4fd2de-3168-4ea7-a51a-7561b92f26d8 · inbound

Tail-Aware HiFloat4: W4A4 Post-Training Quantization for Wan2.2 cites this paper.

Tail-Aware HiFloat4: W4A4 Post-Training Quantization for Wan2.2 OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.628549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T18:20:47.823852Z digest=sha256:ca484c37190da5d1951d041b775950035fd11f71a81d4b68729ec6993087167a

Observation e4cabb42-3932-46e4-9b28-ebaf7eacccd8 · inbound

Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models cites this paper.

Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.320535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T18:58:13.799615Z digest=sha256:f15aaeb232a540d98de1d0ffd72dd16997a23bf13539ec1383ebb89b8aa2fac1

Observation a615470e-d2b6-4529-8694-59cd4b3071a9 · inbound

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation cites this paper.

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:37:36.638712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T13:49:50.272650Z digest=sha256:e0a5819926728102831f398a0f3224c5ccf9266e581236dddc05ad242b501a6c

Observation 90ac3907-af02-4025-9542-4cd1e31b21b3 · inbound

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation cites this paper.

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:09.696509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-25T19:00:23.260939Z digest=sha256:c33ecfe76b292b555bea26e16f056b1080cf94234d58c1017210a35be8584f56

Observation 415ecded-5571-409f-9c9f-beb34b30e50a · inbound

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment cites this paper.

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T20:11:31.576642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:11:31.576642Z digest=sha256:71ea1d77660a22134f37e8c243ae8feeace1b3a264dcecd258e5d3a452215317

Observation 276855ba-c617-4634-a6d8-21ed7da399bb · inbound

Vera: Identity-Faithful Human Subject-to-Video Generation cites this paper.

Vera: Identity-Faithful Human Subject-to-Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T10:24:52.920746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:24:52.920746Z digest=sha256:bba1a7180dae3c5af6480b1d84ce9c3a6e4e4f07e982f38da43b0212940e71a0

Observation 1fdba910-7052-43bc-a0ce-292612471ffe · inbound

RefCaptioner: Multi-Reference Image-Grounded Video Captioning cites this paper.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.144127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.144127Z digest=sha256:86a64288fb1164d03995cea2105841a371a813c956d5f5dd39b8361e74fc405a