Pith. sign in

Paper Citation Record · LEDGER

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models

As of 10 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2502.06812.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06812 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:26:17.427689Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:55:02.514189Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T00:14:04.743463Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8c5a285f-8ac6-42b5-8765-f6c555b337b3 · outbound

This paper cites write newline.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.243081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.243081Z digest=sha256:b0b9e1dfcdfe35af3dbb5245df996684c60adfbd4a1c64e4aa0b78efe0982322

Observation 3fbd6934-8faf-4eb0-b755-3fdc48153444 · outbound

This paper cites URL https://pictory.ai/blog.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models URL https://pictory.ai/blog

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.157160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.248155Z digest=sha256:c0316dc9f9055ca7cf2b6fb238e131820ac2495a96121bb4927947ec0933c01f

Observation 8b75cea8-a0a0-4062-8e95-6c007e4a12b1 · outbound

This paper cites URL https://runwayml.com/.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models URL https://runwayml.com/

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.146504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.253340Z digest=sha256:c2824fafac279bdc78cfb1f5adbf5b3f4a596cc37cfc47d053da8434a71c71d6

Observation aedda03a-87d7-4a68-ab79-768bf9587477 · outbound

This paper cites URL https://www.synthesia.io/tools/ai-video-editor.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models URL https://www.synthesia.io/tools/ai-video-editor

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.135495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.257049Z digest=sha256:432cd250fe5778899890a252f7dc970e09832a4b632f75ca1f76342fc44a7013

Observation efeb00fe-29df-4a0d-8be1-5a03cdde91c2 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.125190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.260790Z digest=sha256:8c128bb132f687d6f9313b7f16d8947a9c9a809e4c62b832552cff860628204e

Observation 5df60e09-8171-4268-82fe-fb82643f38ca · outbound

This paper cites Training diffusion models with reinforcement learning.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Training diffusion models with reinforcement learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.114384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.264626Z digest=sha256:a0e3909613886b9fcd9b81c1e614456334f1982fe0cf8c2f79efd11c361ab7f1

Observation de1572f6-76a2-441a-be2d-add4be29ed80 · outbound

This paper cites W., Fidler, S., and Kreis, K.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models W., Fidler, S., and Kreis, K

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.102381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.268305Z digest=sha256:7a553fbcce942494c14b3c7856a19cfbe09c17198e68a7d1c3d2cbf1ec47cccf

Observation dc14affc-bd09-4e31-b33f-bc07333ecdc7 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.271948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.271948Z digest=sha256:be8751bfdda189db9d7e39386000cf40b6d01b1c7cd4f7d65d8cc5b1b15487f7

Observation e2fed0ae-baa9-4c8f-b9b3-22ac36d071b5 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.091279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.275902Z digest=sha256:a3a3eaf79b5b5f67139971f3c4f8ecd596fe4b787f48b5915fecaef6b03b3f45

Observation fde46d5c-986b-4c56-ab52-c3ff07c40c11 · outbound

This paper cites S., Brox, T., and Ronneberger, O.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models S., Brox, T., and Ronneberger, O

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.080150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.279374Z digest=sha256:43194015b347656542afeb36169402f98fd8c18fc232b69c4253682bdd759d8f

Observation 1bac0149-6a1c-49f8-8c9d-6579e6e16b07 · outbound

This paper cites an unresolved cited work.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:26:18.069161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.283020Z digest=sha256:02b4070379dd53b51aa2d72459bc735dad8a923f359bd823168620c579e1a427

Observation ced1d457-8368-476b-af66-8257e700416c · outbound

This paper cites DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.286546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.286546Z digest=sha256:cea330190e0fd6dfbecc743e29e85830d338f8454a3f174077dba094fcfc6cd1

Observation 7245801a-4267-4d2f-8344-a3abe0d74e8a · outbound

This paper cites D., Ni, Y., Lyu, B., Narsupalli, Y., Fan, R., Lyu, Z., Lin, B.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models D., Ni, Y., Lyu, B., Narsupalli, Y., Fan, R., Lyu, Z., Lin, B

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.058047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.290440Z digest=sha256:3afceaf3459a99d4251524f3774fc4d25c5ce794f5a34978e5eea6b096378fd5

Observation 12601fad-88be-4d2b-b42d-9f27ac66d592 · outbound

This paper cites Denoising diffusion probabilistic models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Denoising diffusion probabilistic models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.293856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.293856Z digest=sha256:1576bccf3cfbad8dc7a57d65bb107456f31cc2b9fd53b320f9c253d2cc6bc1c9

Observation 5ef611e7-864f-491f-8fe7-6842261106a1 · outbound

This paper cites Video Diffusion Models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Video Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.297240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.297240Z digest=sha256:ab61ae82757b5f5df2d5bf2ad20485ff90e22ea339d4bba455a3229fb467c5fc

Observation 01ef9893-5173-45df-a4f6-c633fc7e7f41 · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to-video generation via transformers.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Cogvideo: Large-scale pretraining for text-to-video generation via transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.300911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.300911Z digest=sha256:1bd8b41fa7c79a28ac251fb6b819e4235d7f9232f7695b081d841236fabaa049

Observation 41c20b07-a1e8-4943-aedf-ea363127e0ab · outbound

This paper cites J., Shen, Y., Wallis, P., Allen - Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models J., Shen, Y., Wallis, P., Allen - Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.304046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.304046Z digest=sha256:a2e0e7946a186a43e3967fd60560197672b49b50243d9dba41738155e1057eb9

Observation 1c883af6-977d-49db-b3e8-a0d5e6802840 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Vbench: Comprehensive benchmark suite for video generative models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.027581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.307223Z digest=sha256:20e0cd7d7cd7a85366da52ea1786403e14d3cc479bb1f21da42ef5c5cd45f329

Observation 13f1cb82-f460-4ad9-a828-d0af78269364 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.310948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.310948Z digest=sha256:fc682ec20f242df3a36401cf5d5fab360937a36578f4017d40f47ec99143b33f

Observation e90345d2-d1b0-462c-9b14-49de74ab4ccc · outbound

This paper cites Text2video-zero: Text-to-image diffusion models are zero-shot video generators.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Text2video-zero: Text-to-image diffusion models are zero-shot video generators

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.015879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.314581Z digest=sha256:ac3ca0166023dfb348b6a41e20ab907a1543223acc057fec92f4afc052bf6a23

Observation bf72edf6-0785-4f72-8c9e-10a2e351ee92 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.317958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.317958Z digest=sha256:158f5ce2d3979a251aa32b6be81d534fa9aaf14ed3ee2946ecb71eacc16c1330

Observation 290e26b9-8ff6-43c0-ada4-d2130980bf9c · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.321272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.321272Z digest=sha256:aa303f26de0ecde40fa300098d484a26d10cc8ef8bf745aff5414b32e024c4fb

Observation baf6ea88-7f3d-473c-98db-8daeaf0af1a1 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Aligning Text-to-Image Models using Human Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.324956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.324956Z digest=sha256:cb1f789658ecfa5500940dca8351d585c623a55731f22a845961f1704cdb25e1

Observation 9bc82d02-afd8-4e98-b2cb-a35855234460 · outbound

This paper cites T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.328708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.328708Z digest=sha256:39d963a44e28d172ff1e7458b393c4d0567f3f3a1d8ae72dc995e866bd9179c9

Observation 9579da8f-f846-4c91-84c6-1b1559d92b32 · outbound

This paper cites an unresolved cited work.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.332358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.332358Z digest=sha256:4862a4afecaf35adc4d25d908269dc50a0aaf74fb75064780852544c6dacecc8

Observation 80cabdce-42f1-4842-bbbe-f9b87b4f4220 · outbound

This paper cites Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.335567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.335567Z digest=sha256:c80e12c21f60a9e8036575b6cb1c52ad8867717d89dd511007e7c77e284c764f

Observation e20643f3-bd49-4963-8475-adfa24984a24 · outbound

This paper cites Training diffusion models towards diverse image generation with reinforcement learning.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Training diffusion models towards diverse image generation with reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.998359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.339251Z digest=sha256:bea21baab58841e332fd4cd7b9a47b721f52dab8a6e7f3f2f178cae11208c10c

Observation 092329e7-6779-4fc7-b795-96716892445d · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.342550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.342550Z digest=sha256:0955464bc3b304d90f60c8db85a5419c944cd286513d3b8ab4b333565705b375

Observation c94f1892-48bb-460f-8567-a32f4cd29ebc · outbound

This paper cites GPT-4 Technical Report.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models GPT-4 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.346201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.346201Z digest=sha256:de5fddea0d65b06a5938e66c2471bd83cbef9cf32d28d9ee4b69dbb406255282

Observation cb54f3f0-f509-4844-a63a-e9f8678a61ed · outbound

This paper cites L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.987745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.349544Z digest=sha256:7b70b4890d4aeb8b6957c754051e3b864dd39947053093c6abb2bf3a632f4a62

Observation 4891e545-c561-4147-b1d1-866eb9da518e · outbound

This paper cites L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.976260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.352961Z digest=sha256:e3cd066bb66972ad1b2a366db6c80e9a71809677622a45cff7fcb8148207dd5e

Observation 463148d6-1883-4822-b165-b519192fe8ae · outbound

This paper cites and Xie, S.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models and Xie, S

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.356423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.356423Z digest=sha256:a2521b862187f2313778b9816bc0985e4d44fac272e2a49f298c3c5df0e612bf

Observation f1b94478-f388-4ac6-80cc-a501d47f4708 · outbound

This paper cites SDXL: improving latent diffusion models for high-resolution image synthesis.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models SDXL: improving latent diffusion models for high-resolution image synthesis

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.958648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.359920Z digest=sha256:2c5822c5ea651a81fadd23250854c5fe613e3876eb16c33b9a00c788b5949544

Observation 2ee32771-19e3-4a1e-8b84-ee6453f5bf6c · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models High-resolution image synthesis with latent diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.947151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.363268Z digest=sha256:9adc1050f81d34351c741386a4dc298ebc8693a935844a23b2c5441cb457860b

Observation da93ef58-bc8b-4450-a102-4989ded7d2f2 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models U-net: Convolutional networks for biomedical image segmentation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.366845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.366845Z digest=sha256:a2220aa6dc88b306df428f6c12ad67a3be8af1bc16b3ec1c23e0835b1d2a3681

Observation 2a54d9d5-d202-4512-890d-61b9cf982b8d · outbound

This paper cites LAION-5B: an open large-scale dataset for training next generation image-text models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models LAION-5B: an open large-scale dataset for training next generation image-text models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.929516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.370367Z digest=sha256:ce52424087fa11bf413166ebefe0c52366cdb32309dc279e0821a8e04ce5e8a4

Observation f2c659b3-a8b7-4766-bd25-9870c2910a2d · outbound

This paper cites A., Maheswaranathan, N., and Ganguli, S.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models A., Maheswaranathan, N., and Ganguli, S

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.918690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.373782Z digest=sha256:3db46ff377fe4706f989c3b8d1569d211584f885b66cbcae5bab13a3b01ee3eb

Observation 74c99577-e516-4fe2-a861-162206aba917 · outbound

This paper cites Denoising diffusion implicit models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Denoising diffusion implicit models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.377376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.377376Z digest=sha256:ee16c0a4d9c6ad365e72d9cd0f4d50042bc2af1a9cc400d6ad4ee6c0ac0a2b9b

Observation 81da2080-a620-444a-93a4-19e85f845eb5 · outbound

This paper cites Consistency models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Consistency models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.899497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.380794Z digest=sha256:ff0639081f3e5ab90af78da0c1b906e401cb698e799a2a6239a354e644318a7a

Observation fd1b764d-4f11-42f3-bc79-9d06f29c65a9 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.384105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.384105Z digest=sha256:76f8c9ad873b85748f36dfe28e97a72780d7c916db83d4cfbd548073e890e06a

Observation 8752b132-1c6b-4e5d-ad17-277fb32a4aee · outbound

This paper cites Diffusion model alignment using direct preference optimization.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Diffusion model alignment using direct preference optimization

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.887636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.387754Z digest=sha256:d22df3e7958003fc2d84e0ca0292c6b36c5b477a8601690ec6f37b3825fb19e9

Observation 0b6d9d1b-668d-4bd5-9932-1cf822a576b5 · outbound

This paper cites Diffusion model alignment using direct preference optimization.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Diffusion model alignment using direct preference optimization

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.876517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.391188Z digest=sha256:ee4e948cf6352a3565b2cd69b70c854c0b81cc7e39f68c2ec511498a0d14aff8

Observation 3d57ade5-7816-4cc7-800d-c48ab8671811 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models ModelScope Text-to-Video Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.394937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.394937Z digest=sha256:9f89d85a340fa981eb03eb3ba81bf0e92a2f53e6aadaa8999747be03772882aa

Observation baa090bd-7c97-4d66-a500-6a3906b8df0d · outbound

This paper cites RichRAG: Crafting Rich Responses for Multi-faceted Queries in Retrieval-Augmented Generation.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models RichRAG: Crafting Rich Responses for Multi-faceted Queries in Retrieval-Augmented Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.398677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.398677Z digest=sha256:61c374ab322f147e392aaf761d43dc5b1b928185271cb75675679782b215f63f

Observation 7d231b12-a8c4-46ee-b5e2-0186e4827714 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.402512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.402512Z digest=sha256:53d7d02b4d178449df21af23b8cf0fb0361bf1ed92aea8c4c1e54e5032eab2a8

Observation 75a0a4aa-c41d-4012-82b4-d9c5bb0410e9 · outbound

This paper cites A., Khashabi, D., and Hajishirzi, H.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models A., Khashabi, D., and Hajishirzi, H

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.865449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.406014Z digest=sha256:832243a3b9ec10ffaa3b4552eac72a4a66ea15dc0237781199a785f4348db48b

Observation e9cf16ff-c0b8-4e7a-a393-d939dc595ad5 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.409465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.409465Z digest=sha256:ed2c2fd67692ec5210ece23c6c525e7c2279080b946ce52d7795483b0592cfab

Observation 5551404b-fa29-4bb9-99ce-2a951dd855d8 · outbound

This paper cites Easyanimate: A high-performance long video generation method based on transformer architecture.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Easyanimate: A high-performance long video generation method based on transformer architecture

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.413483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.413483Z digest=sha256:63ee4e245e946f42fa30d686cc4a538f8aa36df8cd3ab10b375868a9045e0ea0

Observation ac826ea7-c472-40d0-a0f3-b374c6d65701 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.416856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.416856Z digest=sha256:1d4f7644dc21eb6f168036c6c22bf68539b68288c8683b582bbfd57a42b406f6

Observation 2e0eadc8-e38b-4b08-9360-074146c24f07 · outbound

This paper cites S., Eom, S., Han, G., Nam, D.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models S., Eom, S., Han, G., Nam, D

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.854512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.420575Z digest=sha256:8193294650aacbe353f0b1b0c34db5863875817b8dde697b7e60d4e27dea7246

Observation 333f10b4-35ed-40bb-bb5a-faa9522fd818 · outbound

This paper cites Instructvideo: Instructing video diffusion models with human feedback.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Instructvideo: Instructing video diffusion models with human feedback

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.842663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.424192Z digest=sha256:7e3fb823aa1a537aa76d6bd364e0fce6fcf6f435c3a5cdf26ac285824ef35bfc

Observation 0dc15b45-3499-4c62-8900-cadfb86ff712 · outbound

This paper cites Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.427689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.427689Z digest=sha256:fb191ad74f21fcfc1438c842f0a736842bcbdebaf5838fa9f48c91ecb1688673

Pith citing papers

Observation 00c016e6-4207-42fb-87cc-2dd859cd3def · inbound

Towards Anatomically Plausible Human Image Generation via Synthetic Localized Preferences cites this paper.

Towards Anatomically Plausible Human Image Generation via Synthetic Localized Preferences Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:14:04.744750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:55:02.514189Z digest=sha256:db882843da64cfb608b6f666c93f37684182e2666c30c297abb12320e5e8dac7