Pith. sign in

Paper Citation Record · LEDGER

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models

As of 10 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2502.06812.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06812 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:26:17.427689Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:55:02.514189Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T00:14:04.743463Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8c5a285f-8ac6-42b5-8765-f6c555b337b3 · outbound

This paper cites write newline.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.243081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.243081Z digest=sha256:81f469358f1f2d1caf3352c79a0266c67cf4a3ae25f39121c48fba4c02693c91

Observation 3fbd6934-8faf-4eb0-b755-3fdc48153444 · outbound

This paper cites URL https://pictory.ai/blog.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models URL https://pictory.ai/blog

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.157160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.248155Z digest=sha256:be8d8700000f0321af9f6fa408c3eae93bb471460b43d6ad0b5d5e122d4fdbda

Observation 8b75cea8-a0a0-4062-8e95-6c007e4a12b1 · outbound

This paper cites URL https://runwayml.com/.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models URL https://runwayml.com/

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.146504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.253340Z digest=sha256:8dafdd6e4a45d92e9bc3a283d964ac0d4376df536dea29e1ddf17aa72a3aaf5a

Observation aedda03a-87d7-4a68-ab79-768bf9587477 · outbound

This paper cites URL https://www.synthesia.io/tools/ai-video-editor.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models URL https://www.synthesia.io/tools/ai-video-editor

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.135495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.257049Z digest=sha256:b88fcb8ba4f79cb0f96d55e4ee8fdd4019d0c1b8b7e215cf9d1790a220da9e32

Observation efeb00fe-29df-4a0d-8be1-5a03cdde91c2 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.125190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.260790Z digest=sha256:9d1c5fa565298e90d01474f9acfd60f663ca1c5d0b570bff6db74b33c18af32e

Observation 5df60e09-8171-4268-82fe-fb82643f38ca · outbound

This paper cites Training diffusion models with reinforcement learning.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Training diffusion models with reinforcement learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.114384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.264626Z digest=sha256:3ecee69833c90bff886c4531c556f202da42de31e0ded51aeb6dafd4afcbf774

Observation de1572f6-76a2-441a-be2d-add4be29ed80 · outbound

This paper cites W., Fidler, S., and Kreis, K.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models W., Fidler, S., and Kreis, K

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.102381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.268305Z digest=sha256:2919a172ca86533870bdd0ab9932fb080a9ee508abdd70e8e443b0d8a23badcf

Observation dc14affc-bd09-4e31-b33f-bc07333ecdc7 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.271948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.271948Z digest=sha256:2e5641229fc1f85a125154108fb6843cc06d9eafa7f62cf1a133008bd2af23c9

Observation e2fed0ae-baa9-4c8f-b9b3-22ac36d071b5 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.091279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.275902Z digest=sha256:9026ddb72cee82adc19cdd50da2f532230ae6cf7c6c22896b10c1b8d426576c5

Observation fde46d5c-986b-4c56-ab52-c3ff07c40c11 · outbound

This paper cites S., Brox, T., and Ronneberger, O.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models S., Brox, T., and Ronneberger, O

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.080150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.279374Z digest=sha256:19ade186595b249de5da1eb3b0e16466e791d6bbd0d19c2084f837deb7a85fb4

Observation 1bac0149-6a1c-49f8-8c9d-6579e6e16b07 · outbound

This paper cites an unresolved cited work.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:26:18.069161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.283020Z digest=sha256:54acce4f5fe22bbdc2611fe024ec15eb20fccd797a750bca3bea4ae177e59810

Observation ced1d457-8368-476b-af66-8257e700416c · outbound

This paper cites DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.286546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.286546Z digest=sha256:c8a365e1c38e3ca04ecc78b1909318eae57c1802e56a5a4f31c6c595fd78e08f

Observation 7245801a-4267-4d2f-8344-a3abe0d74e8a · outbound

This paper cites D., Ni, Y., Lyu, B., Narsupalli, Y., Fan, R., Lyu, Z., Lin, B.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models D., Ni, Y., Lyu, B., Narsupalli, Y., Fan, R., Lyu, Z., Lin, B

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.058047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.290440Z digest=sha256:22c02784a8e1b25c516e6f1e3448465b71ee90aaa37761214dde20178b13714d

Observation 12601fad-88be-4d2b-b42d-9f27ac66d592 · outbound

This paper cites Denoising diffusion probabilistic models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Denoising diffusion probabilistic models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.293856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.293856Z digest=sha256:0fc55986ed4097fd21479e7d88e22259f6ac22512fc30c35e974f2d037df2450

Observation 5ef611e7-864f-491f-8fe7-6842261106a1 · outbound

This paper cites Video Diffusion Models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Video Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.297240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.297240Z digest=sha256:d3e23849fdb2f8fc4fe9e6dbacd9211758bb403532834438afe7cf4097b6f2f2

Observation 01ef9893-5173-45df-a4f6-c633fc7e7f41 · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to-video generation via transformers.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Cogvideo: Large-scale pretraining for text-to-video generation via transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.300911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.300911Z digest=sha256:110ed8aec7eb892e566ab332c1fbdef663b1b7815d0e1fbc04f9e88a959e8345

Observation 41c20b07-a1e8-4943-aedf-ea363127e0ab · outbound

This paper cites J., Shen, Y., Wallis, P., Allen - Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models J., Shen, Y., Wallis, P., Allen - Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.304046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.304046Z digest=sha256:e7b98f766c0ea5271f1ed4408b4294f6557867ed36532b59a9e9f8a9ce14d282

Observation 1c883af6-977d-49db-b3e8-a0d5e6802840 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Vbench: Comprehensive benchmark suite for video generative models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.027581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.307223Z digest=sha256:0eb726830bbed73cda804b74ccb174f6987254df74686f2fe751033e8322cad7

Observation 13f1cb82-f460-4ad9-a828-d0af78269364 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.310948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.310948Z digest=sha256:9fe1997181b611ee4a879971e7ce8a7c6f40f4db82846070fd5ae99d455e21fc

Observation e90345d2-d1b0-462c-9b14-49de74ab4ccc · outbound

This paper cites Text2video-zero: Text-to-image diffusion models are zero-shot video generators.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Text2video-zero: Text-to-image diffusion models are zero-shot video generators

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:18.015879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.314581Z digest=sha256:e6de02b6719602a761c8b6eb3768b6152e698ff639fedd31cb6ff4ab8cec5461

Observation bf72edf6-0785-4f72-8c9e-10a2e351ee92 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.317958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.317958Z digest=sha256:980a63dd36ac31fbad3e9c3f346494cac4cdba011e9c2aae4e62ae1f32a4aa65

Observation 290e26b9-8ff6-43c0-ada4-d2130980bf9c · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.321272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.321272Z digest=sha256:8b0381683ea2cc707f3befaee77fe5c4b717df72c05db1c7addb65eaeb348ab2

Observation baf6ea88-7f3d-473c-98db-8daeaf0af1a1 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Aligning Text-to-Image Models using Human Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.324956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.324956Z digest=sha256:412cfdceda356f515bf66d1b5f0ee86f5cf4e65e98dbad49e692fcee70283aee

Observation 9bc82d02-afd8-4e98-b2cb-a35855234460 · outbound

This paper cites T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.328708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.328708Z digest=sha256:6ad2f2e7d67d4d1e9aa98e5cb34260f5d39afb15018a1587471be0719a48039a

Observation 9579da8f-f846-4c91-84c6-1b1559d92b32 · outbound

This paper cites an unresolved cited work.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.332358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.332358Z digest=sha256:aba7c4a1ca3474e35ed3054636abbd65218976289ce136420396fb9c238bf8c8

Observation 80cabdce-42f1-4842-bbbe-f9b87b4f4220 · outbound

This paper cites Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.335567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.335567Z digest=sha256:2fad2998e6eb97be790e0bc1dcd1b420b312369bd6c26d87b008ce00620fc4f5

Observation e20643f3-bd49-4963-8475-adfa24984a24 · outbound

This paper cites Training diffusion models towards diverse image generation with reinforcement learning.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Training diffusion models towards diverse image generation with reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.998359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.339251Z digest=sha256:e7222d9010f5e474c933609cc3d775d121e830f0accd70799e2d588a6ee2164a

Observation 092329e7-6779-4fc7-b795-96716892445d · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.342550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.342550Z digest=sha256:7a0c3db780d4e233283ac689dc30efd8e52c3485483c7ab6a57d4d7d5cc8912d

Observation c94f1892-48bb-460f-8567-a32f4cd29ebc · outbound

This paper cites GPT-4 Technical Report.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models GPT-4 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.346201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.346201Z digest=sha256:17426d2c4a10c7923b613d752cc9b7dd486e32148a9b9b703ec767ab4e8d5f3d

Observation cb54f3f0-f509-4844-a63a-e9f8678a61ed · outbound

This paper cites L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.987745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.349544Z digest=sha256:3ef1bf957c59df763f0cb4366eed2519df23485fc941315f183279ba4cb15483

Observation 4891e545-c561-4147-b1d1-866eb9da518e · outbound

This paper cites L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.976260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.352961Z digest=sha256:7cd26d263741f48150d1b6ed9e4c2043a13243929702dea1cae2364a9cc211ba

Observation 463148d6-1883-4822-b165-b519192fe8ae · outbound

This paper cites and Xie, S.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models and Xie, S

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.356423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.356423Z digest=sha256:e45205a3ac3a1835e2162aaf351a1aab957be0c921147211c0c663ab1c7fd782

Observation f1b94478-f388-4ac6-80cc-a501d47f4708 · outbound

This paper cites SDXL: improving latent diffusion models for high-resolution image synthesis.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models SDXL: improving latent diffusion models for high-resolution image synthesis

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.958648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.359920Z digest=sha256:e6186ec19bd9ce83a65236f31706bcd211e02f0851713becaf1fb2e445b6f782

Observation 2ee32771-19e3-4a1e-8b84-ee6453f5bf6c · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models High-resolution image synthesis with latent diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.947151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.363268Z digest=sha256:c277a38fb151054efc5be0c809151a888bfe8f038d1952268e6d8d9bbdf799c8

Observation da93ef58-bc8b-4450-a102-4989ded7d2f2 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models U-net: Convolutional networks for biomedical image segmentation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.366845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.366845Z digest=sha256:7d593f1bf4e3aa5576885c2a79e6b327d680208d169c126ef21b08184795d741

Observation 2a54d9d5-d202-4512-890d-61b9cf982b8d · outbound

This paper cites LAION-5B: an open large-scale dataset for training next generation image-text models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models LAION-5B: an open large-scale dataset for training next generation image-text models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.929516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.370367Z digest=sha256:327d488fad9545f854e69f91e170595658554217d3343fdeaa2a1c16ee2854db

Observation f2c659b3-a8b7-4766-bd25-9870c2910a2d · outbound

This paper cites A., Maheswaranathan, N., and Ganguli, S.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models A., Maheswaranathan, N., and Ganguli, S

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.918690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.373782Z digest=sha256:ba42f3a67d5a3013a44b30229a1b9798a5aabae7ed656b5056e8c7772c3f4c3f

Observation 74c99577-e516-4fe2-a861-162206aba917 · outbound

This paper cites Denoising diffusion implicit models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Denoising diffusion implicit models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.377376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.377376Z digest=sha256:09167261270f0a0aa4ea3d8ed2e3bb216324807430731dd2e72c599f7cf72e65

Observation 81da2080-a620-444a-93a4-19e85f845eb5 · outbound

This paper cites Consistency models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Consistency models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.899497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.380794Z digest=sha256:ae98c644748c3a39102a6d194e0f7fe138b726344e2aa53a930d216d2a0df8d9

Observation fd1b764d-4f11-42f3-bc79-9d06f29c65a9 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.384105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.384105Z digest=sha256:a14271c5fce34ee9a03103aa8c92d5d9d5700115cf6b9cd342c20e352e3a4bc0

Observation 8752b132-1c6b-4e5d-ad17-277fb32a4aee · outbound

This paper cites Diffusion model alignment using direct preference optimization.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Diffusion model alignment using direct preference optimization

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.887636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.387754Z digest=sha256:bfb1a4b1f0a886fe2bbd649928166e1218a1bf3ec4ceaebc9746bb6f9faf463e

Observation 0b6d9d1b-668d-4bd5-9932-1cf822a576b5 · outbound

This paper cites Diffusion model alignment using direct preference optimization.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Diffusion model alignment using direct preference optimization

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.876517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.391188Z digest=sha256:036e3e40f7339d02157cadea89f91fb2557da06231457714345ebc86af50b1a3

Observation 3d57ade5-7816-4cc7-800d-c48ab8671811 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models ModelScope Text-to-Video Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.394937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.394937Z digest=sha256:57840f05b3f36e3036c2fff52994f38ba9949f054d2f8bec1917c5b1a0b418c2

Observation baa090bd-7c97-4d66-a500-6a3906b8df0d · outbound

This paper cites RichRAG: Crafting Rich Responses for Multi-faceted Queries in Retrieval-Augmented Generation.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models RichRAG: Crafting Rich Responses for Multi-faceted Queries in Retrieval-Augmented Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.398677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.398677Z digest=sha256:c1b39f0656e2c62808da4cb4e065ffb7a52886a7165ab1f10546690ada9caa47

Observation 7d231b12-a8c4-46ee-b5e2-0186e4827714 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.402512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.402512Z digest=sha256:ed1c7a9493c8a4b3564a6b4bca531c56f219363efe69f16da614da649332f0b7

Observation 75a0a4aa-c41d-4012-82b4-d9c5bb0410e9 · outbound

This paper cites A., Khashabi, D., and Hajishirzi, H.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models A., Khashabi, D., and Hajishirzi, H

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.865449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.406014Z digest=sha256:1f4bdcd9191180293529ca292088b3c8b113ba48de1f93322735f68f34cc4431

Observation e9cf16ff-c0b8-4e7a-a393-d939dc595ad5 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.409465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.409465Z digest=sha256:26a142a1b145e6c77a409c722f61419c26046f829b45001447e4c5e8acf3e9bf

Observation 5551404b-fa29-4bb9-99ce-2a951dd855d8 · outbound

This paper cites Easyanimate: A high-performance long video generation method based on transformer architecture.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Easyanimate: A high-performance long video generation method based on transformer architecture

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.413483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.413483Z digest=sha256:2c9769aabdb4bd0e608898f0d80e8bd510b45d509d632810f0290e3e30168207

Observation ac826ea7-c472-40d0-a0f3-b374c6d65701 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.416856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.416856Z digest=sha256:69e31f8edcf417be3df555f7fd85b36ef60da720e5a1ad7d92bc74c49ef96275

Observation 2e0eadc8-e38b-4b08-9360-074146c24f07 · outbound

This paper cites S., Eom, S., Han, G., Nam, D.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models S., Eom, S., Han, G., Nam, D

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.854512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.420575Z digest=sha256:dff6fca487e7f0862ee6db29245870f5c60fa2833d78a45282647a45244702db

Observation 333f10b4-35ed-40bb-bb5a-faa9522fd818 · outbound

This paper cites Instructvideo: Instructing video diffusion models with human feedback.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Instructvideo: Instructing video diffusion models with human feedback

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:26:17.842663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T11:26:17.424192Z digest=sha256:dc7313fa47ca21230c06cbf1c4d55c928b98e1489b391bb0af7d8b501c69aee3

Observation 0dc15b45-3499-4c62-8900-cadfb86ff712 · outbound

This paper cites Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.427689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.427689Z digest=sha256:e482260089014b8d3cbf3ce60a22cecc175a4a414435b2ca98cf7e202220e7fc

Pith citing papers

Observation 00c016e6-4207-42fb-87cc-2dd859cd3def · inbound

Towards Anatomically Plausible Human Image Generation via Synthetic Localized Preferences cites this paper.

Towards Anatomically Plausible Human Image Generation via Synthetic Localized Preferences Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:14:04.744750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:55:02.514189Z digest=sha256:5da58f5bbdca4aab87de57a3ebeec540511eceff170de59d9e9f109fcb087898