Pith. sign in

Paper Citation Record · LEDGER

DIVE: Taming DINO for Subject-Driven Video Editing

As of 22 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 2 inbound Pith citation observations for arXiv:2412.03347.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03347 v2

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:33:37.215743Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:49:15.079375Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T14:32:36.203118Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact1
  • verified fuzzy36
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1195cb44-1cec-402d-bd32-1dea98310b8d · outbound

This paper cites Deep ViT Features as Dense Visual Descriptors.

DIVE: Taming DINO for Subject-Driven Video Editing Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.253762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.253762Z digest=sha256:cca620009c89feb82c5ebb5783e9dcd3c13f40758f6510c407997ee8eb4fc099

Observation 83d79157-5d24-4beb-bf97-b1d33882e33d · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

DIVE: Taming DINO for Subject-Driven Video Editing Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.845732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.258885Z digest=sha256:e5ad1b77828bdf144bbc234ad97cde7482b732add4c97cbdceb258520b9c9dc3

Observation 5b175b5b-a0ab-49be-bb3e-468fa28613f5 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

DIVE: Taming DINO for Subject-Driven Video Editing In- structpix2pix: Learning to follow image editing instructions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.830689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.263226Z digest=sha256:24de7f2fa8d8e0796e4fec8c44375c5ac13c94d33d03c35aef0677e0bfc4fe04

Observation 8ba869ae-f3cf-4226-830e-b2f43a01d42b · outbound

This paper cites DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation.

DIVE: Taming DINO for Subject-Driven Video Editing DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.267881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.267881Z digest=sha256:66bb488498e08110442d3ac6a00860a3abee96432545ea41d16b259410818ce8

Observation 570ee406-cda8-456d-8ffb-2d03a2bc319f · outbound

This paper cites Subject-driven Text-to-Image Generation via Apprenticeship Learning.

DIVE: Taming DINO for Subject-Driven Video Editing Subject-driven Text-to-Image Generation via Apprenticeship Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.272670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.272670Z digest=sha256:8ee45aea90da3c88377a70901dbcac695cdb92320762ef9b8f546439a4b883f4

Observation 6dfa0a8f-f6e8-4b7f-933c-b55439e3d058 · outbound

This paper cites Slicedit: Zero- shot video editing with text-to-image diffusion models using spatio-temporal slices.

DIVE: Taming DINO for Subject-Driven Video Editing Slicedit: Zero- shot video editing with text-to-image diffusion models using spatio-temporal slices

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.816516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.277065Z digest=sha256:2b280bb93cba738a1a73e7eff2f4e56d9a8aeb0375fc8c5a5f1199a7b2064160

Observation ab698145-20a9-4768-8da1-3258c4ecfad3 · outbound

This paper cites FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing.

DIVE: Taming DINO for Subject-Driven Video Editing FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.280747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.280747Z digest=sha256:5ff5aae23a470cc40ffb3feb8fe0d264695d8ba5b13ff96a0a49305c686fcaaa

Observation dfbb1d24-336a-4915-a815-56cfca20ab7d · outbound

This paper cites Diffusion models beat gans on image synthesis.

DIVE: Taming DINO for Subject-Driven Video Editing Diffusion models beat gans on image synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.285452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.285452Z digest=sha256:b1a296ca0f3e36ae9c202b54f22ce5e5f520ec9d6188d7eb7946c69dc3e9dc75

Observation bf0232df-116d-444b-8263-c7f2738daf03 · outbound

This paper cites Structure and content-guided video synthesis with diffusion models.

DIVE: Taming DINO for Subject-Driven Video Editing Structure and content-guided video synthesis with diffusion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.792503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.289426Z digest=sha256:f1e6d5904fde9309424ebbb3d6091413c62613d5ca8eb0f83eef63734de67f8f

Observation dc55732d-6484-48b7-9dfb-294086c86f77 · outbound

This paper cites Ccedit: Creative and controllable video editing via diffusion models.

DIVE: Taming DINO for Subject-Driven Video Editing Ccedit: Creative and controllable video editing via diffusion models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.777177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.293245Z digest=sha256:9acd9a02de51987063fa17a9b26bdef031431d28b6f78dbd6979ebfd74bcf7a3

Observation 73fa0a72-d29b-49e7-8aa9-0220f62686e7 · outbound

This paper cites An image is worth one word: Personalizing text-to-image generation using textual inversion.

DIVE: Taming DINO for Subject-Driven Video Editing An image is worth one word: Personalizing text-to-image generation using textual inversion

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.762890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.298143Z digest=sha256:bf2764cfebc16d80086bd079707d83d3ec2c095ed303c4e8deaf3aed076447ae

Observation 0dfd8851-1615-4b15-a521-c4f8bfc2cd8c · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

DIVE: Taming DINO for Subject-Driven Video Editing Preserve your own correlation: A noise prior for video diffusion models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.747155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.301732Z digest=sha256:a0f4dab65ac7a5f66b4927a30c172a14a263cc804876693e5f69422273fe7910

Observation c77b4c38-bb13-4ea4-b9f1-93683dd42248 · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

DIVE: Taming DINO for Subject-Driven Video Editing TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.305651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.305651Z digest=sha256:e16f5c185f4b13e7104d9fc3f215be771330eaf055893e6797f29014b8c54642

Observation 904e81fe-f636-406e-89ef-1d69ed5ee284 · outbound

This paper cites Photoswap: Personalized Subject Swapping in Images.

DIVE: Taming DINO for Subject-Driven Video Editing Photoswap: Personalized Subject Swapping in Images

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-11T22:33:37.606539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.310176Z digest=sha256:787af535272072644f18dee1327644c120ebcb6988124d4b327c3e87d1a39a60

Observation 7800bab4-26ed-4c6d-b1e1-30c5e63b37cc · outbound

This paper cites Mix-of-show: Decentralized low- rank adaptation for multi-concept customization of diffusion models.

DIVE: Taming DINO for Subject-Driven Video Editing Mix-of-show: Decentralized low- rank adaptation for multi-concept customization of diffusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.682548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.314406Z digest=sha256:92a97e288cd9965357d50b2b01ca15c6da277cf9c6e5fe664f936c67214983a1

Observation 74cae92f-f09f-4d48-8e35-2b450e4ff96f · outbound

This paper cites Videoswap: Customized video subject swapping with interactive semantic point cor- respondence.

DIVE: Taming DINO for Subject-Driven Video Editing Videoswap: Customized video subject swapping with interactive semantic point cor- respondence

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.591014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.318843Z digest=sha256:756188b4e282595b552916f732bc7201e059094a613a4857443ccd4112ff55ac

Observation 8ecb986c-ae37-4f66-8daf-a82a32185f08 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

DIVE: Taming DINO for Subject-Driven Video Editing AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.323746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.323746Z digest=sha256:48b25ed749d9d23127eec3f2a9dcbc0b1105f44a1ab7475588910b93253db8b8

Observation 9b1711d7-348c-4964-a78f-158a59d07465 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

DIVE: Taming DINO for Subject-Driven Video Editing Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.342467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.342467Z digest=sha256:7bae043a5a089240deeac714f684169b6bc1b8a2a5e883dc26879016956d3250

Observation bfd66e7c-0eac-4e97-bdb6-a3047e9df413 · outbound

This paper cites Denoising dif- fusion probabilistic models.

DIVE: Taming DINO for Subject-Driven Video Editing Denoising dif- fusion probabilistic models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.373545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.373545Z digest=sha256:e3dcecbcab04aefbcdff21192c44fb4b85d2818d7826908bb249819fee88204b

Observation 0f87215b-ab25-48b6-a834-5897043dcc5a · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

DIVE: Taming DINO for Subject-Driven Video Editing Imagen Video: High Definition Video Generation with Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.438404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.438404Z digest=sha256:56dd022d0cd012d5122838b0a6156c135908b5cd028291b2fd8428f1760b1473

Observation b9c0573d-6b6e-4066-acea-9c19eddee8fb · outbound

This paper cites Cascaded diffu- sion models for high fidelity image generation.

DIVE: Taming DINO for Subject-Driven Video Editing Cascaded diffu- sion models for high fidelity image generation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.500496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.538495Z digest=sha256:a248d23f6708415b3e42358ba79db9ac5ca7b4988336e62bdd8d36714e81f2c4

Observation 3da87eee-3a64-4228-9437-03e975f3fd7b · outbound

This paper cites Video dif- fusion models.

DIVE: Taming DINO for Subject-Driven Video Editing Video dif- fusion models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.483058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.543494Z digest=sha256:a719d55695d28f8d6b33eb3405ed0530f6789bb39279a5d17f88e0ab2e7ad7ca

Observation cf9cfdab-ebd2-4dbf-8ed6-50daa4637d1b · outbound

This paper cites VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet.

DIVE: Taming DINO for Subject-Driven Video Editing VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.547376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.547376Z digest=sha256:f6114a15b1e39a1d1ce495ffd6eba7d4ace982cb64c5f13e181f0ca8f9e363d9

Observation 046c8a49-5f92-4e6c-97d6-9232396d9780 · outbound

This paper cites SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models.

DIVE: Taming DINO for Subject-Driven Video Editing SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.551431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.551431Z digest=sha256:172ee49a87a1331fe9ec9c435b2fff59f24b4bbb7057c9b85d5746063d049915

Observation d63188c8-2c47-44b7-a7e8-68e2060dc459 · outbound

This paper cites Diffusion Model-Based Image Editing: A Survey.

DIVE: Taming DINO for Subject-Driven Video Editing Diffusion Model-Based Image Editing: A Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.556016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.556016Z digest=sha256:e525f0203b9fec72e127f0eb8a61f1a4ac56aaaa0f6d406b97a4bc5d911b07b3

Observation 2045473e-f9b8-4257-af79-a24f0d24c57b · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

DIVE: Taming DINO for Subject-Driven Video Editing Vbench: Comprehensive bench- mark suite for video generative models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.468511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.560745Z digest=sha256:f7ba703c8607da4d10d78c9679e6ed7a9982022982faffd188ff175ccda82574

Observation 3e0d087e-86df-4345-869b-3fbf19accd4e · outbound

This paper cites Rave: Randomized noise shuf- fling for fast and consistent video editing with diffusion mod- els.

DIVE: Taming DINO for Subject-Driven Video Editing Rave: Randomized noise shuf- fling for fast and consistent video editing with diffusion mod- els

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.453562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.565582Z digest=sha256:3d96d1dc0464b0debab1b74ff17de4330e422e66ffefa08e35e6e6fcb56431f6

Observation 18ffaf68-87fb-48e0-aa79-6df3438496cc · outbound

This paper cites Text2video-zero: Text- to-image diffusion models are zero-shot video generators.

DIVE: Taming DINO for Subject-Driven Video Editing Text2video-zero: Text- to-image diffusion models are zero-shot video generators

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.438681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.569368Z digest=sha256:884592f724810d8952e31f5fbc7ba5d7a924f0128cb4301343af6c146c557c64

Observation a597bea9-c146-43fd-87a0-f4092d822ce9 · outbound

This paper cites OMG: Occlusion-friendly Personalized Multi-concept Generation in Diffusion Models.

DIVE: Taming DINO for Subject-Driven Video Editing OMG: Occlusion-friendly Personalized Multi-concept Generation in Diffusion Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.573699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.573699Z digest=sha256:1e1f7bf42eea39fd97b4ba2d6a16ac9fedb19f6d2ac0696189ac14fc8a10169b

Observation 2af268f0-17eb-43bb-bbb9-0047269231d5 · outbound

This paper cites Anyv2v: A tuning-free framework for any video-to- video editing tasks.

DIVE: Taming DINO for Subject-Driven Video Editing Anyv2v: A tuning-free framework for any video-to- video editing tasks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.425114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.578185Z digest=sha256:571054a40cec767a544e60b148e6f4edd6b0306e77748911e4e06788543cc4ce

Observation dacb8d6f-1208-4446-83d5-87d5eebe6787 · outbound

This paper cites Multi-concept customiza- tion of text-to-image diffusion.

DIVE: Taming DINO for Subject-Driven Video Editing Multi-concept customiza- tion of text-to-image diffusion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.582476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.582476Z digest=sha256:f03471dcefea86aee944f8eb9deb80a6014709f83a7d51f1054aa73979242a44

Observation e98ebf46-fe6e-446d-a132-760e29408db7 · outbound

This paper cites MagicEraser: Erasing Any Objects via Semantics-Aware Control.

DIVE: Taming DINO for Subject-Driven Video Editing MagicEraser: Erasing Any Objects via Semantics-Aware Control

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.586309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.586309Z digest=sha256:0451a366a0ab3b5b6644d7c3c79e43e84a8c4ea8722ad6879ab37b18fadfa981

Observation 7692b6be-0d7c-4c13-a30b-0b183463d686 · outbound

This paper cites A video is worth 256 bases: Spatial-temporal expectation-maximization inversion for zero-shot video editing.

DIVE: Taming DINO for Subject-Driven Video Editing A video is worth 256 bases: Spatial-temporal expectation-maximization inversion for zero-shot video editing

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.402001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.590293Z digest=sha256:8f43499faad38a3528a41d3a3994c78ea982dae0adb1748c6460643659d07497

Observation ee223534-aa57-40f9-a659-7bd45ab6937c · outbound

This paper cites Vidtome: Video token merging for zero-shot video editing.

DIVE: Taming DINO for Subject-Driven Video Editing Vidtome: Video token merging for zero-shot video editing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.317705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.594270Z digest=sha256:cb8d2af23217887cedb9a5ed9eb91cd9e1692386568c4e7017135fd022bdc45b

Observation 4fb6d6a3-d86c-4121-b5fe-a8eb177c61b9 · outbound

This paper cites Flowvid: Taming imperfect opti- cal flows for consistent video-to-video synthesis.

DIVE: Taming DINO for Subject-Driven Video Editing Flowvid: Taming imperfect opti- cal flows for consistent video-to-video synthesis

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.285380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.598320Z digest=sha256:3acb27dc5e1d7fa85730a22444ec05738e35bed07a921dff0ceead1602bb89f3

Observation 4a64455a-7147-49fe-a770-8434e970e0a2 · outbound

This paper cites MagicEdit: High-Fidelity and Temporally Coherent Video Editing.

DIVE: Taming DINO for Subject-Driven Video Editing MagicEdit: High-Fidelity and Temporally Coherent Video Editing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.602077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.602077Z digest=sha256:7305b2456ff292fb038d51ddf4fbec8a30ee6500af439086ee4a2c440420c54f

Observation 94e780c1-abac-4247-915a-83105a5170e6 · outbound

This paper cites Video-p2p: Video editing with cross-attention control.

DIVE: Taming DINO for Subject-Driven Video Editing Video-p2p: Video editing with cross-attention control

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.605952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.605952Z digest=sha256:7a2dc775ce58f51cfee37511e4b4d9a8f2d3b0d5a29d5b2a6aaa44e5417e48ac

Observation 178be755-c851-4db2-8eaf-57708e4cf501 · outbound

This paper cites Cones: Concept neurons in diffusion models for customized generation.

DIVE: Taming DINO for Subject-Driven Video Editing Cones: Concept neurons in diffusion models for customized generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.259667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.609656Z digest=sha256:2ff1e557cbec6aeae0e6b43e2a84942c8bbd88020d46286bd2bf374796b9a536

Observation 15c260c8-e47f-40a6-9601-8aaeffbdb276 · outbound

This paper cites Spe- cialist diffusion: Plug-and-play sample-efficient fine-tuning of text-to-image diffusion models to learn any unseen style.

DIVE: Taming DINO for Subject-Driven Video Editing Spe- cialist diffusion: Plug-and-play sample-efficient fine-tuning of text-to-image diffusion models to learn any unseen style

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.227844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.613779Z digest=sha256:f13ce8e6ffe57503941e1da50541084ac6f8073400a6cada86e7100a5bde36e8

Observation 1a598fc6-ca88-4eaf-95d6-383260de3bce · outbound

This paper cites Gpt4motion: Scripting physical motions in text-to-video generation via blender-oriented gpt planning.

DIVE: Taming DINO for Subject-Driven Video Editing Gpt4motion: Scripting physical motions in text-to-video generation via blender-oriented gpt planning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.189761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.617825Z digest=sha256:ef884474ba6e6a8596bb76b74b73923e4cc7e3a1cbca26f94a732f9f681b953e

Observation ead4acb3-1b76-48d7-8cc4-7be013e63f3f · outbound

This paper cites Deep spectral methods: A surprisingly strong baseline for unsupervised semantic segmentation and localization.

DIVE: Taming DINO for Subject-Driven Video Editing Deep spectral methods: A surprisingly strong baseline for unsupervised semantic segmentation and localization

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.175529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.627507Z digest=sha256:ece8fe3c3742dc6b710f02a2ab1d34771f64e2a15e671ad083bdba53bf3611f9

Observation 18c2f70a-58a3-4014-8f0b-225a7da4c425 · outbound

This paper cites T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models.

DIVE: Taming DINO for Subject-Driven Video Editing T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.663038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.663038Z digest=sha256:605ca88d6dbd2480a37e733e661c4b46926837b01d358e1db2f9fede31ccfeae

Observation cbd05bc3-9677-4d74-897e-59d0d76ce216 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

DIVE: Taming DINO for Subject-Driven Video Editing DINOv2: Learning Robust Visual Features without Supervision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.721166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.721166Z digest=sha256:e4e463b090d402dba7bb418a14d56f7884ab57b6e0c4cfa23545e7428e178a74

Observation ac0593e1-a0b8-4568-9039-d7aefbd7567a · outbound

This paper cites I2vedit: First-frame-guided video editing via image-to- video diffusion models.

DIVE: Taming DINO for Subject-Driven Video Editing I2vedit: First-frame-guided video editing via image-to- video diffusion models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.161653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.785841Z digest=sha256:077f80a4df2cd682765c9ac0908c49e7a07dbeaa49e1f5e6ee844fab1e849dd2

Observation 72263605-9c78-477d-a313-3b5de24eac14 · outbound

This paper cites A benchmark dataset and evaluation methodology for video object segmentation.

DIVE: Taming DINO for Subject-Driven Video Editing A benchmark dataset and evaluation methodology for video object segmentation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.147419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.848875Z digest=sha256:400946270175742ee24d3554a39471d26379173ac74a5d1f829090b7db52862d

Observation c77e20cd-deb1-4e98-b0a6-be13e6438114 · outbound

This paper cites Fatezero: Fus- ing attentions for zero-shot text-based video editing.

DIVE: Taming DINO for Subject-Driven Video Editing Fatezero: Fus- ing attentions for zero-shot text-based video editing

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.853310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.853310Z digest=sha256:9b8f1d15997f0397948164c657a1c5886f50cfbb05ae32f64be468a95a1a709d

Observation 50bb8db0-b9ca-487c-a8ef-9f17cc043c9b · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

DIVE: Taming DINO for Subject-Driven Video Editing High-resolution image syn- thesis with latent diffusion models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.857935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.857935Z digest=sha256:c2e942a716969800f3a8714cd0a3ed0709fbcc0246287f42c501ec358c2c2268

Observation babc799a-a365-4b0e-b288-88792eb2c953 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

DIVE: Taming DINO for Subject-Driven Video Editing Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.116222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.861789Z digest=sha256:b244a3ee581e1dc958788c147e0a76bf578003bc05ba4da59c7c7738df9ca550

Observation 4a7bdab8-c8a4-4037-9d11-bae0d83a32aa · outbound

This paper cites Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models.

DIVE: Taming DINO for Subject-Driven Video Editing Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.102698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.865949Z digest=sha256:f2d7386a0bae30b0867247a0d4993b8a568a17b5b07126fa813948a56461b5f7

Observation 6fa9e447-4c55-4f91-8355-df5655bbae4d · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

DIVE: Taming DINO for Subject-Driven Video Editing Photorealistic text-to-image diffusion models with deep language understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.869699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.869699Z digest=sha256:fe167fcb9f902dd3fd57713d5e61f169f81d1df3557d101a6a362772cfdc6547

Observation ebbe8ea9-60bd-47c7-84eb-e34301811c10 · outbound

This paper cites Emu Edit: Precise Image Editing via Recognition and Generation Tasks.

DIVE: Taming DINO for Subject-Driven Video Editing Emu Edit: Precise Image Editing via Recognition and Generation Tasks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.873906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.873906Z digest=sha256:24cd4496ff69e3cc479d95172b0222f1c8d156620a5bf68f6ecc18fd60686e43

Observation a56a2e16-0d6e-4eb2-b5a8-1fe63d7a2378 · outbound

This paper cites InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning.

DIVE: Taming DINO for Subject-Driven Video Editing InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.878296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.878296Z digest=sha256:13f613d41bd2c478fec19574e28356d2a1bbe952a373fd36d9ffa4243baa76ad

Observation f6a68224-9d6d-4291-b8ab-2eaefe04a2ad · outbound

This paper cites Learning universal semantic correspondences with no supervision and automatic data curation.

DIVE: Taming DINO for Subject-Driven Video Editing Learning universal semantic correspondences with no supervision and automatic data curation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:38.026217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.882192Z digest=sha256:44e3e62f35e5fc40d6dcc4e75ec3391c7c8b4b4d48d71fab88f43a67baf79975

Observation d88a466f-58c3-405a-8257-70876ccf8aad · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

DIVE: Taming DINO for Subject-Driven Video Editing Make-a-video: Text-to-video generation without text-video data

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.886104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.886104Z digest=sha256:4028d4d4150474c02f887be4f9f3a3efb404dde648dcf45e33d473d99d209f05

Observation f3d93aec-a5ef-469d-a0bc-5e4c8270059a · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

DIVE: Taming DINO for Subject-Driven Video Editing Deep unsupervised learning using nonequilibrium thermodynamics

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.890061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.890061Z digest=sha256:c9420dd9b2849552f8af32916f8792c224efbc94f2d59bf52b06fa1d8b496599

Observation 915da755-535b-4959-8363-b89e911ec62f · outbound

This paper cites Denois- ing diffusion implicit models.

DIVE: Taming DINO for Subject-Driven Video Editing Denois- ing diffusion implicit models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.893719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.893719Z digest=sha256:b644724083351ee7f8e0230924bb969f7f15302aa5758f015692492e9b43b423

Observation d28adcaa-876c-4802-9f1a-92e9465fe74a · outbound

This paper cites Generative modeling by esti- mating gradients of the data distribution.

DIVE: Taming DINO for Subject-Driven Video Editing Generative modeling by esti- mating gradients of the data distribution

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:37.918099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.897916Z digest=sha256:3e6d36bbadf8b2be1cad958c62c9201cc0ab844398239871288b45e27e09eb5d

Observation fd2120ae-c2c6-4c63-bf15-63c288643322 · outbound

This paper cites Score-based generative modeling through stochastic differential equa- tions.

DIVE: Taming DINO for Subject-Driven Video Editing Score-based generative modeling through stochastic differential equa- tions

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.901428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.901428Z digest=sha256:3e7a9350c25db1af14e2ae3490a3e3f34b4490d7c9ecce9507692f8ad037cfda

Observation 54710915-ee0f-47e4-b33e-128e47ffcf97 · outbound

This paper cites Diffusion Model-Based Video Editing: A Survey.

DIVE: Taming DINO for Subject-Driven Video Editing Diffusion Model-Based Video Editing: A Survey

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.905208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.905208Z digest=sha256:c80f6753a8e783c26fcb9e5a31448076f38e36ffcff5782fa7f0f75c295df6e1

Observation 34f7c607-2ddd-4715-a15a-c6e21fb08343 · outbound

This paper cites Emergent correspondence from image diffusion.

DIVE: Taming DINO for Subject-Driven Video Editing Emergent correspondence from image diffusion

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:37.896618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.919490Z digest=sha256:55e61467b0ab179009e1e97e2a8a22957aaae2367bdff78d96092e51c08b5d36

Observation c07b31e9-5d6a-40a8-a8f2-a1fc5d39e31b · outbound

This paper cites Key-locked rank one editing for text-to-image personaliza- tion.

DIVE: Taming DINO for Subject-Driven Video Editing Key-locked rank one editing for text-to-image personaliza- tion

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:37.883251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:36.972664Z digest=sha256:ad87fa4efadc1e21776b72f739ce3dac5076efae4c3f8e772b8a95ba012b1cdb

Observation 9e31dc42-e6a6-4d52-b2aa-15108840c360 · outbound

This paper cites Is this loss informative? faster text-to-image customization by tracking objective dynamics.

DIVE: Taming DINO for Subject-Driven Video Editing Is this loss informative? faster text-to-image customization by tracking objective dynamics

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:37.869427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:37.014064Z digest=sha256:cae8e2d800520492dcf7b4f8111a1a4855e9e453d527801f0f40a7eb69ae966f

Observation e4535634-422d-4850-bd77-1f3ea039e314 · outbound

This paper cites COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video Editing.

DIVE: Taming DINO for Subject-Driven Video Editing COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video Editing

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:37.047250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:37.047250Z digest=sha256:57286d869e9597dac3eab7175fb84e196fd96b83434e54ab9583623d60ff45a5

Observation d2747a08-5227-4d24-bc07-e7d6e47e6f8c · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

DIVE: Taming DINO for Subject-Driven Video Editing LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:37.100164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:37.100164Z digest=sha256:3c7d52df67a0fb51e6178305c06efe5e9b96adf6ca79acb10a08affc79cb7759

Observation 0c299ff2-ac62-49a1-b468-4e7ae9dce14e · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

DIVE: Taming DINO for Subject-Driven Video Editing InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:37.130216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:37.130216Z digest=sha256:e8ee77bdddd4be2ba96bab76b0e5c32fbb219493dc68f378e91cd6a988eb1a8e

Observation 9ab75395-0db3-43fa-88b7-226aaedf1e72 · outbound

This paper cites ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation.

DIVE: Taming DINO for Subject-Driven Video Editing ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:37.157503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:37.157503Z digest=sha256:68c025ee7f949af5dd7fa9ae4a5b0b8b980d11fe864a460339b671c83abc253e

Observation ea719f2c-8733-4b24-b1e6-253b0e68e18a · outbound

This paper cites Dreamvideo: Composing your dream videos with customized subject and motion.

DIVE: Taming DINO for Subject-Driven Video Editing Dreamvideo: Composing your dream videos with customized subject and motion

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:37.162495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:37.162495Z digest=sha256:886884a62bfd00021b087153cd1c40dae046ab2bed59869cc44a14ed07146d13

Observation 5537b191-1b84-40be-a4e3-790abcc8eb6c · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

DIVE: Taming DINO for Subject-Driven Video Editing Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:37.846780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:37.167401Z digest=sha256:cb232cda7e4598a7c325ea5bc04ff8c740ed63f67616de6ec85a66010b5ed4ee

Observation 0848e07f-1e90-4a3c-bb71-1d679e3fdbe8 · outbound

This paper cites CVPR 2023 Text Guided Video Editing Competition.

DIVE: Taming DINO for Subject-Driven Video Editing CVPR 2023 Text Guided Video Editing Competition

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:37.171823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:37.171823Z digest=sha256:d9575d3a442683efd91506912b934bd44d8293bb2ac8c5c60bd71f31a31b01e9

Observation 836f975d-d8c5-46c1-9f8a-31e4105e1b64 · outbound

This paper cites COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision Models.

DIVE: Taming DINO for Subject-Driven Video Editing COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:37.176530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:37.176530Z digest=sha256:d28495f9fed1919906cbbb2056871599d7d3a1e2057faba3e8e7e6d768284db2

Observation 31526673-141e-4b39-a5c5-a99699481481 · outbound

This paper cites A Survey on Video Diffusion Models.

DIVE: Taming DINO for Subject-Driven Video Editing A Survey on Video Diffusion Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:37.180569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:37.180569Z digest=sha256:141d1355c203662331031f518631767034517be092f003a00eae9731d5e9376a

Observation 99520b70-f89f-4eb2-8cc4-52a489c2aafe · outbound

This paper cites Ad- vancing high-resolution video-language representation with large-scale video transcriptions.

DIVE: Taming DINO for Subject-Driven Video Editing Ad- vancing high-resolution video-language representation with large-scale video transcriptions

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:37.832855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:37.185095Z digest=sha256:36b81e954d7dfe28856c35a9d471fca183b586adb17379fcf5b7d3afd83dc7b0

Observation ed17c25a-0067-442f-9afe-d33acb8fc16a · outbound

This paper cites Rerender a video: Zero-shot text-guided video-to-video translation.

DIVE: Taming DINO for Subject-Driven Video Editing Rerender a video: Zero-shot text-guided video-to-video translation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:37.818858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:37.189773Z digest=sha256:713703494111646bdc5bbefc5f619ff74e629e755773d045432f31410cd828b9

Observation 4eba989d-dcbf-4ea4-a6c3-b5fce6b180bd · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

DIVE: Taming DINO for Subject-Driven Video Editing IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:37.193927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:37.193927Z digest=sha256:203c1e9b8bdcb4cc9ec1649215a59800f79cef2c9dc1a53aaeb51dddb73106c9

Observation fefefe19-1914-464d-99ea-675884856c8f · outbound

This paper cites CustomNet: Zero-shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models.

DIVE: Taming DINO for Subject-Driven Video Editing CustomNet: Zero-shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:37.198319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:37.198319Z digest=sha256:023dce4286b8d57bbd8af7a99462a7f0e1b48e4a4f13c64319b9b43cb00e4350

Observation 50dd65dd-eddd-4779-bf9f-8de86b3efc75 · outbound

This paper cites A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence.

DIVE: Taming DINO for Subject-Driven Video Editing A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:37.805282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:37.202851Z digest=sha256:dbd498d156eab566f50fafe672d2d2cfc106f387ce5590e5353a913692c53d26

Observation d1748a41-a750-4449-a4e8-b76e22e393bf · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

DIVE: Taming DINO for Subject-Driven Video Editing Adding conditional control to text-to-image diffusion models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:37.207143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:37.207143Z digest=sha256:07fc549793f990cc9b8789adbcf281aff53b9aa8b409d77f6db913d18e502131

Observation a177d390-5a6a-4124-983d-e2e92ce533c1 · outbound

This paper cites Towards consistent video edit- ing with text-to-image diffusion models.

DIVE: Taming DINO for Subject-Driven Video Editing Towards consistent video edit- ing with text-to-image diffusion models

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:33:37.781600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:33:37.211236Z digest=sha256:c115470f2ccf92605b9506e59d417a1f7e80ecfd4e119e3db323923858965b4e

Observation dae329fb-b23d-4ca2-af67-02a78ed947cc · outbound

This paper cites ControlVideo: Conditional Control for One-shot Text-driven Video Editing and Beyond.

DIVE: Taming DINO for Subject-Driven Video Editing ControlVideo: Conditional Control for One-shot Text-driven Video Editing and Beyond

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:37.215743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:37.215743Z digest=sha256:636b0c7ba9b1b7e7dee107540e061da43ed9ed7807aba2404c7f0194fc1ce116

Pith citing papers

Observation 3b9050b3-9dce-45eb-94cf-192beaf90d9a · inbound

Component Adaptive Clustering for Generalized Category Discovery cites this paper.

Component Adaptive Clustering for Generalized Category Discovery DIVE: Taming DINO for Subject-Driven Video Editing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:49:15.079375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:49:15.079375Z digest=sha256:42f2e35bb28a4bab730846d8f3831901499d8c5bfc64604fc80f6e7e5b4ef9d9

Observation a6290e1d-7f53-4fd3-aaae-333fb8d2bff5 · inbound

Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance cites this paper.

Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance DIVE: Taming DINO for Subject-Driven Video Editing

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:32:36.204932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T14:30:18.778618Z digest=sha256:f81322c6131cacb8e95423c30d8bf39e56d2daf3e830283afef9d7f09c863ded