Pith. sign in

Paper Citation Record · LEDGER

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation

As of 5 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2606.03168.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.03168 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T11:04:49.366127Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact27
  • verified fuzzy0
  • unresolved8
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7eeccabf-dd94-47ca-8a6f-ab26a3f47d90 · outbound

This paper cites Insvie-1m: Effective instruction-based video editing with elaborate dataset construction.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Insvie-1m: Effective instruction-based video editing with elaborate dataset construction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T11:04:49.366127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:62971289a356a60361b157854d16c358d4f1f71b0c1245ccd7042ba35787792b

Observation 18a0ee24-8692-4f78-8519-c5aa8200a789 · outbound

This paper cites arXiv preprint arXiv:2512.07826 , year=.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation arXiv preprint arXiv:2512.07826 , year=

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.743190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:ee40e179b458bb1a5a324b80e3f352d49cd33312b82db95b29135351ed269fc2

Observation f1e7ddff-77f5-44a3-9b6f-c9a8c5c9d552 · outbound

This paper cites arXiv preprint arXiv:2510.15742 , year=.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation arXiv preprint arXiv:2510.15742 , year=

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:16:26.745870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:4775505b01f4c4fa6fe2cff0d56552a3476ea366c639b62a6f72ab1641fc2459

Observation 5f227e3c-16a9-46de-8416-4b5504f19ecd · outbound

This paper cites Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.764211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:914706a5552ff5f0fd77b63239d5e054f029b3b1f42e3a781de581ae7a87ff28

Observation 512d2ac6-b1bb-4edd-839e-0009085e3bd9 · outbound

This paper cites Av-edit: Multimodal generative sound effect editing via audio-visual semantic joint control.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Av-edit: Multimodal generative sound effect editing via audio-visual semantic joint control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T11:04:49.366127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:03d86670591bcc27269cb47750e6c0286de73d032f2c3eb9bd6a2f272e6a26ff

Observation a2f8037c-11a4-4a69-a147-76ccc93413ca · outbound

This paper cites AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.759100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:b0d898333f3fb11793f038c6890aacde855bb5bc60b3e97e3df3192787da67b9

Observation dbcecc2b-82a8-4dd0-b96a-955cd28af8cb · outbound

This paper cites Openhumanvid: A large-scale high-quality dataset for enhancing human-centric video generation.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Openhumanvid: A large-scale high-quality dataset for enhancing human-centric video generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T11:04:49.366127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:00db4f3b6e9fea7afe976ac80482688938a980cdba9519cff3da9e980d4a100e

Observation f9d4710d-4d57-43be-8766-893118d790e6 · outbound

This paper cites VidGen-1M: A Large-Scale Dataset for Text-to-video Generation.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.776236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:62b3ba42f73fc29cfbb68481949f2b4b9684b0bb92709f0690d805cbfe819030

Observation 47715f08-aabd-4995-a65b-740a9166c7d4 · outbound

This paper cites VGGSound: A Large-scale Audio-Visual Dataset.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation VGGSound: A Large-scale Audio-Visual Dataset

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.809374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:2391c916bf4cae1b39eeeeaecefa86476dae08b33f289719d35037c37dc5294b

Observation 12d17b61-ec44-445a-9033-1d22877cfc74 · outbound

This paper cites Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T11:04:49.366127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:2e83d3da0bd891fbf8385db96a424aac06b26707184c8d6fd37ec7a26966fe45

Observation 733d0f7c-5215-441b-bd55-a469b9bdf976 · outbound

This paper cites Qwen3-Omni Technical Report.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Qwen3-Omni Technical Report

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.750991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:8c2f0357a4d2c4a46ad881273e2f6f25fb7bfa27986a1780c28e5961c83658df

Observation 5f42ee24-a374-4854-a97f-1a3ca3ed2eab · outbound

This paper cites Mel-RoFormer for Vocal Separation and Vocal Melody Transcription.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Mel-RoFormer for Vocal Separation and Vocal Melody Transcription

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.773429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:3b1a0e73cfa3f707004ab78c28f486bfbe50fc505a05242260749085c094535f

Observation 3a3124cf-28ce-4e0f-9763-d40762330633 · outbound

This paper cites ZeroSep: Separate Anything in Audio with Zero Training.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation ZeroSep: Separate Anything in Audio with Zero Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.754001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:aa1010cf141cbdf486699c343238fed4581a56bb444283a1a412d4b805451ef0

Observation 3b7734ac-fe9d-442d-a325-c37dfa514464 · outbound

This paper cites Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian, Mike Zheng Shou, and Haizhou Li.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian, Mike Zheng Shou, and Haizhou Li

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:16:26.732613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:416bd8239ec4d7b5332fd3bfa18f37789f1fde7de8772657bb5ff30b8e87d7d7

Observation 935d5965-417a-4a9e-b7d6-f02832cef876 · outbound

This paper cites Qwen3 Technical Report.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Qwen3 Technical Report

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.761375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:a7347065e1a04f7b9f2a02354bcf51088a2ae7e072f1dab03de32f8557f5491b

Observation 80f263da-5825-4738-b7af-39bff0ff780e · outbound

This paper cites HunyuanImage 3.0 Technical Report.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation HunyuanImage 3.0 Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.756423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:4c32cf4567c6720980216dea90b4cb2cd73d187ef174b08b3b1988d9590d0b47

Observation 9ac78d33-4697-4e05-8e13-51fc775aab86 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.770215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:f973f428805554b017b4092e646a50398939f6e8ccc8acbf622ec7a35b0a94ec

Observation d655ff30-74f8-4484-96db-181507b2e4d5 · outbound

This paper cites DreamVoice: Text-Guided Voice Conversion.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation DreamVoice: Text-Guided Voice Conversion

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.797282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:0bd572ce2037e7e86917341f5aef050b60c70d19b40227cf3f1bf953214d055b

Observation 965398af-e867-4dfa-9a4f-5e3aaec6a554 · outbound

This paper cites Ffp-300k: Scaling first-frame propagation for generalizable video editing.arXiv preprint arXiv:2601.01720, 2026.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Ffp-300k: Scaling first-frame propagation for generalizable video editing.arXiv preprint arXiv:2601.01720, 2026

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.747725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:222c02e428f90840eb7e87c3b1bd4a414c2cb55580b410aee4a75402ff1803ea

Observation aa346e1f-6b7a-4698-84f3-0395901c5310 · outbound

This paper cites HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.744449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:6e7bdb6382e04f80da318139eceb65eff3b2b93985c4b9acfd465512af268fc1

Observation 806f97f2-d01d-4411-a459-559b7a892109 · outbound

This paper cites MiniMax-Remover: Taming Bad Noise Helps Video Object Removal.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation MiniMax-Remover: Taming Bad Noise Helps Video Object Removal

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.812069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:858107ee4f01215b3dcd1ad956f439173c82b34d08c7f8a00cb7b3f0db3061a4

Observation 2069844f-6987-4678-8437-c4648e3b40ad · outbound

This paper cites SAM 3: Segment Anything with Concepts.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation SAM 3: Segment Anything with Concepts

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.804202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:28a5bb8db05ba7fbbb7cbf87a6712b7db41b4ce9d5ae5dab27ec77779413b407

Observation 0c992d99-9660-4583-94db-41731bc63462 · outbound

This paper cites Qwen3-TTS Technical Report.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Qwen3-TTS Technical Report

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.741492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:fa04d9d0ae1f09428969cad1249ac25fef05e14b2ca1baec82d2eaec21def3c7

Observation ed43b2f1-d428-4817-bc4d-5b6c76f75484 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.806856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:a037a18e003f2110962e9b24509782a850f79f55c5b74d553d92ec27e37a377f

Observation ec41fd86-9262-41f2-9019-b5c3cabf12bd · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.791937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:ecfaaa44d35fbb2ae70077e7f1c04f6906f946529b7f0d1aa56deb2cb198ff39

Observation 47ff6187-ee6a-405d-a300-ccca218c79d6 · outbound

This paper cites The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.789450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:4c8b1dff0267bb05b51dfdd0eebcf559f4e612256269b09774e54da11f2c0acd

Observation 9c1beeb9-2b51-4a73-ab23-95bbac9f59d7 · outbound

This paper cites Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.783808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:d6edcbec8d667226ab47f6e1f0415e33c5164e2982cd2cdca19f201042df64d5

Observation 6f9db44e-0142-48b3-96b7-56b4b8944e60 · outbound

This paper cites Metaclaw: Just talk–an agent that meta-learns and evolves in the wild.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Metaclaw: Just talk–an agent that meta-learns and evolves in the wild

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.778819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:21076b34525f364f26c70a4ed999abf4d4641c5f8f1a98c6e83990af632a9c0c

Observation 72ff8f0f-2504-47d9-a95b-02adaf134aa1 · outbound

This paper cites 2023 , keywords =.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation 2023 , keywords =

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T11:12:00.839746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:1e96a86c4290512311564352f391106545f542444dd59f28f7b8cc6ac8295920

Observation 1ae7171d-17c2-4077-88e4-8db78d3feb49 · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T11:04:49.366127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:2320e108cc71587140bf6e1d985a4b020956e22bdd8322e1812190da791a938a

Observation 9ccf29e2-36ea-42ba-a99e-cb35d05c7cee · outbound

This paper cites Scalable diffusion models with transformers.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Scalable diffusion models with transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T11:04:49.366127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:c7183e6efc99391384c74831f3d9fe16dc857ea35340da7b9fb544459a380389

Observation fc659e16-7cc5-4acd-963a-3e5a40732356 · outbound

This paper cites arXiv preprint arXiv:2511.18822 (2025).

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation arXiv preprint arXiv:2511.18822 (2025)

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.786773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:1b66ea73efc2562251f2616e48c3af1cf3d1271b5dabd3276203ed2822785ab8

Observation 3fc6c7ef-5ca3-43f1-9db4-0b6139758611 · outbound

This paper cites Classifier-Free Diffusion Guidance.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Classifier-Free Diffusion Guidance

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.781353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:2b97743a3032a40f16f2cc39ced781018a8c287d1a1b05ac79a5577bba28f7f6

Observation 0e6a18fc-fd70-4e50-a77b-975544d28f03 · outbound

This paper cites Ragd: Regional-aware diffusion model for text-to-image generation.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Ragd: Regional-aware diffusion model for text-to-image generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T11:04:49.366127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:9f8e3815a7999a7b5493536007066c0e1b66df857454acbe48a8cc1408fb5f8e

Observation c496c4b1-d363-43df-8fa6-283a2691de02 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Score-Based Generative Modeling through Stochastic Differential Equations

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.794697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:238e722cda354d257b1b776dc59b1d8fb8bda93726adadd9a2bec659aa5444f0

Observation c26f449c-96fa-42ca-bdec-25df508e6230 · outbound

This paper cites L2P: Unlocking Latent Potential for Pixel Generation.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation L2P: Unlocking Latent Potential for Pixel Generation

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.799539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:b2f846abf1657a162f09d23ff9733efe0d2a87b854ad3c7515e86c26c9e698b1

Observation e6b72f66-a01c-4fde-9929-6efabc1cdd15 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation SAM 2: Segment Anything in Images and Videos

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.766566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:72f54088d6daa2994134626a13aaacbe461c11e973697f0c8dfeaaa27afd2275

Observation e63be31f-b2b5-4d49-b4cd-f47b0ab46695 · outbound

This paper cites Ivebench: Modern benchmark suite for instruction-guided video editing assessment.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Ivebench: Modern benchmark suite for instruction-guided video editing assessment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T11:04:49.366127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:58407a31fb54538c8aa401eaacf0c37d1775b18b1b1ecaa830f8d04cb91d83fe

Observation b2f06081-76d7-47ec-8aa8-c52e4c44a8d0 · outbound

This paper cites hard to distinguish.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation hard to distinguish

Reference 42

Resolution
malformed identifier
arxiv_id, observed 2026-07-02T02:16:26.801917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:fa441109c7d77311fece903565ef3ec63f7f27f0d47d6e63fdb44f19f791cead

Pith citing papers

No inbound Pith citation observations are available.