{"paper":{"title":"Active Sampling for Ultra-Low-Bit-Rate Video Compression via Conditional Controlled Diffusion","license":"http://creativecommons.org/licenses/by/4.0/","headline":"Sparse adaptive keyframes and tracked trajectories let a conditional diffusion model reconstruct video at much lower bitrates while keeping perceptual quality high.","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Amirhosein Javadi, Shirin Saeedi Bidokhti, Tara Javidi","submitted_at":"2026-05-04T17:25:14Z","abstract_excerpt":"Diffusion models provide a powerful generative prior for perceptual reconstruction at ultra-low bitrates, but effective video compression requires controlling the generative process using highly compact conditioning signals. In this work, we present ActDiff-VC, a diffusion-based video compression framework for the ultra-low-bitrate regime. Our method partitions videos into variable-length segments, transmits keyframes only when needed, and summarizes temporal dynamics using a compact set of tracked point trajectories. Conditioned on these sparse signals, a conditional diffusion decoder synthes"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Experiments on the UVG and MCL-JCV benchmarks show that ActDiff-VC achieves up to 64.6% bitrate reduction at matched NIQE, improves KID by up to 64.6% and FID by up to 37.7% at comparable bitrates against strong learned codecs, and delivers favorable perceptual rate--distortion trade-offs relative to learned and diffusion-based baselines in the ultra-low-bitrate regime.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"The assumption that sparse conditioning signals consisting of content-adaptive keyframes and budget-aware tracked point trajectories are sufficient for a conditional diffusion decoder to synthesize perceptually realistic and temporally coherent video frames without introducing major artifacts or inconsistencies.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"ActDiff-VC achieves up to 64.6% bitrate reduction at matched NIQE and improves perceptual metrics like KID and FID by using content-adaptive keyframe selection and budget-aware sparse trajectory selection to condition a diffusion decoder for ultra-low-bitrate video reconstruction.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Sparse adaptive keyframes and tracked trajectories let a conditional diffusion model reconstruct video at much lower bitrates while keeping perceptual quality high.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"bf103f9fbb213939e4ef03d1163e2a8991c4a35cb28042eaefe15c8207158738"},"source":{"id":"2605.02849","kind":"arxiv","version":2},"verdict":{"id":"7f10d38f-55be-4bf4-8e4b-f4df579d6d53","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-08T18:26:03.996352Z","strongest_claim":"Experiments on the UVG and MCL-JCV benchmarks show that ActDiff-VC achieves up to 64.6% bitrate reduction at matched NIQE, improves KID by up to 64.6% and FID by up to 37.7% at comparable bitrates against strong learned codecs, and delivers favorable perceptual rate--distortion trade-offs relative to learned and diffusion-based baselines in the ultra-low-bitrate regime.","one_line_summary":"ActDiff-VC achieves up to 64.6% bitrate reduction at matched NIQE and improves perceptual metrics like KID and FID by using content-adaptive keyframe selection and budget-aware sparse trajectory selection to condition a diffusion decoder for ultra-low-bitrate video reconstruction.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"The assumption that sparse conditioning signals consisting of content-adaptive keyframes and budget-aware tracked point trajectories are sufficient for a conditional diffusion decoder to synthesize perceptually realistic and temporally coherent video frames without introducing major artifacts or inconsistencies.","pith_extraction_headline":"Sparse adaptive keyframes and tracked trajectories let a conditional diffusion model reconstruct video at much lower bitrates while keeping perceptual quality high."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2605.02849/integrity.json","findings":[],"available":true,"detectors_run":[{"name":"ai_meta_artifact","ran_at":"2026-05-20T14:40:45.661336Z","status":"completed","version":"1.0.0","findings_count":0},{"name":"doi_title_agreement","ran_at":"2026-05-20T02:31:21.924063Z","status":"completed","version":"1.0.0","findings_count":0},{"name":"doi_compliance","ran_at":"2026-05-19T15:55:17.674955Z","status":"completed","version":"1.0.0","findings_count":0}],"snapshot_sha256":"6aefba89733ed96ab4089d47b778b55735389cbbc42786d54ef89d9fc742b5c1"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":2,"snapshot_sha256":"744c47153be2d16120b2d4a0d3803aaaa7abfa578eca5578722b526936e1bd13"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}