Pith. sign in

Paper Citation Record · LEDGER

Mimir: Improving Video Diffusion Models for Precise Text Understanding

As of 13 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 3 inbound Pith citation observations for arXiv:2412.03085.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03085 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:51:11.861299Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:09:07.991313Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T01:53:28.925576Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa3d78dd-059b-486f-854a-f07c0d1cf60d · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.612873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.612873Z digest=sha256:52ec2daaba9b53e258edf6ae703b6540c6fa6ac828833ce2148fa301fd591a04

Observation 174ce6c1-871c-49fe-be62-2eee06cdd298 · outbound

This paper cites Qwen Technical Report.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.617764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.617764Z digest=sha256:668a4f7b2ac6c40e39d2bfa75d1bb9fbf19f8b0f51faf85a1d5e94e0b864e02a

Observation 93c2fe9c-38a0-42f0-b0c2-532a69bedd52 · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

Mimir: Improving Video Diffusion Models for Precise Text Understanding eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.621761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.621761Z digest=sha256:bcd1b857e43348c7c6752673946e2600770b0ecb0741fff0f0170679aebe4472

Observation 31fbc9c4-a06d-4694-896e-45314cf6334a · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

Mimir: Improving Video Diffusion Models for Precise Text Understanding LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.626353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.626353Z digest=sha256:c3ef01dd4535f0c254c8dc053e05adc341c622effb706e3f7b868b24e59fc988

Observation a5492bb5-849a-4f41-a4fb-73ab32b8f523 · outbound

This paper cites Improving image generation with better captions.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Improving image generation with better captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.630859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.630859Z digest=sha256:8896424a27ab9b5b5b1ab4e0a5a8e34181606409dddb51e36e4fed5d8d26de26

Observation dc2703c4-4df9-446b-ae3f-c6d6f1b7fdb4 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.634791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.634791Z digest=sha256:a14105626b616fe65f58314e1aa00c4552c4cc952ee6f2ede8f2c0165d9c16d6

Observation f009921e-e3c9-4f5c-aa70-1002c1168767 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Mimir: Improving Video Diffusion Models for Precise Text Understanding PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.639492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.639492Z digest=sha256:2ae428c9dc87613280faf5a2a97dbbd72dfd460863ec69238a826946a0c444ff

Observation 5b167563-b1c0-4aff-a3e9-c7d1ffa9c571 · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

Mimir: Improving Video Diffusion Models for Precise Text Understanding PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.643736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.643736Z digest=sha256:b9907bb0a2ba4a701ce579771cf393831ec277dc08d2967f8a4b4aca2edbc26c

Observation de3a697d-3f46-4bb7-ad16-68da13c74268 · outbound

This paper cites PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.648408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.648408Z digest=sha256:04f8c785f3e9189368f36750141bc0e7af08d354a168348e192ccda31b6d2de8

Observation 66913913-901d-4840-89c9-15821d23ebca · outbound

This paper cites OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model.

Mimir: Improving Video Diffusion Models for Precise Text Understanding OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.651909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.651909Z digest=sha256:52d6b2517141e97b9bb0976c4e89a8f1edd3768214a6f051d2c821c8ff046027

Observation b244d4ae-c421-41b0-a980-c1150df05139 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Taming transformers for high-resolution image synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.656089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.656089Z digest=sha256:e718656b410e150bca0b41af032f7fe4220df9d66a4f92c9e030162088f71b68

Observation 62d3c6c2-d215-4a35-b549-8674f9f8b614 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Scaling rectified flow transformers for high-resolution image synthesis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.504267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.659739Z digest=sha256:036a289fd966888637a91b9329ddee6ca4236dee7766988077b912aa368a4be5

Observation a619a1c5-5332-4ef9-ba53-7b5841425048 · outbound

This paper cites Perceptual quality assessment of smartphone photography.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Perceptual quality assessment of smartphone photography

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.663489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.663489Z digest=sha256:07b5aa7f9e253f7a887ee186a47dfe0aab6d82cf1d863d2d6316c5f524b4da80

Observation e2a60417-8b0d-4c08-9027-32ec10e61a03 · outbound

This paper cites Ranni: Taming text-to-image diffusion for accurate instruction following.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Ranni: Taming text-to-image diffusion for accurate instruction following

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.488580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.667579Z digest=sha256:b0eb3bc9197219ef505498fffc3de7a61a30f5af0e5131e4c1f0545ddfa8d7ea

Observation 45b05fbf-8760-4267-bc57-45d5b7dfed17 · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.670823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.670823Z digest=sha256:0fd5fb21268be6ac3547e20a6490e0d73aa921613c4eff50ed5bda36ec025806

Observation c3cf3bea-0787-4c18-a181-d61f21012c25 · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Preserve your own correlation: A noise prior for video diffusion models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.477864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.674621Z digest=sha256:b507aadabc700040922395ed8a17899141b80b16df440aea4a25d18d9d526398

Observation 8db5b61a-cb49-4097-bf14-00676a7ecaf5 · outbound

This paper cites Check locate rectify: A training- free layout calibration system for text-to-image generation.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Check locate rectify: A training- free layout calibration system for text-to-image generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.466479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.678765Z digest=sha256:76b532266d3de805b8fa18e6c2d4feb4fe90a5c8ef6bae59c7a9edb72509962c

Observation 7a1b328f-6ac9-4d11-a578-f0ca4f9f0ed1 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Mimir: Improving Video Diffusion Models for Precise Text Understanding AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.682448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.682448Z digest=sha256:298d7a5e70fac445e0848ee4c0d2bb4d663f9fced468f3da0492cc6c7e3af592

Observation 7d66e66d-b89d-4fc4-8936-e018f8b2aada · outbound

This paper cites Denoising dif- fusion probabilistic models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Denoising dif- fusion probabilistic models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.686779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.686779Z digest=sha256:f1f9e4eba682fe585adc3e327b70b9212233f6c1ba994aeb6a0268606619e6e6

Observation d8128cdc-9f30-49be-9d2c-1ecf7b7ddc93 · outbound

This paper cites Video dif- fusion models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Video dif- fusion models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.690080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.690080Z digest=sha256:65d2a47a72029780005c3243f51d707b1521def5bef96d10b154a2b0cf4f92fe

Observation 72645895-e98b-4a7f-a8a6-0a3b48f4dc0d · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

Mimir: Improving Video Diffusion Models for Precise Text Understanding CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.693552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.693552Z digest=sha256:b76917a2e3820b80cc55134ec0b18f3cac0ed3a986f2f408c4dacd88e797726d

Observation a21b01a0-49ff-4cb7-b178-40b54e03990d · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Mimir: Improving Video Diffusion Models for Precise Text Understanding CogVLM2: Visual Language Models for Image and Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.697842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.697842Z digest=sha256:cc8fc0ac6ae5409d57def19de455dbfab16a6d1f3e03240b97ea674e2beb34b2

Observation 22e6d8b2-de04-450e-83c8-346d24e7b021 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Mimir: Improving Video Diffusion Models for Precise Text Understanding ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.700744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.700744Z digest=sha256:bb2af417fe5448d9a7f009055aec203b3fefd5ae7d22aef9156d1bbbe84edf7c

Observation 6baa9656-9778-4f21-8fb8-8ab3caa37b8b · outbound

This paper cites T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- 9 tion.

Mimir: Improving Video Diffusion Models for Precise Text Understanding T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- 9 tion

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.444116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.703692Z digest=sha256:7432691f8cd1a7618dd489613ab3861340c04663ad100c17f4edec8d2eed92f8

Observation ebc51882-8616-4c0c-bf7d-49852386588f · outbound

This paper cites VBench: Com- prehensive benchmark suite for video generative models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding VBench: Com- prehensive benchmark suite for video generative models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.431850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.706469Z digest=sha256:0d282e1e75ecc934375eaf4937f3c26773f6d0ca8a7ded28844f4b4f3d51150f

Observation 6239cb00-0320-49f7-9520-3e90afdea811 · outbound

This paper cites Musiq: Multi-scale image quality transformer.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Musiq: Multi-scale image quality transformer

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.420034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.709085Z digest=sha256:73df299c167a7a9a04e93b61cc67fcc4315c2c55156ef86fda59aad993dba48a

Observation b88ade9b-eb5f-425a-a074-e21a22a4ae57 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Mimir: Improving Video Diffusion Models for Precise Text Understanding VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.711501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.711501Z digest=sha256:7885f685ffbf220ef2193f41c5f3c88089730055f90e5fbe26aebc337b990d05

Observation 397fdc8b-b211-4ac7-a6b7-8bfe676f77f5 · outbound

This paper cites Open-sora-plan, 2024.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Open-sora-plan, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.409609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.714777Z digest=sha256:91eae64d6e9783c59f2bf1dc919df9be9164c5430394e823ee6a4044aaf879d4

Observation 2220ec05-9a82-44ac-a17d-68ab422d66f6 · outbound

This paper cites Common diffusion noise schedules and sample steps are flawed.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Common diffusion noise schedules and sample steps are flawed

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.718305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.718305Z digest=sha256:01e48cf93d93a1a039ea1ef6df7c31c96edfd2591b40613004d909180e019ab6

Observation d22d68cb-0418-4f31-a267-d067f09854cd · outbound

This paper cites Videofusion: Decomposed diffusion models for high-quality video generation.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Videofusion: Decomposed diffusion models for high-quality video generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.394518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.721349Z digest=sha256:450f86331cb0013e99766be9a5412242dffa38477c1be00b8bfbbbd51fc72adf

Observation f1f1dd64-9159-4d62-bedc-24cca4f2eeb3 · outbound

This paper cites Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.725237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.725237Z digest=sha256:38ead558561b44ac5968f98ed44f88f139c340cc5b5abab803f3578ff2cc6b56

Observation 94af8bdb-1df9-49ce-be09-ab6e4bf4e9aa · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Mimir: Improving Video Diffusion Models for Precise Text Understanding SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.729407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.729407Z digest=sha256:f2202dd78c061c27154e7a857c0e6d103066cd684025c33b06fd26d262e9c2ef

Observation caf6a1f7-5e7c-418c-b5eb-ce2e29ce478f · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Learning transferable visual models from natural language supervi- sion

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.384419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.733125Z digest=sha256:63cbe235d4606910378e19771bf263dcd2b6554576c10fcf0b860b05a51e191b

Observation 9cdd973b-fe4b-4b4c-ba09-61c8c29dd0c8 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.736887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.736887Z digest=sha256:7cadb300205e741eb7d35536560b07cd01e873485a163e8fb2a6e3761ab81f7b

Observation ca035896-122e-4992-b697-583963c91b40 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.740143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.740143Z digest=sha256:2443466890675ce7e8167b9c2c71dd9621d20dba1f0616354afa4550c5d205ac

Observation f3ae12f5-5f68-42c8-b040-8a8916b2c07b · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding High-resolution image synthesis with latent diffusion models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.367811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.743465Z digest=sha256:48276a67e8200bf52cd5860e20d49b4316be8d765779e1f743157f346f81b309

Observation 3cc13530-d148-47f9-a88c-90caab8fef4e · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Photorealistic text-to-image diffusion models with deep language understanding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.358061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.746344Z digest=sha256:2656c3c74633f5ea44b99db866ad4d66e624ff231d9c58052d74f730abd347ec

Observation 604fe72a-990a-44bf-a5b3-5b2f919b16ff · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Progressive Distillation for Fast Sampling of Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.749400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.749400Z digest=sha256:4e6aab17f67b90ff2e77fb7e4d87d48c038dd8bf8983600b7a51e957db18d596

Observation 51edff70-7685-45e7-81c3-482dcdb6b7cb · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Deep unsupervised learning using nonequilibrium thermodynamics

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.753568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.753568Z digest=sha256:ed1f47ab9abd711855d079ad24a6d0358c7a50be73c7479c7d90afddeb7f112f

Observation 2d940805-e791-4ff3-8f01-eb3f97354397 · outbound

This paper cites Stable Diffusion 2.0 Release, 2022.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Stable Diffusion 2.0 Release, 2022

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.341114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.757006Z digest=sha256:2caa5d12309671bf7a6f4dda59990bba14ffa6a11be685334f8590ad297806ba

Observation 8c4b294e-7380-4d32-bce9-ee3cf89529b9 · outbound

This paper cites Galip: Generative adversarial clips for text-to-image synthesis.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Galip: Generative adversarial clips for text-to-image synthesis

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.328642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.760544Z digest=sha256:e7e27ce94ff20bd3ae68fa1fdd14fff0caff11011a118751175cd68d076d6110

Observation 9b869e0b-9656-42a2-a5a3-2275c644060a · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.763823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.763823Z digest=sha256:23273e3c4c3d5e2ac2548a302d0a9080e8b918f7c8d453d8d1bd9c04f62b177d

Observation 14280fb3-5cc0-48cf-b81b-a137f4750877 · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities, 2023.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Internlm: A multilingual language model with progressively enhanced capabilities, 2023

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.767558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.767558Z digest=sha256:a3cd809af08e36a2ca5174b71bcc06046d9159384662dc20890209ab4dd48f7f

Observation 8c8393dd-ca0a-4920-b893-cfee5060a49e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.770731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.770731Z digest=sha256:e1ed8f3aae73eaada60c15b91503bdcd89e3738eb3f3a31175aa900c477196e1

Observation ebbd1558-1adf-4a4b-8507-97a12bac64dc · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.774507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.774507Z digest=sha256:a6684c7e92945403caf239f8ff64859d13aeac97f9c3e69ae21743c2d89d7a48

Observation d83c2f09-c7e3-4a8f-ac8b-42d26310f4ac · outbound

This paper cites Visualizing data using t-sne.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Visualizing data using t-sne

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.778460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.778460Z digest=sha256:f215058e4dcecef3ab762d7351cc73ad2fa43873ff0d0cbb89e5a38979fa6f39

Observation 370095dc-e1e5-452a-ba97-170d31db0697 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models, 2023.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Cogvlm: Visual expert for pretrained language models, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.305573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.781986Z digest=sha256:89597b5e22c459e88e4ccabf1505ed6e9a6eb86f0c66d1fe9fa0d356090a20ca

Observation b4c52a21-8243-4808-973b-25604af2bee3 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Mimir: Improving Video Diffusion Models for Precise Text Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.786109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.786109Z digest=sha256:172edba7c84fac86e21f2ec221bfd614acf531f7f003bec1c964c33e9837a5f3

Observation a17778bf-b914-4a82-96f6-5f13fd06f7bf · outbound

This paper cites Grit: A gener- ative region-to-text transformer for object understanding.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Grit: A gener- ative region-to-text transformer for object understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.295080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.790031Z digest=sha256:739f01a1f855975cd846ac647f709465e8fd61e975032c00fc54a8826b482a86

Observation 078ce21b-e4e4-407b-9b82-e25718a751a9 · outbound

This paper cites Paragraph-to-Image Generation with Information-Enriched Diffusion Model.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Paragraph-to-Image Generation with Information-Enriched Diffusion Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.793880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.793880Z digest=sha256:1322b94deee3d8554de1e9cb3e3a087b856b52498002c1c079fd94776e3311ae

Observation 02eba41d-9f32-41c5-9774-7c3e884688df · outbound

This paper cites SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers.

Mimir: Improving Video Diffusion Models for Precise Text Understanding SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.797701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.797701Z digest=sha256:d117f152c149247e0849ee325224986a8cd03b820c013ede7742ec2ccba36e93

Observation b644eab3-2884-4903-a021-e10675749e39 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Baichuan 2: Open Large-scale Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.800770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.800770Z digest=sha256:3f757e46e3a685eed65d669663128319b8ad5f7bc13b61a26af3f9c811763976

Observation c6430e2a-71bf-494e-8f17-c1d6284c3712 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Mimir: Improving Video Diffusion Models for Precise Text Understanding CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.804287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.804287Z digest=sha256:96f60efda31a51242f36e9572642e56e6bd17d4c83d241c4caef964f45eb4ba0

Observation f7e04b10-0996-4aa7-9f3e-29316a62060f · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Yi: Open Foundation Models by 01.AI

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.807591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.807591Z digest=sha256:939822c1afb36128f97c48f57b6e9e3350362f11acd9c405c640962d767589da

Observation 2a3bf3b5-821a-428c-8285-f8604899fcdb · outbound

This paper cites Magvit: Masked generative video transformer.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Magvit: Masked generative video transformer

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.284845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.810768Z digest=sha256:8e617a388e5d325e49acd7dc91ae599e52b51e17911cfe0e84015b2ac3e40d38

Observation d22537aa-a24e-42e8-bc1b-85ac610b88c6 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.813821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.813821Z digest=sha256:aa4e330c6de86069951d5e7dd4a571b7935dfd9c6ad7a426d2b6cfd7f2698f2b

Observation 68908d34-bc99-438d-bc82-1be8c17d0042 · outbound

This paper cites Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.274098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.818765Z digest=sha256:18d20ef10a2a7d050da31bbd9b3fa91e8b83dec3628bdd6280e46083c08e4edb

Observation 560ddeec-0fa3-42e1-8d32-b12b9fe9caf4 · outbound

This paper cites Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.822595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.822595Z digest=sha256:b86b328c7dba694410bb246792c3c76da5e27c7dc43f5fa799d61daefe75a3b1

Observation f36d674f-0dae-4596-95b0-33d22e2afff5 · outbound

This paper cites CV-VAE: A Compatible Video VAE for Latent Generative Video Models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding CV-VAE: A Compatible Video VAE for Latent Generative Video Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.827049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.827049Z digest=sha256:86da723a41c45d8e18b1ab51eab7b213b55ee4fbe0327fd45b19467fba0db89a

Observation 1199afe8-f167-4d7b-abc3-248dd0a6b825 · outbound

This paper cites Open-sora: Democratizing efficient video production for all, march 2024.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Open-sora: Democratizing efficient video production for all, march 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.263393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.831286Z digest=sha256:e5634a5c93a5209a78e5c4d54e276cc354dbe3b557c20db0236ebd3db378ecaf

Observation 2fde7757-3226-4491-89de-2b0d958a469c · outbound

This paper cites an unresolved cited work.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:51:12.254031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.835420Z digest=sha256:0b702b96ce98d4a23cdf9382847f8240fd983be8b1eed439a548ab8b09323953

Observation f956b5f7-5c53-49a0-8f7e-41a835d0d893 · outbound

This paper cites • Videos with a motion score of 0, determined using optical flow, are excluded.

Mimir: Improving Video Diffusion Models for Precise Text Understanding • Videos with a motion score of 0, determined using optical flow, are excluded

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.243362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.839179Z digest=sha256:6a4d1e447313b3e67c7070454afd58aa519d9e8858b1ef7de6e544d4dbe9bbd9

Observation 8515fc69-f386-4409-8cea-ba987206a161 · outbound

This paper cites an unresolved cited work.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:51:12.231669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.843069Z digest=sha256:97e0dbd7b0bc37e3f10a2659ff57d7918e71edce75adbf12f1136dfe5dc2fed4

Observation ed2bb0e4-2c56-40f2-a982-cc8cd9bf3c35 · outbound

This paper cites Input text prompt.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Input text prompt

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.217479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.847314Z digest=sha256:e591a821c539ffd6ffa3a6c20ac875fdb7c3d7ad2023171b75ecd2995a36df57

Observation 537e7d6e-a8cd-4598-83db-03bae41de380 · outbound

This paper cites an unresolved cited work.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:51:12.206404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.850609Z digest=sha256:4b553aa76e05958bdf49751abf1a2009ebe3e3f623f1c00fa19c3ad13983cc99

Observation a4417213-4970-4ef3-9ecb-4ba35cbc2b33 · outbound

This paper cites containing watermarks.

Mimir: Improving Video Diffusion Models for Precise Text Understanding containing watermarks

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.196304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.854420Z digest=sha256:a466bf295f37d890ee27fcc2f2bbcaaa955946b199ee1d676eb8208e69849502

Observation b87fc7b9-720d-48af-8501-26cd4788d9ce · outbound

This paper cites an unresolved cited work.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:51:12.186046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.858303Z digest=sha256:33d2921d5d373ec512289f7544617b3437789135bd60729f95774b414f42f4e9

Observation a99ad1bb-e85d-4473-823a-b87f226ba4cb · outbound

This paper cites top”, “ below.

Mimir: Improving Video Diffusion Models for Precise Text Understanding top”, “ below

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.173566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:51:11.861299Z digest=sha256:882ce030f319794df018c2c781c67c78ebbdc8abd53c7fcee5602b796a2cbe03

Pith citing papers

Observation 71f99619-f953-4740-9ac9-36f852f104e7 · inbound

Animate-X++: Universal Character Image Animation with Dynamic Backgrounds cites this paper.

Animate-X++: Universal Character Image Animation with Dynamic Backgrounds Mimir: Improving Video Diffusion Models for Precise Text Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T21:09:07.991313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:09:07.991313Z digest=sha256:e68dd911a167759afc04c9059dd0b68be73e43f73717c1843aebbc7ee21c8dbd

Observation ed581e4b-c3b5-4117-9062-2b5d37bed3de · inbound

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis cites this paper.

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis Mimir: Improving Video Diffusion Models for Precise Text Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T19:07:34.083942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:07:34.083942Z digest=sha256:2a685fee509b30aff392a3d40312da67af71e079e1895c7c57b963641b56f872

Observation 8513a963-5d84-4bb1-b70d-40f6565b1a0b · inbound

Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction cites this paper.

Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction Mimir: Improving Video Diffusion Models for Precise Text Understanding

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:53:28.927343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T01:51:57.018809Z digest=sha256:a29171ac1702fd9666508aa1437e3bcdd6a6e048b6e6656d85c4b7c9658758df