Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T20:34:09.157279Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 1 inbound Pith citation observation for arXiv:2411.09439.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T20:34:09.157279Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:30:16.892900Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T20:30:17.239571Z
75 of 75 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fc0976e2-2d16-46de-8495-5023b430a0c4 · outbound
Spider: Any-to-Many Multimodal LLM In OpenAI, 2023
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a5c16b3d-3e44-4993-810e-d5b224dabe9b · outbound
Spider: Any-to-Many Multimodal LLM Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 520e7c5f-8bff-49fd-b0ed-d1902156d88c · outbound
Spider: Any-to-Many Multimodal LLM Latent-shift: Latent diffusion with temporal shift for efficient text-to-video generation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ab0e6d23-bf05-4c63-a50d-13f985ca4cdf · outbound
Spider: Any-to-Many Multimodal LLM Blended latent diffusion
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 712f352d-1d74-4a97-91db-04d942dc3165 · outbound
Spider: Any-to-Many Multimodal LLM Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 766ab9fb-82fe-4737-b6b1-610e9841c718 · outbound
Spider: Any-to-Many Multimodal LLM Lan- guage models are few-shot learners
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4cd44ab-966a-4103-bd18-7c4e22a4e498 · outbound
Spider: Any-to-Many Multimodal LLM Zeroscope: Diffusion-based text-to-video syn- thesis
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 77f94e95-45bc-4bbd-99bd-86df5b734659 · outbound
Spider: Any-to-Many Multimodal LLM Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c0b5f94e-37e5-4f08-a6d4-162ab1eef096 · outbound
Spider: Any-to-Many Multimodal LLM Gonzalez, Ion Stoica, and Eric P
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 166ba334-f9fe-48ed-b9f7-f0a1aa2cb932 · outbound
Spider: Any-to-Many Multimodal LLM Palm: Scaling language modeling with pathways
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d933fbab-fd8d-48fb-9252-bc6e1bfbf8ad · outbound
Spider: Any-to-Many Multimodal LLM Diffedit: Diffusion-based semantic image editing with mask guidance
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1622171f-0a85-46e8-8131-a66a4707f045 · outbound
Spider: Any-to-Many Multimodal LLM Bert: pre-training of deep bidirectional trans- formers for language understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f304cf2b-ec49-44e2-85e8-fdecea0f4b0c · outbound
Spider: Any-to-Many Multimodal LLM Cogview: Mastering text-to- image generation via transformers
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dd802106-ee98-464d-bf2c-faea0d87b241 · outbound
Spider: Any-to-Many Multimodal LLM Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61cc2a65-b053-4325-b271-ded83992ac79 · outbound
Spider: Any-to-Many Multimodal LLM Preserve your own correlation: A noise prior for video diffusion models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 39a849bd-c566-46a8-98b9-f487de559b85 · outbound
Spider: Any-to-Many Multimodal LLM AudioSet: An ontology and human- labeled dataset for audio events
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a0437587-179c-4cf3-ac12-9dbd763ccc13 · outbound
Spider: Any-to-Many Multimodal LLM Imagebind: One embedding space to bind them all
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 075def10-a591-4d96-9e10-a00ab5279689 · outbound
Spider: Any-to-Many Multimodal LLM Au- tomated audio captioning by fine-tuning BART with audioset tags
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 238794fc-1f6f-427d-be6b-d283a977b5e6 · outbound
Spider: Any-to-Many Multimodal LLM Onellm: One framework to align all modalities with language
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 60c8f2f2-e864-4675-bf50-59d2540e185e · outbound
Spider: Any-to-Many Multimodal LLM Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1f73eac1-81cf-4018-8b8e-7701e949f239 · outbound
Spider: Any-to-Many Multimodal LLM Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9ccbb427-5e81-4b33-b843-abd08b6b23fc · outbound
Spider: Any-to-Many Multimodal LLM Audiogpt: Understanding and generating speech, music, sound, and talking head
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4c7ea450-4049-4413-9bd7-ba35c5f49eb9 · outbound
Spider: Any-to-Many Multimodal LLM Language is not all you need: Aligning perception with language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2afe10a7-cb84-4c36-a1d9-04953cbf11fb · outbound
Spider: Any-to-Many Multimodal LLM Pfb-diff: Progres- sive feature blending diffusion for text-driven image editing
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 50b21481-6238-4476-85d3-157deef15429 · outbound
Spider: Any-to-Many Multimodal LLM Audiocaps: Generating captions for audios in the wild
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9abc238d-65b3-4dd3-9e7d-8b9ec81434b2 · outbound
Spider: Any-to-Many Multimodal LLM Improving audio-language learning with mixgen and multi- level test-time augmentation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 998a3544-f4e8-44bb-9bda-7d73d6239a6e · outbound
Spider: Any-to-Many Multimodal LLM Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 68ad3f1b-db77-4935-a552-519e7d387678 · outbound
Spider: Any-to-Many Multimodal LLM Gen- erating images with multimodal language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b9d6411b-f207-4397-99d2-f241640dd74e · outbound
Spider: Any-to-Many Multimodal LLM Are diffusion models vision-and- language reasoners? In Thirty-seventh Conference on Neural Information Processing Systems, 2023
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cee7fa5a-80fa-4d82-b2d2-1908f245d9e0 · outbound
Spider: Any-to-Many Multimodal LLM Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 412160ca-95f0-4a8d-92c4-93d154f41327 · outbound
Spider: Any-to-Many Multimodal LLM Videochat: Chat-centric video understanding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a9ca3b2b-d14b-4167-907a-93853157632d · outbound
Spider: Any-to-Many Multimodal LLM Oscar: Object-semantics aligned pre-training for vision-language tasks
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7ccc856e-e814-4b1a-abc4-7aa6c6f16c62 · outbound
Spider: Any-to-Many Multimodal LLM VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8358131f-0679-4d5d-a903-f4dd95f17692 · outbound
Spider: Any-to-Many Multimodal LLM Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a313f540-6bc6-4e18-9d1c-b08a200cf3d5 · outbound
Spider: Any-to-Many Multimodal LLM Mandic, Wenwu Wang, and Mark D
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b33a1fae-9986-4895-9652-26cd808cab83 · outbound
Spider: Any-to-Many Multimodal LLM Improved baselines with visual instruction tuning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5b26ae3c-05fa-40b3-83fd-87c9c6d7faf0 · outbound
Spider: Any-to-Many Multimodal LLM Visual instruction tuning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8b707845-0070-41bd-9cad-17dac61ffe51 · outbound
Spider: Any-to-Many Multimodal LLM Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5be99fd5-d219-44dd-99f5-1424ad17f4c1 · outbound
Spider: Any-to-Many Multimodal LLM Sdedit: Guided image synthesis and editing with stochastic differential equa- tions
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c32d38e-bb4c-40f7-b312-0d01ebcd2f94 · outbound
Spider: Any-to-Many Multimodal LLM Audio-journey: Open domain latent diffusion based text-to-audio genera- tion
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3c27fd9b-7ca4-4505-8eda-e3ca6c3456a5 · outbound
Spider: Any-to-Many Multimodal LLM GLIDE: towards photorealis- tic image generation and editing with text-guided diffusion models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8be672d8-a044-403e-b7aa-e45da14d6a5e · outbound
Spider: Any-to-Many Multimodal LLM Introducing chatgpt
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7ef45537-6ae6-4558-933f-102e5ea72481 · outbound
Spider: Any-to-Many Multimodal LLM Gross, and Alexander Sorkine- Hornung
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bc91d3dd-eaa1-4fa0-ba9b-2b54317a4023 · outbound
Spider: Any-to-Many Multimodal LLM Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ce979068-4b8d-4868-98be-41ca2f008dac · outbound
Spider: Any-to-Many Multimodal LLM Discriminative probing and tuning for text-to-image generation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dbde3744-b669-4155-a1b6-15cbe8c4024c · outbound
Spider: Any-to-Many Multimodal LLM Language models are unsu- pervised multitask learners
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 06808218-3fa6-48cd-afa3-2003b2210368 · outbound
Spider: Any-to-Many Multimodal LLM High-resolution image syn- thesis with latent diffusion models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 72f1d907-2720-4506-a45a-81d1affdd979 · outbound
Spider: Any-to-Many Multimodal LLM Talking about large language models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 51acdada-fb45-4107-a4e1-32e4374a59a6 · outbound
Spider: Any-to-Many Multimodal LLM Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ddb230e4-9917-410a-909a-64c124ad6dfe · outbound
Spider: Any-to-Many Multimodal LLM Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2fc207c6-e93c-4a82-aa49-24603a8472d0 · outbound
Spider: Any-to-Many Multimodal LLM Make-a-video: Text-to-video generation without text-video data
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fd3bdb4c-d298-4238-b49f-be9628a7a14e · outbound
Spider: Any-to-Many Multimodal LLM UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af46836e-fc3f-4fae-8024-d86bddb60e42 · outbound
Spider: Any-to-Many Multimodal LLM Lan- guage models can see: Plugging visual controls in text gen- eration
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 334befef-3ef3-443c-84fb-7b31632802d1 · outbound
Spider: Any-to-Many Multimodal LLM Pandagpt: One model to instruction-follow them all
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c8fb613d-3121-4800-b3c3-ee896ba2d032 · outbound
Spider: Any-to-Many Multimodal LLM Any-to-any generation via composable diffu- sion
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 30870797-96ab-4e91-a31c-e98c2e46fbae · outbound
Spider: Any-to-Many Multimodal LLM Galactica: A large language model for science
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0551c22d-96cd-4b88-b179-a0f161af4e3d · outbound
Spider: Any-to-Many Multimodal LLM Gemini: a family of highly capable multimodal models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2ec57542-e34c-44eb-82f7-014367412fc1 · outbound
Spider: Any-to-Many Multimodal LLM Llama: Open and efficient foundation language models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b3f2d3bf-d972-4d4a-8cbb-1ecb806b36d5 · outbound
Spider: Any-to-Many Multimodal LLM Llama 2: Open foundation and fine-tuned chat models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2a5bec41-d5dd-4191-a617-ce771d08bee8 · outbound
Spider: Any-to-Many Multimodal LLM Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 89d33028-0227-4c98-be1f-56578b5497dd · outbound
Spider: Any-to-Many Multimodal LLM GIT: A generative image-to-text transformer for vision and language
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 461d48b8-71d4-46f7-a40f-15e24bef1b26 · outbound
Spider: Any-to-Many Multimodal LLM Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 290ded71-e602-4753-9df8-80228250801b · outbound
Spider: Any-to-Many Multimodal LLM Campnet: Context-aware mask prediction for end- to-end text-based speech editing
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 820c89e5-9c00-4678-85d9-15ba3204d9d7 · outbound
Spider: Any-to-Many Multimodal LLM Visual chatgpt: Talking, drawing and editing with visual foundation models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f3482a93-9430-4576-9cf7-3e51f9ea277c · outbound
Spider: Any-to-Many Multimodal LLM Tune-a-video: One-shot tuning of im- age diffusion models for text-to-video generation
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation af9eb693-1002-480f-9eb7-50fa7046290a · outbound
Spider: Any-to-Many Multimodal LLM Next-gpt: Any-to-any multimodal llm
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0003e81f-25dd-4ee5-aa48-79f9cea3b850 · outbound
Spider: Any-to-Many Multimodal LLM mplug-2: A modularized multi-modal founda- tion model across text, image and video
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 17f10101-2ba0-4cb5-abe9-0257bed5a4af · outbound
Spider: Any-to-Many Multimodal LLM MSR-VTT: A large video description dataset for bridging video and lan- guage
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6f79a270-668f-4181-b6bc-04a9d9639e7c · outbound
Spider: Any-to-Many Multimodal LLM Diffsound: Discrete dif- fusion model for text-to-sound generation.IEEE ACM Trans
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a3796435-b038-41bd-883e-6a210aba67ed · outbound
Spider: Any-to-Many Multimodal LLM Speechgpt: Empow- ering large language models with intrinsic cross-modal con- versational abilities
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fa1d3c69-b0cc-43c0-97c4-cdd92099050f · outbound
Spider: Any-to-Many Multimodal LLM Video-llama: An instruction-tuned audio-visual language model for video un- derstanding
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cc7c4f99-a4ff-46d0-bcbe-762f9bc2bc39 · outbound
Spider: Any-to-Many Multimodal LLM Object relational graph with teacher-recommended learning for video captioning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e40a7b48-6890-407d-a034-9eb22b966aea · outbound
Spider: Any-to-Many Multimodal LLM Minigpt-4: Enhancing vision-language understanding with advanced large language models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bac3282e-acb2-486b-bdfc-b61e3a3b3a74 · outbound
Spider: Any-to-Many Multimodal LLM Thus, the modalities generation performances of our Spider are limited by the integrated Decoder models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 49b10b23-3b5c-4c6f-920b-b5f861ed235f · outbound
Spider: Any-to-Many Multimodal LLM Image-to-Text generation on COCO-caption [34]
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0fbc1ba3-14af-48e3-b4d2-2f4459be40ff · inbound
A Unified Multi-Agent Framework for Universal Multimodal Understanding and Generation Spider: Any-to-Many Multimodal LLM
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.