Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T00:46:08.276886Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 4 inbound Pith citation observations for arXiv:2501.18096.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T00:46:08.276886Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:41:39.719320Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-16T17:53:11.781235Z
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 07461521-f5d5-4dbc-ba99-b9eca04c3b9d · outbound
LLMs can see and hear without any training write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f40734fc-2f39-4238-b0e9-b4ac7c89a2c0 · outbound
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f5ce520-c345-42cb-8b39-daedf969618c · outbound
LLMs can see and hear without any training SPICE : S emantic propositional image caption evaluation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 259fe07d-7b7f-43c5-89ef-e5b091e28ad5 · outbound
LLMs can see and hear without any training and Lavie, A
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f7f04a0f-559b-4c16-a6f7-d1f502c76582 · outbound
LLMs can see and hear without any training Improving image generation with better captions
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cb5f418e-b65d-4717-9e60-9717e8746d21 · outbound
LLMs can see and hear without any training W., Fidler, S., and Kreis, K
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d80eda5-09d1-4c14-a4db-130921020026 · outbound
LLMs can see and hear without any training Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67571005-a8df-4362-aab6-cacfa53cad31 · outbound
LLMs can see and hear without any training Imagenet: A large-scale hierarchical image database
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ec30906-4b7d-4993-9436-66a84a2a57e2 · outbound
LLMs can see and hear without any training Clotho: An audio captioning dataset
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8c0cfefa-d255-443d-87bb-1d43bbba6d22 · outbound
LLMs can see and hear without any training The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d19565a3-573a-4ece-93c4-787d72b96cfe · outbound
LLMs can see and hear without any training M., Jain, A., Schmidt, L., Toshev, A., and Shankar, V
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8dda7dcb-7a38-4118-9091-8a4fc12923ac · outbound
LLMs can see and hear without any training Interpreting the Second-Order Effects of Neurons in CLIP
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24436c06-49ae-487f-bbbd-1af9ed9e7331 · outbound
LLMs can see and hear without any training A Neural Algorithm of Artistic Style
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee6ca5fe-c59e-4946-ab2d-0d88b34aa8a9 · outbound
LLMs can see and hear without any training On the content bias in fr \'e chet video distance
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 54066178-ac0e-4576-b742-67253af18245 · outbound
LLMs can see and hear without any training F., Ellis, D
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bfefb861-2462-44e1-91b4-ca65e854ed40 · outbound
LLMs can see and hear without any training V., Joulin, A., and Misra, I
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation da8027aa-beb1-4bf2-a24d-36f98151b87b · outbound
LLMs can see and hear without any training S., Shah, A., Yin, X., Parikh, D., and Misra, I
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdd064a1-cf92-4143-9fd5-2bf110ec5155 · outbound
LLMs can see and hear without any training Mmg-ego4d: Multimodal generalization in egocentric action recognition
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2aa4606c-ac40-4d58-b040-be7ec5246a39 · outbound
LLMs can see and hear without any training Audioclip: Extending clip to image, text and audio
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0741e06d-d761-4090-95da-3600bbee57e9 · outbound
LLMs can see and hear without any training Imagen Video: High Definition Video Generation with Diffusion Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20abe8e5-1b44-4608-a80a-2f1ca44b0f92 · outbound
LLMs can see and hear without any training Openclip, 2021
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 50a54a14-ec58-44f2-b3be-4933370cbdc2 · outbound
LLMs can see and hear without any training Rethinking fid: Towards a better evaluation metric for image generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e327ff0a-71c0-455a-8b32-53de2061b349 · outbound
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e04e41e-d73b-49e9-a3d0-baeb2bf2f162 · outbound
LLMs can see and hear without any training and Fei-Fei, L
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 45df32dc-b331-40fc-8f74-81b7e508733e · outbound
LLMs can see and hear without any training What do we learn from inverting CLIP models?
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45345009-5da6-443e-8f01-318e7bb14201 · outbound
LLMs can see and hear without any training Pick-a-pic: An open dataset of user preferences for text-to-image generation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 50f1f474-5d21-4fa5-9df9-987b465c0df9 · outbound
LLMs can see and hear without any training S., Reid, M., Matsuo, Y., and Iwasawa, Y
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 688e3352-7798-4a11-b2f5-d99be38bd7b3 · outbound
LLMs can see and hear without any training Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 36a6ddfb-a057-4d2b-b35b-cbcf5ab1102b · outbound
LLMs can see and hear without any training Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 52c1ea4a-947f-4dd1-ad4c-5160baaccb34 · outbound
LLMs can see and hear without any training Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ac1e12be-fce0-4996-a55a-dcd92156cbe6 · outbound
LLMs can see and hear without any training M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaddedd7-5f35-4e5b-9764-503fad42fcb9 · outbound
LLMs can see and hear without any training DeCap : D ecoding CLIP latents for zero-shot captioning via text-only training
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 14a4ec48-b4ce-41cd-b469-16522d447a2b · outbound
LLMs can see and hear without any training Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cbe38f1-f7c2-43ad-ad6f-6b85627af581 · outbound
LLMs can see and hear without any training Flow Matching for Generative Modeling
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34264a0c-d475-4189-b526-fced26b8aac1 · outbound
LLMs can see and hear without any training Improving Text-to-Image Consistency via Automatic Prompt Optimization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 593a6f57-3e95-44bb-a976-34d8cea1126d · outbound
LLMs can see and hear without any training Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1033a21d-83e2-4edf-aeba-44da51e4bdd5 · outbound
LLMs can see and hear without any training How T o100 M : L earning a T ext- V ideo E mbedding by W atching H undred M illion N arrated V ideo C lips
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aa943601-8824-46bc-8f50-368d8e29c6bf · outbound
LLMs can see and hear without any training Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 396a3532-bf3b-4daf-a936-b2a2bd6c225c · outbound
LLMs can see and hear without any training H., Seybold, B., Hauth, A., Manen, S., Sun, C., and Schmid, C
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5dddf899-95dd-414b-b6d0-038afb215443 · outbound
LLMs can see and hear without any training Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 10d47875-b6e7-4f60-83ca-41e7e4e945bf · outbound
LLMs can see and hear without any training Introducing openai o1-preview
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b7918d94-7097-4632-a4d4-a19cd2c21e94 · outbound
LLMs can see and hear without any training BLEU : A method for automatic evaluation of machine translation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c4876983-387a-4576-a046-bea434a1ba8d · outbound
LLMs can see and hear without any training Movie Gen: A Cast of Media Foundation Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d52ba24e-5c4a-4635-ac20-4c094d5ec01d · outbound
LLMs can see and hear without any training W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9baacafb-a287-4ef2-b54d-1cbb561e5bed · outbound
LLMs can see and hear without any training Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa7d852a-d09a-49d4-bb44-801c457c92ca · outbound
LLMs can see and hear without any training High-resolution image synthesis with latent diffusion models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0874f5f2-fb01-4a21-af90-4158959c883e · outbound
LLMs can see and hear without any training L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., Ho, J., Fleet, D
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7728eb58-c5f2-4394-9cd9-8e7d72f1312f · outbound
LLMs can see and hear without any training Zero-shot audio captioning with audio-language model guidance and audio context keywords
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 566a6112-59e8-49d9-b567-3629ee02ecd8 · outbound
LLMs can see and hear without any training Zero-Shot Audio Captioning via Audibility Guidance
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da075a6e-f406-4cee-b340-a85ccdcbdc6c · outbound
LLMs can see and hear without any training H., Tan, H., Bansal, M., Rohrbach, A., Chang, K.-W., Yao, Z., and Keutzer, K
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 266e9844-c78d-4d8b-9996-41f9ca698977 · outbound
LLMs can see and hear without any training Emu edit: Precise image editing via recognition and generation tasks
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 93a29c0a-a145-4515-bed3-dfa693881bb7 · outbound
LLMs can see and hear without any training and Zisserman, A
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f0924f5-8430-486a-9718-8c1f6a86e3d4 · outbound
LLMs can see and hear without any training Gemma: Open Models Based on Gemini Research and Technology
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45254f0e-ab4f-49be-85de-a6f2c7572252 · outbound
LLMs can see and hear without any training Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5138eade-5d48-4783-bde5-0238c7b8605c · outbound
LLMs can see and hear without any training CIDEr : C onsensus-based image description evaluation
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ba63aaaa-e306-448c-9b55-4807c0e9e945 · outbound
LLMs can see and hear without any training Internvid: A large-scale video-text dataset for multimodal understanding and generation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3c11ff3c-043c-47d7-b8b2-5a1e44f33196 · outbound
LLMs can see and hear without any training V., Zhou, D., et al
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 44818b35-f4e7-42e0-bd36-ffc37988b3c4 · outbound
LLMs can see and hear without any training E., Huang, P.-Y., Howes, R., Sharma, V., Li, S.-W., Ghosh, G., Zettlemoyer, L., and Feichtenhofer, C
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dc653a20-dfca-4130-9302-d9e9532d66bb · outbound
LLMs can see and hear without any training MSR-VTT : A large video description dataset for bridging video and language
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 69bb8738-4a48-403b-ab0e-cbe856820531 · outbound
LLMs can see and hear without any training V., Zhou, D., and Chen, X
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3c9f03ad-4cec-483f-ad2e-a15fc1827c61 · outbound
LLMs can see and hear without any training ConZIC : Controllable zero-shot image captioning by sampling-based polishing
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4ed7780e-fb70-4000-b3f7-e11c6360ea04 · outbound
LLMs can see and hear without any training Meacap: Memory-augmented zero-shot image captioning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fc8b6ceb-0b37-4262-907d-dc8bf1a118a5 · outbound
LLMs can see and hear without any training Sigmoid loss for language image pre-training
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88cde246-7d4e-4957-9ac3-b77d3112e4f5 · outbound
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 07f8e3c3-0c4b-49da-9c4b-4fc59bc741a6 · inbound
SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents LLMs can see and hear without any training
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 234690bb-8c5c-4572-865e-afc404bf07a3 · inbound
Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models LLMs can see and hear without any training
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7cc7435-15c9-461a-aad3-117776b30975 · inbound
It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models LLMs can see and hear without any training
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 03fd9dca-c859-42c1-bd48-dfcfb2d0f5bb · inbound
Personalizing Text-to-Image Generation to Individual Taste LLMs can see and hear without any training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.