Pith. sign in

Paper Citation Record · LEDGER

Testing chatbots on the creation of encoders for audio conditioned image generation

As of 7 August 2026, this Paper Citation Record lists 100 of 123 outbound references and 0 inbound Pith citation observations for arXiv:2509.09717.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09717 v1

Coverage vector

measured 100 of 123 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:25:27.288300Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 123 outbound references displayed

  • verified exact5
  • verified fuzzy22
  • unresolved72
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f2773827-75a3-47c3-820d-11bd236c39a0 · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.568746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.568746Z digest=sha256:b84bb9b832819313fe46c27ac91594115f3205018de45c25f4d9975ff751b1ec

Observation a3196e02-8a17-4eae-b2ca-7b4c4fccbee3 · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.575691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.575691Z digest=sha256:48a3363685eee0bfd93d824dac85a623bc7ed6bac1d80a2050955dcc09d6433f

Observation a05e9955-7c72-4f18-a048-e542adba6200 · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.582775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.582775Z digest=sha256:f0ece32d414f41caad8b71051d945e030da3c1a44247a82e5116066a7bb88df9

Observation 3e732957-04b3-454e-a69a-2081c3d8a33a · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.590262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.590262Z digest=sha256:d4712428a40dff59072c5d98e2c1ee90920841e568c4397fb8e66e2a704782b5

Observation 73eda8d4-c2cc-4b6a-8a19-cbb7a2817182 · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.596268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.596268Z digest=sha256:f67c465130f41b586f664bf428bd61cf74daf9604b543199d05f10561642f299

Observation fdecdf9e-7d72-433a-a58f-cc59b3d5ccaa · outbound

This paper cites MusicLM: Generating Music From Text.

Testing chatbots on the creation of encoders for audio conditioned image generation MusicLM: Generating Music From Text

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.603213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.603213Z digest=sha256:7267addca3207fb94d45405bb815a185fc9d27956b72e64df7c02123776fdd54

Observation 30b1b9fa-3cbc-4db3-8af1-1fa9a452d78d · outbound

This paper cites Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering.

Testing chatbots on the creation of encoders for audio conditioned image generation Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.610281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.610281Z digest=sha256:96fde0231e37903252b186174fee8e8ddd2961ffe3a37e33d2f4e49f437f63c0

Observation 9dc492ba-08cb-4972-9fc4-c7814d9eb3fa · outbound

This paper cites Mistral Models, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Mistral Models, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.616045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.616045Z digest=sha256:06044482aca7a04b456002383326b10b87948d78a3d5b58f3d3e09e43e103f01

Observation a0aa1203-489f-4ccc-a811-9e9ff8a99709 · outbound

This paper cites Transcripter- Generation of the transcript from audio to text using Deep Learning.International Journal of Computer Sciences and Engineering, 7(1):770–773, 2019.

Testing chatbots on the creation of encoders for audio conditioned image generation Transcripter- Generation of the transcript from audio to text using Deep Learning.International Journal of Computer Sciences and Engineering, 7(1):770–773, 2019

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.622613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.622613Z digest=sha256:6f513a8d6ab373982870e892b70a7985da239bf195cd55ee0f01cd20366b438a

Observation f6c6ed46-b1e8-4bda-80ed-e5bb48a66962 · outbound

This paper cites The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.630294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.630294Z digest=sha256:eb8833ae9d6efe2f97dd7fbacfef7ead3f302e20b37645a3183351ffc989650f

Observation 264247d8-e039-464b-8b52-798c1768a358 · outbound

This paper cites Claude 3.7 Sonnet and Claude Code, 2025.

Testing chatbots on the creation of encoders for audio conditioned image generation Claude 3.7 Sonnet and Claude Code, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.635144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.635144Z digest=sha256:583ecdb2013aed55068da124a708e47e1363085721dcd4c2c8ad96cb50bb9495

Observation af19f626-c13b-44ea-a39a-4bc7577bf874 · outbound

This paper cites AudioSetCaps: An Enriched Audio-Caption Dataset using Auto- mated Generation Pipeline with Large Audio and Language Models.

Testing chatbots on the creation of encoders for audio conditioned image generation AudioSetCaps: An Enriched Audio-Caption Dataset using Auto- mated Generation Pipeline with Large Audio and Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.641339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.641339Z digest=sha256:7802915053417e35ec6c9d76f853a5daafae5e6a262dc7d346d3b7cb8464fcfc

Observation 1b3e2439-fd12-4a7b-b049-34ec221eb8f9 · outbound

This paper cites Are Mod- els Biased on Text without Gender-related Language? InProceedings of the 12th International Conference on Learning Representations, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Are Mod- els Biased on Text without Gender-related Language? InProceedings of the 12th International Conference on Learning Representations, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.649251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.649251Z digest=sha256:16c43e82723706f6e2be6faa98de3d71f66930e758487d3e6d1603138af07b26

Observation c9e1f668-242a-4ffd-b03c-3a86f74b18cd · outbound

This paper cites Ballester.

Testing chatbots on the creation of encoders for audio conditioned image generation Ballester

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.654893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.654893Z digest=sha256:6c1fd4bdca3860a514e13477d64d8cf444ceeac8cdebac75c73d8fa37aad3dbf

Observation 1fd800a7-ff39-4098-8915-37b54e294be5 · outbound

This paper cites Improving Image Generation with Better Captions.

Testing chatbots on the creation of encoders for audio conditioned image generation Improving Image Generation with Better Captions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.661088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.661088Z digest=sha256:7a3598887ec0d36a911332bc45525f18b094f4a5ddefd5e6766dcbf133064619

Observation 20547dc6-b008-4d2c-947b-a66803d7055c · outbound

This paper cites RenAIssance: A Survey into AI Text-to-Image Generation in the Era of Large Model.

Testing chatbots on the creation of encoders for audio conditioned image generation RenAIssance: A Survey into AI Text-to-Image Generation in the Era of Large Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.666579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.666579Z digest=sha256:0b2b2559a04c4719cbafcacecdc76eda234fd0c5ff4fc632b39ac3d2f9cd0ad2

Observation fcc6e1ec-fb5d-4440-9ba7-70abb84a231e · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.674010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.674010Z digest=sha256:00259be873810afe8f05605f1db788917ff41cead5dd6e387d3536097fdd89bb

Observation ab37063a-de8f-43b6-bfe1-d16d0dc28503 · outbound

This paper cites A contemporary review on chatbots, AI-powered virtual conversa- tional agents, ChatGPT: Applications, open challenges and future research directions.

Testing chatbots on the creation of encoders for audio conditioned image generation A contemporary review on chatbots, AI-powered virtual conversa- tional agents, ChatGPT: Applications, open challenges and future research directions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.680984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.680984Z digest=sha256:ab997b1d70f4151ee758214ad530917e3d98f3d2429944f548f04743e2682714

Observation 82dcdb46-bde2-4d42-b8ce-53cdd125d978 · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.690305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.690305Z digest=sha256:eaaea8159c4cb8e5c54abb427d0606d9d57c6cf90e044cec0cef349f2eb3a481

Observation bc3c8152-5272-4f11-a4e6-d8dcb2bad7cb · outbound

This paper cites Veo, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Veo, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.697479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.697479Z digest=sha256:6762a7d9abab2cb3944300db9deae8c281f328be24045105e130796ba9437097

Observation 309f9b93-2acc-4a5f-a337-cf7a05d6abc1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Testing chatbots on the creation of encoders for audio conditioned image generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.706108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.706108Z digest=sha256:a7e02f62cda8ed9fffe8a5300402d04b0a1a43c9ef851ac975461f3933fb71a7

Observation 8de0c676-9fd4-422e-ac60-8fa77d590a9b · outbound

This paper cites A Survey of On-Device Machine Learning: An Algorithms and Learning Theory Perspective.ACM Transactions on Internet of Things, 2(3), 2021.

Testing chatbots on the creation of encoders for audio conditioned image generation A Survey of On-Device Machine Learning: An Algorithms and Learning Theory Perspective.ACM Transactions on Internet of Things, 2(3), 2021

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.712700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.712700Z digest=sha256:2bcfb49476ec19d9bef39438e93b50e0edf35c14699475cc7309ca8997da9b7f

Observation 2ef99852-0c8f-45c3-a4b6-81286b25bc31 · outbound

This paper cites Jukebox: A Generative Model for Music.

Testing chatbots on the creation of encoders for audio conditioned image generation Jukebox: A Generative Model for Music

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.718047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.718047Z digest=sha256:6c90f8cfe4b0773c6a75dbad72899d9b634ebd1c1fbcfea45444756f904f6299

Observation fc4d511e-d4f1-4943-bbe4-4729dfd730d9 · outbound

This paper cites The Llama 3 Herd of Models.

Testing chatbots on the creation of encoders for audio conditioned image generation The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.723712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.723712Z digest=sha256:8db797adb4963950b5e8158d59f7869fa85aaf993cf367bc36ad2c4553b77375

Observation 3da38730-450b-4774-b7e8-e809ea7f9be0 · outbound

This paper cites Grok, Gemini, ChatGPT and DeepSeek: Comparison and Applications in Conversational Artificial Intelligence.

Testing chatbots on the creation of encoders for audio conditioned image generation Grok, Gemini, ChatGPT and DeepSeek: Comparison and Applications in Conversational Artificial Intelligence

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.730575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.730575Z digest=sha256:fb15f57658c3832602260c2c6799e3ad19f83fac8a8e60b1d63116f3092e732d

Observation 9f2c13e2-b696-45c9-82cd-af72e946ec29 · outbound

This paper cites Image Generation: A Review.Neural Processing Letters, 54(5):4609–4646, 2022.

Testing chatbots on the creation of encoders for audio conditioned image generation Image Generation: A Review.Neural Processing Letters, 54(5):4609–4646, 2022

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.738911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.738911Z digest=sha256:359291dcd17bd727066bfdc5cec155111b4aa7acc07f7572a92f1c28dcb60533

Observation fa9b58c6-6e37-4aa8-ba87-ba78c7eb8c3d · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Testing chatbots on the creation of encoders for audio conditioned image generation Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.744606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.744606Z digest=sha256:dedb1a80f2da08d5ad9b898c3a739cea489390cd9758f97d5676df0e37a2329a

Observation ee78b31a-84de-4bf3-bd58-33552e137ab4 · outbound

This paper cites Learning From Noisy Correspondence With Tri-Partition for Cross-Modal Matching.IEEE Transactions on Multimedia, 26:3884–3896, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Learning From Noisy Correspondence With Tri-Partition for Cross-Modal Matching.IEEE Transactions on Multimedia, 26:3884–3896, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.750603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.750603Z digest=sha256:82790d680b09550b12308a803ef2ec5ac6a2667e2b0a7cf5356f036cbf1c16b9

Observation 047ae5d0-4bd9-4baf-8a9d-3790b8fe4c05 · outbound

This paper cites Line Goes Up? Inherent Limitations of Benchmarks for Evaluating Large Language Models.

Testing chatbots on the creation of encoders for audio conditioned image generation Line Goes Up? Inherent Limitations of Benchmarks for Evaluating Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.758106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.758106Z digest=sha256:bd98e26d3342e3fa20ef576394cbb851dd8c2bd9a71716e0e051b1d475971eeb

Observation ad942ac4-e177-4d3f-921c-3aa1448bd327 · outbound

This paper cites Creativity and Machine Learning: A Survey.

Testing chatbots on the creation of encoders for audio conditioned image generation Creativity and Machine Learning: A Survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.764195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.764195Z digest=sha256:da8f694ccd4820cba4619eec1a52254c6641f3a7af07ba42f8e018c0646e858b

Observation bc1a4d8b-6bd8-46f1-9916-ab35fab94150 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Testing chatbots on the creation of encoders for audio conditioned image generation The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.772408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.772408Z digest=sha256:9d63aca51f0937ccda01cc4469fc913ed205791c57627da538fa2a62129b19e8

Observation 09e8c540-3a19-48e0-8d72-22511bc0c912 · outbound

This paper cites ImageBind: One Embedding Space To Bind Them All.

Testing chatbots on the creation of encoders for audio conditioned image generation ImageBind: One Embedding Space To Bind Them All

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.782500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.782500Z digest=sha256:e0c010c5e0020fab2d0ba1ead3d32add6a0a933026363ec94dfa2261815197cc

Observation 119adbf1-bf88-4802-853f-34a2ac03bf28 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Testing chatbots on the creation of encoders for audio conditioned image generation Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.792335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.792335Z digest=sha256:7cd920537fc2793ea6965edc6c46caf0dc8d57fc1fbb2c012c78f539b0f41edf

Observation 406eddbf-0f26-4e61-a2d9-0a886814416a · outbound

This paper cites AudioCLIP: Extending CLIP to Image, Text and Audio.

Testing chatbots on the creation of encoders for audio conditioned image generation AudioCLIP: Extending CLIP to Image, Text and Audio

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.798563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.798563Z digest=sha256:de362f90b468078371b61c22ee7ea43cdde4b9b895c35a8bba2f776fc069e3cc

Observation ffbee515-f413-429f-9cdd-85b23af3b599 · outbound

This paper cites Deep Residual Learning for Image Recognition.

Testing chatbots on the creation of encoders for audio conditioned image generation Deep Residual Learning for Image Recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.806461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.806461Z digest=sha256:d21f18b015791517352283c9670ab6def2081b257628e057648f2daf3a3e7f31

Observation d0ee86c4-a0bb-4610-862c-89a99d38fcfb · outbound

This paper cites Ringle, and Rudolf R.

Testing chatbots on the creation of encoders for audio conditioned image generation Ringle, and Rudolf R

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.812512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.812512Z digest=sha256:4d6bf1d505abc0b2a4578fbdd3c4f4e64c2ad161e136aee0ed3db394a151be80

Observation 4d61857a-ddd6-4619-a87e-a0626dc1af99 · outbound

This paper cites Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model.

Testing chatbots on the creation of encoders for audio conditioned image generation Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.818748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.818748Z digest=sha256:9a3af462462112b9c52acaa9c475c31490e202b9f6eec0eef261ade890595865

Observation cc085099-2c40-44a1-8e42-403256fd10a6 · outbound

This paper cites Make-an-audio: text-to-audio generation with prompt-enhanced diffusion models.

Testing chatbots on the creation of encoders for audio conditioned image generation Make-an-audio: text-to-audio generation with prompt-enhanced diffusion models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.825041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.825041Z digest=sha256:64ecd6126402535274a903f696b8365ded75420ff47c6ed7c42111d32e2c86e0

Observation df0d6c3f-633e-4981-8fb9-65f2ecdd678f · outbound

This paper cites NLIP: Noise-Robust Language-Image Pre-training.

Testing chatbots on the creation of encoders for audio conditioned image generation NLIP: Noise-Robust Language-Image Pre-training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.832038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.832038Z digest=sha256:4133c9e1458779282feafe57b30f41be60879ee2561f85d44d271c5f4953100a

Observation 98e98fc0-7756-49d1-a249-36c0ef3a0af8 · outbound

This paper cites Large Language Models for Code Generation: A Comprehensive Survey of Challenges, Techniques, Evaluation, and Applications.

Testing chatbots on the creation of encoders for audio conditioned image generation Large Language Models for Code Generation: A Comprehensive Survey of Challenges, Techniques, Evaluation, and Applications

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.837379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.837379Z digest=sha256:c4e4636b3aa0fcaf7952b8075de65e2a25a6619020c9d0be7ccde3b1835db06f

Observation 46dae24a-4596-4be1-a69b-016b472057f1 · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.844739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.844739Z digest=sha256:db9bef42257d89016085e63ac743a9bd3d835cd96a84bdecf8c2663ca240101b

Observation 0a4d095f-b77c-4ee5-a5aa-9c072ebfc696 · outbound

This paper cites A Survey on Large Language Models for Code Generation.

Testing chatbots on the creation of encoders for audio conditioned image generation A Survey on Large Language Models for Code Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.851346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.851346Z digest=sha256:6d995ce7367d038a912d3d718f9e1d38c18d0cb7fb8b54d54a8824b6a7abd458

Observation e2250698-490c-43ca-9507-c52ddad3006b · outbound

This paper cites TimbreCLIP: Connecting Timbre to Text and Images.

Testing chatbots on the creation of encoders for audio conditioned image generation TimbreCLIP: Connecting Timbre to Text and Images

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:25:28.343551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:26.858805Z digest=sha256:9ed11ef90c63029d1ad01e397e41ed1cbfa3bdeec17b5d8ec935e564789d159a

Observation 6c058880-0ef4-4c78-af9a-c283e4103de4 · outbound

This paper cites Noise-Aware Learning from Web-Crawled Image-Text Data for Image Captioning.

Testing chatbots on the creation of encoders for audio conditioned image generation Noise-Aware Learning from Web-Crawled Image-Text Data for Image Captioning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.864893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.864893Z digest=sha256:d4aebc066e5e97076e3f76f9ba93349ad37ed45c2ab510787a5a91e0e40b5700

Observation 69b1f822-cf4d-45df-ac9c-c41495f623ae · outbound

This paper cites Gemini 2.5: Our most intelligent AI model, 2025.

Testing chatbots on the creation of encoders for audio conditioned image generation Gemini 2.5: Our most intelligent AI model, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.871127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.871127Z digest=sha256:4cc234f93ef8bff850049a4a8b116b33024534e46520e6fbf7a40e6f9ee3a3f1

Observation 4d7c3d45-3127-4197-aa5a-f7c0b94c6f3a · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.877668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.877668Z digest=sha256:9a39c23b364771a46b8afff7e093118004c179ce8ef8eb1d2f30f7a68f14c814

Observation d253f024-d21b-4968-b2e5-759b970f9219 · outbound

This paper cites Kingma and Max Welling.

Testing chatbots on the creation of encoders for audio conditioned image generation Kingma and Max Welling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.887383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.887383Z digest=sha256:9f4a3916fe1ab01b51c01f884b5c99e7fad9815df1d033eac9637f8f33b2cd72

Observation e4a5ec78-d9dc-4d66-ba76-3f0446d240da · outbound

This paper cites Benchmarking Cognitive Biases in Large Language Models as Evaluators.

Testing chatbots on the creation of encoders for audio conditioned image generation Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.892235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.892235Z digest=sha256:383b8c3abe538f780fbb24f577e2ea7bbb72a96ae08faeb10d36dc9d9663cde3

Observation 5d376a71-001d-4b19-8ad8-e663da3d7499 · outbound

This paper cites Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?Proceedings of the ACM on Software Engineering, 1(FSE):2261–2284, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?Proceedings of the ACM on Software Engineering, 1(FSE):2261–2284, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.898143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.898143Z digest=sha256:b5f20453c9f62829f12ac32c249f9786a2774d92d3003ef99dbb80f1783bbbd7

Observation b2ed3cc5-127e-42e8-a0e2-220132449bde · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

Testing chatbots on the creation of encoders for audio conditioned image generation AudioGen: Textually Guided Audio Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.909866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.909866Z digest=sha256:1ee7208c88aa89d85ac9ca276b23ccd802962118dff5897081cbfc3ee7b21a73

Observation 435325f9-0ec9-4167-a8c6-b8d4486fe4ad · outbound

This paper cites BindDiffusion: One Diffusion Model to Bind Them All, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation BindDiffusion: One Diffusion Model to Bind Them All, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.916774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.916774Z digest=sha256:d661b695f489ccceeee15f4047be099f277ac834101962625b82ec72086f4a87

Observation 976fec45-ace3-47a1-b262-0ec5f103148f · outbound

This paper cites FLUX, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation FLUX, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.923550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.923550Z digest=sha256:03e462ec0b8e33ef4c5f551799ad6469223fa4ca8ab9ed0ec1893051b1e4ea9d

Observation 7f772987-5f54-46cf-9866-ec6efe4b5788 · outbound

This paper cites Effectively obtaining acoustic, visual and textual data from videos.

Testing chatbots on the creation of encoders for audio conditioned image generation Effectively obtaining acoustic, visual and textual data from videos

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T21:25:28.242608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:26.928879Z digest=sha256:e12b57d9cb155b5804c58ed64ad2024745325baa49a1365cad0e059ce548ed0d

Observation e5a97f0b-c847-4cce-9647-fee7cfe9f742 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Testing chatbots on the creation of encoders for audio conditioned image generation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.940451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.940451Z digest=sha256:6c10f127ec124ca625d43fda946fa6cc3fb358809dd307b443c9deea2ef11475

Observation 41c23dbc-f4c8-4775-a078-abb7d87613f1 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

Testing chatbots on the creation of encoders for audio conditioned image generation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.951332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.951332Z digest=sha256:c429bcb36af3230d694b4ded339a08449e13e356a39922d4718628486683226d

Observation d24a96c5-b7d6-4c71-ab90-f8bde6247612 · outbound

This paper cites Word-Level Explanations for Analyzing Bias in Text-to-Image Models.

Testing chatbots on the creation of encoders for audio conditioned image generation Word-Level Explanations for Analyzing Bias in Text-to-Image Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:25:28.146878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:26.959197Z digest=sha256:f82b2e9d253ee5b62a930883ebbf21ccc7542c74725fb010929aee3bef19ef6f

Observation 1294e260-46d8-432a-bb08-7ca16290258f · outbound

This paper cites Plumbley.

Testing chatbots on the creation of encoders for audio conditioned image generation Plumbley

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.966726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.966726Z digest=sha256:8711a1afa3e5c0907e20ea9fcc0e69190d09ee583505fcda995a9285d7a89ed0

Observation 21019a34-676b-45d6-8db4-d5e6457671e8 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

Testing chatbots on the creation of encoders for audio conditioned image generation Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.973235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.973235Z digest=sha256:27d3139ecb6a38fa7ef8c26be58f0adddc0259a027be98fdec14457f203454e3

Observation 34c387b4-362c-40dd-be69-e07d440041c2 · outbound

This paper cites Michaud, Max Tegmark, and Mike Williams.

Testing chatbots on the creation of encoders for audio conditioned image generation Michaud, Max Tegmark, and Mike Williams

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.980025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.980025Z digest=sha256:fd8286211c533fa140e252b5ea7b29fea6cc62241ff5c992834dd75efacf4b38

Observation 9724b24e-3ecf-4c90-bd92-f7a47ee7a40e · outbound

This paper cites BLAP: Bootstrapping Language-Audio Pre-training for Music Captioning.

Testing chatbots on the creation of encoders for audio conditioned image generation BLAP: Bootstrapping Language-Audio Pre-training for Music Captioning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.985740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.985740Z digest=sha256:fff7d77b976bcb9cc2038bb6e7f929048c179f4e1d1411197fdb3b54b464c413

Observation 2da851cd-2682-42f5-88a5-89d4d333f690 · outbound

This paper cites Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering.

Testing chatbots on the creation of encoders for audio conditioned image generation Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.993264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.993264Z digest=sha256:d7d0609347c8869734b8fc88d73590958d008d38080e754c4726826b32ad869c

Observation 8d17c90f-f472-4576-a9dc-85ca822f9bd4 · outbound

This paper cites Stable Diffusion Akashic Records, 2023.

Testing chatbots on the creation of encoders for audio conditioned image generation Stable Diffusion Akashic Records, 2023

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.003073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.003073Z digest=sha256:8b0ba72db658f0d2161aab25cafb0ba6cdad913a358c5e8dacaf63a02a626bc6

Observation 3a15c995-acde-409f-ae15-874ae5110b08 · outbound

This paper cites Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.IEEE Transactions on Artificial Intelli- gence, pages 1–18, 2025.

Testing chatbots on the creation of encoders for audio conditioned image generation Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.IEEE Transactions on Artificial Intelli- gence, pages 1–18, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.008405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.008405Z digest=sha256:c487772e093534c2764260e1eb590f202a410f6521f5d9005f5a07a1b0f834d5

Observation 9df21f03-dd77-4c91-b5d1-001243792a5a · outbound

This paper cites Mustango: Toward Controllable Text-to-Music Generation.

Testing chatbots on the creation of encoders for audio conditioned image generation Mustango: Toward Controllable Text-to-Music Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.015737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.015737Z digest=sha256:9475f6f5b6f00f6e81205274e456cfce552875c7020e37b21a09864c28553b5b

Observation 4ff0f8b1-b540-4fad-a911-2893145cd957 · outbound

This paper cites Mukhamediev, Adilkhan Symagulov, Yan Kuchin, Kirill Yakunin, and Ma- rina Yelis.

Testing chatbots on the creation of encoders for audio conditioned image generation Mukhamediev, Adilkhan Symagulov, Yan Kuchin, Kirill Yakunin, and Ma- rina Yelis

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.028569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.028569Z digest=sha256:3438cd73ba348dc69e45fee845ec20e2e06976a0d4929a2f74cee24f3b355984

Observation 1e474672-964f-487d-9555-73187522931d · outbound

This paper cites DALL·E 3 System Card, 2023.

Testing chatbots on the creation of encoders for audio conditioned image generation DALL·E 3 System Card, 2023

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.040043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.040043Z digest=sha256:f91192f52ea8c169a637aac71b3fc685d51ce8c8cb670f5376f0c7a1466460dc

Observation 68eebaa9-7df4-4b8f-932d-d2b1201983c5 · outbound

This paper cites Video generation models as world simulators, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Video generation models as world simulators, 2024

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.046611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.046611Z digest=sha256:76b8ea547fd88693104218d57b03371b168428a50b0c7e34ac994e181d1f947d

Observation d34ea8b4-de07-4fa4-9651-94139d0d5480 · outbound

This paper cites OpenAI o3-mini, 2025.

Testing chatbots on the creation of encoders for audio conditioned image generation OpenAI o3-mini, 2025

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:30.124640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.055357Z digest=sha256:5e15d21a45a2e5a529f5e68aaa0023206bd5afc370387811a1267b738bddb6e7

Observation fc5b0e53-a0ec-41a7-bd91-a40affe9a61e · outbound

This paper cites Image-to-Image Translation: Methods and Applications.IEEE Transactions on Multimedia, 24:3859–3881, 2022.

Testing chatbots on the creation of encoders for audio conditioned image generation Image-to-Image Translation: Methods and Applications.IEEE Transactions on Multimedia, 24:3859–3881, 2022

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.944903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.064624Z digest=sha256:e938b433ce4002f30cad46f8a984bb561a7048e03f0d1e53c61276759bccb8c4

Observation d2ad6525-3af7-4fbf-bffa-451c752a362e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Testing chatbots on the creation of encoders for audio conditioned image generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.072098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.072098Z digest=sha256:e6f43b919a115827d96696ad9192f288ae15b17f0066ba381465d6deab1c27de

Observation 9873f815-964c-4990-a342-ec5cc674ded1 · outbound

This paper cites Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.

Testing chatbots on the creation of encoders for audio conditioned image generation Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.802763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.079736Z digest=sha256:85314a7f793ededc76f79252af7a55a338923bc954d9f7f073f63fe9397b8f81

Observation ae62acc3-d4f7-43ab-8426-d82bfc8b4266 · outbound

This paper cites MirrorGAN: Learning Text-To-Image Generation by Redescription.

Testing chatbots on the creation of encoders for audio conditioned image generation MirrorGAN: Learning Text-To-Image Generation by Redescription

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.668998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.087641Z digest=sha256:e9f2cdeb393f3be75003656f15e462a392a4138f8c91cf5cbd5269ac3b5f8497

Observation 94515e86-5902-4e92-a229-dc1711ab73c1 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Testing chatbots on the creation of encoders for audio conditioned image generation Learning Transferable Visual Models From Natural Language Supervision

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.100683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.100683Z digest=sha256:4722e98362a74e6e1ff64bfb398e55b234ecd44670ec08187cf71718ee4fccae

Observation 560ca1e8-28fb-46bf-b316-0d2626334408 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Testing chatbots on the creation of encoders for audio conditioned image generation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.604367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.108332Z digest=sha256:be61660957f51eca6609551ffd48859d270f352e37117e18d1ab5dca191cde3e

Observation e4bfc21b-1c53-4578-98e0-b9f4ab7a1ca6 · outbound

This paper cites Zero-Shot Text-to-Image Generation.

Testing chatbots on the creation of encoders for audio conditioned image generation Zero-Shot Text-to-Image Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.114736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.114736Z digest=sha256:b6e71fb7cba1661d27343797b94efbf3c07ef8a8434c2d17c8f335537aea34a6

Observation 189ec3ff-a078-41c5-bb52-ac6561aa6189 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Testing chatbots on the creation of encoders for audio conditioned image generation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.120753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.120753Z digest=sha256:df0c1ace5d6cce4b1b1f42a80f8d410879c993f30c43a048b890f5ac39a96af3

Observation cd2fb3e2-711a-4f2f-83cc-c9d162c93154 · outbound

This paper cites Stable Diffusion, 2021.

Testing chatbots on the creation of encoders for audio conditioned image generation Stable Diffusion, 2021

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.560658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.128148Z digest=sha256:9c8fc7df688b4098d2159832a54c69680cc6a8febf28bc25678b30e23834d239

Observation b390fbb8-d714-4c5f-b665-0ee1f85481c4 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Testing chatbots on the creation of encoders for audio conditioned image generation High-Resolution Image Synthesis with Latent Diffusion Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.135223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.135223Z digest=sha256:52941dab733b80742d1b86b89452ebc92223d9237d8cc1a45ef9ba6e1a9fc5ec

Observation 876a173b-c180-4813-9a4a-8f98572b99b7 · outbound

This paper cites Stable Diffusion v1-5 Model Card, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Stable Diffusion v1-5 Model Card, 2024

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.532726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.141626Z digest=sha256:aafd909d5aa184885cb330a5044ab9b2d88549590ab29656038f304f88cbd029

Observation 427732d4-70e3-434b-8784-53b9af8ba850 · outbound

This paper cites U-Net: Convolutional Net- works for Biomedical Image Segmentation.

Testing chatbots on the creation of encoders for audio conditioned image generation U-Net: Convolutional Net- works for Biomedical Image Segmentation

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.498665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.146424Z digest=sha256:f4bfb655ad56f24086ba788e134594c78650c7df9aefc500582484ac5657bc5d

Observation f5a8a677-2f4f-4297-aa56-0c07c574c3ff · outbound

This paper cites Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.468796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.151806Z digest=sha256:bbbf13fd51cdefcc6c4fb3232453941ba8d959eaf8c8b2bb110249620112104c

Observation 580452de-ae29-454a-a195-0e68d6aabc2e · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

Testing chatbots on the creation of encoders for audio conditioned image generation Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.447853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.157197Z digest=sha256:881c5eaeb6483b495e9f13ef893de279a75cf09677d27568b923bfff8e2048c3

Observation 448a132b-e5c9-440d-8066-d6aa0d5b81b5 · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

Testing chatbots on the creation of encoders for audio conditioned image generation A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.165116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.165116Z digest=sha256:676b061d9ca0912806291c61d8bd84cd0b0a78d62ec4831330cd3a032585267d

Observation 039102f3-8421-4ccd-b322-a7b9d1b4189a · outbound

This paper cites Comparison and Analysis of Image-to-Image Generative Adversarial Networks: A Survey.

Testing chatbots on the creation of encoders for audio conditioned image generation Comparison and Analysis of Image-to-Image Generative Adversarial Networks: A Survey

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:25:27.892925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.175663Z digest=sha256:2d43a36dc8d85d6713c9cdba8b5f95b8fe192642a38f976518feded8a22dcec1

Observation 4df9dea7-e41e-41d1-82cf-cc5dbb4adae1 · outbound

This paper cites What is noise?Geophysics, 63(4):1122–1124, 1998.

Testing chatbots on the creation of encoders for audio conditioned image generation What is noise?Geophysics, 63(4):1122–1124, 1998

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.427153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.182122Z digest=sha256:c01b515fedd66239b58a3f1ba1b1200636986331e7559815a943f180f3f06bd2

Observation 9e8efcb1-4fd6-484c-a962-df22d26cfcad · outbound

This paper cites Large pre-trained language models contain human- like biases of what is right and wrong to do.Nature Machine Intelligence, 4:258–268, 2022.

Testing chatbots on the creation of encoders for audio conditioned image generation Large pre-trained language models contain human- like biases of what is right and wrong to do.Nature Machine Intelligence, 4:258–268, 2022

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.406716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.188129Z digest=sha256:b991c2e3b1f07f4fbd7df115002150e10b215d7fcc9d2830877573e3c0342212

Observation ca7199ee-0ff4-4c83-836d-9c28b6654a4b · outbound

This paper cites A comprehensive review of large language models: issues and solutions in learning environments.Discover Sustainability, 6, 2025.

Testing chatbots on the creation of encoders for audio conditioned image generation A comprehensive review of large language models: issues and solutions in learning environments.Discover Sustainability, 6, 2025

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.386873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.193122Z digest=sha256:a26c9ea5acf04e14f1cbc6c387aa84fb04664fe3460f0266727583e776182aa4

Observation 09179d60-4806-4adb-a7d8-ec64697aef52 · outbound

This paper cites I Hear Your True Colors: Image Guided Audio Generation.

Testing chatbots on the creation of encoders for audio conditioned image generation I Hear Your True Colors: Image Guided Audio Generation

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:25:27.862179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.198569Z digest=sha256:6ab539368f8c4d2f350246d7acfd78148790ff9bc86f95bf6249713b3bd89388

Observation a0fc0fbc-155f-4321-8f35-c0266598a0db · outbound

This paper cites A Survey on Audio Synthesis and Audio-Visual Multimodal Processing.

Testing chatbots on the creation of encoders for audio conditioned image generation A Survey on Audio Synthesis and Audio-Visual Multimodal Processing

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:25:27.823348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.206894Z digest=sha256:a3bb59200bb99e994322f8d8d74f013a0ef2441285d89e4561a01d44e3486e86

Observation d7b109c7-918a-4162-b654-d7f697ebdd39 · outbound

This paper cites Audio-to-Visual Cross-Modal Generation of Birds.IEEE Access, 11:27719–27729, 2023.

Testing chatbots on the creation of encoders for audio conditioned image generation Audio-to-Visual Cross-Modal Generation of Birds.IEEE Access, 11:27719–27729, 2023

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.353725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.214946Z digest=sha256:bedcdb900a4208b4262e41ba813401c49be8b6074810ebd5e406a2ed7a9580a4

Observation 9c67e0b6-73e8-4529-9a42-bf3c2d3cf6b2 · outbound

This paper cites Outpainting Images and Videos using GANs.International Journal of Computer Trends and Tech- nology, 68(5):24–29, 2020.

Testing chatbots on the creation of encoders for audio conditioned image generation Outpainting Images and Videos using GANs.International Journal of Computer Trends and Tech- nology, 68(5):24–29, 2020

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.323392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.221269Z digest=sha256:b8ea7c623b1c994294be103fbcb7a62519287314a72e427ef428d225db35082e

Observation 8ee53418-5151-4228-92ba-0de62cba72a1 · outbound

This paper cites Pre-trained Speech Processing Models Contain Human-Like Biases that Propagate to Speech Emo- tion Recognition.

Testing chatbots on the creation of encoders for audio conditioned image generation Pre-trained Speech Processing Models Contain Human-Like Biases that Propagate to Speech Emo- tion Recognition

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.294833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.227831Z digest=sha256:4e56ad395b6434967b9f7b38f4887e507e12ed6e7e37cf02cd761f82ab894e40

Observation e529c811-477d-4dc6-8775-99272ed96449 · outbound

This paper cites A survey of multimodal deep generative models.

Testing chatbots on the creation of encoders for audio conditioned image generation A survey of multimodal deep generative models

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.262507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.234737Z digest=sha256:c1fd27929e974d37247519ecc30e1f43df2b549d766bb4534546606d6ffa0b03

Observation e31a462e-99d7-463c-b0e8-11cae17856c8 · outbound

This paper cites CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation.

Testing chatbots on the creation of encoders for audio conditioned image generation CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.241054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.241054Z digest=sha256:89505037de07c8966e93c85f2d166a3aaac701204db23611e1074519b02090e0

Observation 1abe22b2-cfb2-4ecf-b81c-80a6af345805 · outbound

This paper cites Any- to-any generation via composable diffusion.

Testing chatbots on the creation of encoders for audio conditioned image generation Any- to-any generation via composable diffusion

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.242371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.255462Z digest=sha256:1838c9eb5b05bb32534edc782e7032c4d82e9c65636a6053517f49a95809cab6

Observation 13a05cb0-55f7-4dbb-991c-ba76d164d6bf · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Movie Gen: A Cast of Media Foundation Models, 2024

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.222820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.262025Z digest=sha256:0a84a108d47ef499fcd1fd5d2252e4422078a09ff29907816d1a96a7de02fc74

Observation 58e740db-a03e-4fc7-a7ac-05156baf126a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Testing chatbots on the creation of encoders for audio conditioned image generation LLaMA: Open and Efficient Foundation Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.268025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.268025Z digest=sha256:87365924f6543c7c0dc704d4404e993ca43f33961c34dcfd5dd7298c6ce9ae29

Observation f6aec77f-4114-4f64-a806-6f059afab01a · outbound

This paper cites Structural Equation Modeling in Information Systems Research Using Partial Least Squares.Journal of Information Technology Theory and Application, 11(2):5–40, 2010.

Testing chatbots on the creation of encoders for audio conditioned image generation Structural Equation Modeling in Information Systems Research Using Partial Least Squares.Journal of Information Technology Theory and Application, 11(2):5–40, 2010

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.200270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.274583Z digest=sha256:f5f59e5142972623927f44e536ea7672b2d0179d54c998c58d800057ec70f6f8

Observation ef90ac3f-497b-487e-8f25-76419f114f45 · outbound

This paper cites Fugatto 1 - Foundational Genera- tive Audio Transformer Opus 1, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Fugatto 1 - Foundational Genera- tive Audio Transformer Opus 1, 2024

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.174044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.281670Z digest=sha256:ba19b4fae67bc0c6c02115ccbd7ccb85822a1cfce3aedd4a7ff9a9f3deedb5e9

Observation 1310c348-71c0-43bc-81d7-ac37937b96b5 · outbound

This paper cites Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N.

Testing chatbots on the creation of encoders for audio conditioned image generation Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.152219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.288300Z digest=sha256:783f7fbd285c91852c88e15db4471f660140a60ce001d2d55867f7c3d5171df9

Pith citing papers

No inbound Pith citation observations are available.