Pith. sign in

Paper Citation Record · LEDGER

Spider: Any-to-Many Multimodal LLM

As of 13 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 1 inbound Pith citation observation for arXiv:2411.09439.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.09439 v2

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:34:09.157279Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:30:16.892900Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T20:30:17.239571Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact0
  • verified fuzzy64
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc0976e2-2d16-46de-8495-5023b430a0c4 · outbound

This paper cites In OpenAI, 2023.

Spider: Any-to-Many Multimodal LLM In OpenAI, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.920020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.910955Z digest=sha256:f37a2d355b650aece45c6fc524a36fcc472077d8f886957bb5b47e46cd55bc21

Observation a5c16b3d-3e44-4993-810e-d5b224dabe9b · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Spider: Any-to-Many Multimodal LLM Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T20:34:08.915231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:34:08.915231Z digest=sha256:9899e6f5e63615819e982f44180b54436440831c76990454f482803e19e6f9d4

Observation 520e7c5f-8bff-49fd-b0ed-d1902156d88c · outbound

This paper cites Latent-shift: Latent diffusion with temporal shift for efficient text-to-video generation.

Spider: Any-to-Many Multimodal LLM Latent-shift: Latent diffusion with temporal shift for efficient text-to-video generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.902864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.919034Z digest=sha256:3cf19cc626107faae923e1deec985007ef7c0d9d6b3b14431b3f5abcfefc4df6

Observation ab0e6d23-bf05-4c63-a50d-13f985ca4cdf · outbound

This paper cites Blended latent diffusion.

Spider: Any-to-Many Multimodal LLM Blended latent diffusion

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T20:34:08.922655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:34:08.922655Z digest=sha256:f767d9dff5bb143d40f1e1d3c228ec3ff2405caee2cbd9e50e39cdd7226b6737

Observation 712f352d-1d74-4a97-91db-04d942dc3165 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Spider: Any-to-Many Multimodal LLM Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.886471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.926056Z digest=sha256:5fcae19fdf24bea6750772a7b5cb4248435a11a193353c66015b186cb30e5f87

Observation 766ab9fb-82fe-4737-b6b1-610e9841c718 · outbound

This paper cites Lan- guage models are few-shot learners.

Spider: Any-to-Many Multimodal LLM Lan- guage models are few-shot learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T20:34:08.929277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:34:08.929277Z digest=sha256:fcc39c5c24dac941db42d9f9345af1009cac6e5944833da86a89a02a167d65fd

Observation a4cd44ab-966a-4103-bd18-7c4e22a4e498 · outbound

This paper cites Zeroscope: Diffusion-based text-to-video syn- thesis.

Spider: Any-to-Many Multimodal LLM Zeroscope: Diffusion-based text-to-video syn- thesis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.871972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.933476Z digest=sha256:68ecc185a6a23da38b8aa1614f98466f7456633e57c8608ca8ad1ea2045e9790

Observation 77f94e95-45bc-4bbd-99bd-86df5b734659 · outbound

This paper cites an unresolved cited work.

Spider: Any-to-Many Multimodal LLM Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-12T20:34:09.862443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.936631Z digest=sha256:415c5eed61b5e5fddb11b893870a095fe696f792e5905c8e2b1f76d88cd36b59

Observation c0b5f94e-37e5-4f08-a6d4-162ab1eef096 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Spider: Any-to-Many Multimodal LLM Gonzalez, Ion Stoica, and Eric P

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.852114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.939959Z digest=sha256:39408075d0c509fddd7a73b747e9556852fc5f9a989c46438a8a54d53a252f3d

Observation 166ba334-f9fe-48ed-b9f7-f0a1aa2cb932 · outbound

This paper cites Palm: Scaling language modeling with pathways.

Spider: Any-to-Many Multimodal LLM Palm: Scaling language modeling with pathways

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.842316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.943297Z digest=sha256:e099a12c99d58ae866c99d2fa2f8d2ec6832de1a3b844b89efb51b7368ee60af

Observation d933fbab-fd8d-48fb-9252-bc6e1bfbf8ad · outbound

This paper cites Diffedit: Diffusion-based semantic image editing with mask guidance.

Spider: Any-to-Many Multimodal LLM Diffedit: Diffusion-based semantic image editing with mask guidance

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.832263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.947137Z digest=sha256:4a06d9e2bb59fa860dc36290a1e31ef1d61d7404c455f522266bcc753925b765

Observation 1622171f-0a85-46e8-8131-a66a4707f045 · outbound

This paper cites Bert: pre-training of deep bidirectional trans- formers for language understanding.

Spider: Any-to-Many Multimodal LLM Bert: pre-training of deep bidirectional trans- formers for language understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.820907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.950391Z digest=sha256:44a0eee0c3de4129ba264035fb62c155c95ec6aa18b7d0c00a364dfdf5a79df3

Observation f304cf2b-ec49-44e2-85e8-fdecea0f4b0c · outbound

This paper cites Cogview: Mastering text-to- image generation via transformers.

Spider: Any-to-Many Multimodal LLM Cogview: Mastering text-to- image generation via transformers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.810948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.953687Z digest=sha256:31f0bc9caab0b03cf5ab360d306351c97818e8330f99b0d5a154e77dcb2cdd39

Observation dd802106-ee98-464d-bf2c-faea0d87b241 · outbound

This paper cites Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis.

Spider: Any-to-Many Multimodal LLM Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T20:34:08.956846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:34:08.956846Z digest=sha256:1e301092cc4837d55eda1c368ec796c6d0155ba542e8260bacc9327b59d30849

Observation 61cc2a65-b053-4325-b271-ded83992ac79 · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

Spider: Any-to-Many Multimodal LLM Preserve your own correlation: A noise prior for video diffusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.801537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.960414Z digest=sha256:d8653b8caa9cf4a5036d01584898fe2a8a6a321913309da73024d98355121711

Observation 39a849bd-c566-46a8-98b9-f487de559b85 · outbound

This paper cites AudioSet: An ontology and human- labeled dataset for audio events.

Spider: Any-to-Many Multimodal LLM AudioSet: An ontology and human- labeled dataset for audio events

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.791977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.963703Z digest=sha256:09f17f4fb72b944026c7178c4df80f5a0bf439e44110d7222992fb9861dd2ef6

Observation a0437587-179c-4cf3-ac12-9dbd763ccc13 · outbound

This paper cites Imagebind: One embedding space to bind them all.

Spider: Any-to-Many Multimodal LLM Imagebind: One embedding space to bind them all

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.782698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.966767Z digest=sha256:93c7d54b4b2c9cdaab6a728ed1b3036240d71d59dcd02e158d6540ffc57a3e28

Observation 075def10-a591-4d96-9e10-a00ab5279689 · outbound

This paper cites Au- tomated audio captioning by fine-tuning BART with audioset tags.

Spider: Any-to-Many Multimodal LLM Au- tomated audio captioning by fine-tuning BART with audioset tags

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.773589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.969875Z digest=sha256:aa14f2d6fc5519951dbcc12d66915a3ccfe67c6b5ae9b5fc92275d06fbea0848

Observation 238794fc-1f6f-427d-be6b-d283a977b5e6 · outbound

This paper cites Onellm: One framework to align all modalities with language.

Spider: Any-to-Many Multimodal LLM Onellm: One framework to align all modalities with language

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.764515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.973064Z digest=sha256:312a2440fbffe23d916d8640ab4ffed930242e3af419f4706b29909b2fb73eac

Observation 60c8f2f2-e864-4675-bf50-59d2540e185e · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to-video generation via transformers.

Spider: Any-to-Many Multimodal LLM Cogvideo: Large-scale pretraining for text-to-video generation via transformers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.754673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.976175Z digest=sha256:fce6dc34efa9fd43fc412c7a332757cdbc138ea0cf6632a6b6576e1ff10642e0

Observation 1f73eac1-81cf-4018-8b8e-7701e949f239 · outbound

This paper cites Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models.

Spider: Any-to-Many Multimodal LLM Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.745294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.979279Z digest=sha256:ab2391cfbe4e9424f22a1e1895effec8308d38fa584e1e5ae94f4d24b0d880d9

Observation 9ccbb427-5e81-4b33-b843-abd08b6b23fc · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head.

Spider: Any-to-Many Multimodal LLM Audiogpt: Understanding and generating speech, music, sound, and talking head

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.735378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.982334Z digest=sha256:c7e5f968ecc0e8f1689934a221e400a7d4cc0c3d9f4f1cc8103be42544425e35

Observation 4c7ea450-4049-4413-9bd7-ba35c5f49eb9 · outbound

This paper cites Language is not all you need: Aligning perception with language models.

Spider: Any-to-Many Multimodal LLM Language is not all you need: Aligning perception with language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.725295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.986077Z digest=sha256:292d0c15db41e3ab43ec62d3acf657ccf19c1c67b83cda434d92690ad0585c48

Observation 2afe10a7-cb84-4c36-a1d9-04953cbf11fb · outbound

This paper cites Pfb-diff: Progres- sive feature blending diffusion for text-driven image editing.

Spider: Any-to-Many Multimodal LLM Pfb-diff: Progres- sive feature blending diffusion for text-driven image editing

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.714461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.989357Z digest=sha256:86f096da498dff6d3e66402881ff44e528fffe04a7b0b54caff1a3f7531cd0c5

Observation 50b21481-6238-4476-85d3-157deef15429 · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

Spider: Any-to-Many Multimodal LLM Audiocaps: Generating captions for audios in the wild

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.703904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.992419Z digest=sha256:4ceeb626d08634f95b0b42923d5dbb6ecc3a7c6ea7dc942399dcf0dfe6f1c50c

Observation 9abc238d-65b3-4dd3-9e7d-8b9ec81434b2 · outbound

This paper cites Improving audio-language learning with mixgen and multi- level test-time augmentation.

Spider: Any-to-Many Multimodal LLM Improving audio-language learning with mixgen and multi- level test-time augmentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.694162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.995786Z digest=sha256:7ccf30fa7684c05bd62a687fc605253c7da0a1e55be498ca06bec106b208da33

Observation 998a3544-f4e8-44bb-9bda-7d73d6239a6e · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick.

Spider: Any-to-Many Multimodal LLM Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.683521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:08.999251Z digest=sha256:68dfb85f2085f174796c147a3efd98c8a9154bb4cb201b2de8f1731140003769

Observation 68ad3f1b-db77-4935-a552-519e7d387678 · outbound

This paper cites Gen- erating images with multimodal language models.

Spider: Any-to-Many Multimodal LLM Gen- erating images with multimodal language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.673855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.002398Z digest=sha256:dfbcbd16c1514df02604b2d3383e3f56393a54bcd86b6aa95745e3b7e00b00ba

Observation b9d6411b-f207-4397-99d2-f241640dd74e · outbound

This paper cites Are diffusion models vision-and- language reasoners? In Thirty-seventh Conference on Neural Information Processing Systems, 2023.

Spider: Any-to-Many Multimodal LLM Are diffusion models vision-and- language reasoners? In Thirty-seventh Conference on Neural Information Processing Systems, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.663225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.006763Z digest=sha256:0557cbff2b5902899c6d4e1b6c9a30e7ca66e7d83329bc74fceb9f3ee6f36cf6

Observation cee7fa5a-80fa-4d82-b2d2-1908f245d9e0 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Spider: Any-to-Many Multimodal LLM Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T20:34:09.009891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:34:09.009891Z digest=sha256:090fa1ec4d381a5fe3ec61c67a11757a2f448ac8481b08d4964ab13f3f731c51

Observation 412160ca-95f0-4a8d-92c4-93d154f41327 · outbound

This paper cites Videochat: Chat-centric video understanding.

Spider: Any-to-Many Multimodal LLM Videochat: Chat-centric video understanding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.644685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.012982Z digest=sha256:c34fec4522f9a1dd75fd33b78e78d230f0935e8073ad616b43165b33512fd998

Observation a9ca3b2b-d14b-4167-907a-93853157632d · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

Spider: Any-to-Many Multimodal LLM Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.633561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.016310Z digest=sha256:1cfc0268d22ffbafd77b2cd648c31f2ed67103cf2e50e910cc8899c1f6896982

Observation 7ccc856e-e814-4b1a-abc4-7aa6c6f16c62 · outbound

This paper cites VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation.

Spider: Any-to-Many Multimodal LLM VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T20:34:09.019445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:34:09.019445Z digest=sha256:2c89472fe1cf539e92f4153798d0d2b2de23a2ed18a13893e2de079dfdcfb181

Observation 8358131f-0679-4d5d-a903-f4dd95f17692 · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C.

Spider: Any-to-Many Multimodal LLM Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.622125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.022974Z digest=sha256:448f989c3d0a41b8f935d60e64e16f16fd540fb89ece9d43198e753b3d50f346

Observation a313f540-6bc6-4e18-9d1c-b08a200cf3d5 · outbound

This paper cites Mandic, Wenwu Wang, and Mark D.

Spider: Any-to-Many Multimodal LLM Mandic, Wenwu Wang, and Mark D

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.611024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.026277Z digest=sha256:cbc4a1f037984bf9bec4dd146e7b08eb45e37ceb4ad90c4c22741f9dbdc07e34

Observation b33a1fae-9986-4895-9652-26cd808cab83 · outbound

This paper cites Improved baselines with visual instruction tuning.

Spider: Any-to-Many Multimodal LLM Improved baselines with visual instruction tuning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.599872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.029588Z digest=sha256:ee6b4dda120fc6314216a25fdd8ae931d19adf68fa2fdb81b3deba3ab6a02b9f

Observation 5b26ae3c-05fa-40b3-83fd-87c9c6d7faf0 · outbound

This paper cites Visual instruction tuning.

Spider: Any-to-Many Multimodal LLM Visual instruction tuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.589030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.032891Z digest=sha256:1919c749ae57bbea204c738a5ce39a3e624ee2e03a74a192e4c4944b44a024a4

Observation 8b707845-0070-41bd-9cad-17dac61ffe51 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Spider: Any-to-Many Multimodal LLM Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.578355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.036143Z digest=sha256:abe5472be8ead620d9a773d5cf3643987b28da4ca9cf8e13f4146b7ebcc6794b

Observation 5be99fd5-d219-44dd-99f5-1424ad17f4c1 · outbound

This paper cites Sdedit: Guided image synthesis and editing with stochastic differential equa- tions.

Spider: Any-to-Many Multimodal LLM Sdedit: Guided image synthesis and editing with stochastic differential equa- tions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T20:34:09.039426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:34:09.039426Z digest=sha256:552fb6802885f31887d5e61891f91c8c630596b7c188cc34ddf3037b4874a99d

Observation 3c32d38e-bb4c-40f7-b312-0d01ebcd2f94 · outbound

This paper cites Audio-journey: Open domain latent diffusion based text-to-audio genera- tion.

Spider: Any-to-Many Multimodal LLM Audio-journey: Open domain latent diffusion based text-to-audio genera- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.562516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.042401Z digest=sha256:535546cd056428abc0581abcd8829de0db5c219240bcbf5992d8b265c020cd5f

Observation 3c27fd9b-7ca4-4505-8eda-e3ca6c3456a5 · outbound

This paper cites GLIDE: towards photorealis- tic image generation and editing with text-guided diffusion models.

Spider: Any-to-Many Multimodal LLM GLIDE: towards photorealis- tic image generation and editing with text-guided diffusion models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.552335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.045558Z digest=sha256:d0fe072d9ebdc59f4695bfb7f2dcc52693c7062bd8d69c3361b5d3d17bf505d7

Observation 8be672d8-a044-403e-b7aa-e45da14d6a5e · outbound

This paper cites Introducing chatgpt.

Spider: Any-to-Many Multimodal LLM Introducing chatgpt

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.542294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.048694Z digest=sha256:6439f6ba85f4c108959c44a7cf92e4d6ab39dd25b366c59de278ff790104da63

Observation 7ef45537-6ae6-4558-933f-102e5ea72481 · outbound

This paper cites Gross, and Alexander Sorkine- Hornung.

Spider: Any-to-Many Multimodal LLM Gross, and Alexander Sorkine- Hornung

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.532138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.051925Z digest=sha256:2891d1f72d0f76ed67a79d2c5ac518f8daef2aa914230591bd60e90857b6c8ca

Observation bc91d3dd-eaa1-4fa0-ba9b-2b54317a4023 · outbound

This paper cites Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation.

Spider: Any-to-Many Multimodal LLM Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.521061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.055083Z digest=sha256:fbef021d72bd143012739c8215fbfb6097b7a4bb2dbb6fd0b5419af91fd85ef2

Observation ce979068-4b8d-4868-98be-41ca2f008dac · outbound

This paper cites Discriminative probing and tuning for text-to-image generation.

Spider: Any-to-Many Multimodal LLM Discriminative probing and tuning for text-to-image generation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.509792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.058144Z digest=sha256:ac45a51430f309ec54fa2ad65bcd0c5e26d7b92747e70b73acaeab4c509e4ba0

Observation dbde3744-b669-4155-a1b6-15cbe8c4024c · outbound

This paper cites Language models are unsu- pervised multitask learners.

Spider: Any-to-Many Multimodal LLM Language models are unsu- pervised multitask learners

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.499022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.061340Z digest=sha256:0fae3d4669f37e7bb41cbe1d733f5e70f7122a4d0a6939ce6d5d18e1cce12593

Observation 06808218-3fa6-48cd-afa3-2003b2210368 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Spider: Any-to-Many Multimodal LLM High-resolution image syn- thesis with latent diffusion models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.488018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.064435Z digest=sha256:2f987915a2b766e922fa919078823e9a5e6b3069711cfd8cc2e3959b40cb29a4

Observation 72f1d907-2720-4506-a45a-81d1affdd979 · outbound

This paper cites Talking about large language models.

Spider: Any-to-Many Multimodal LLM Talking about large language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.477978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.067505Z digest=sha256:97a3c44e3ca6be43f357c55f16305a4fb8aae36f8431b502a57b9ffb02f2239a

Observation 51acdada-fb45-4107-a4e1-32e4374a59a6 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

Spider: Any-to-Many Multimodal LLM Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.467572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.070632Z digest=sha256:3ef23d34f8812ca322dba240bb29beffd8d4cd88bddc56106f530a8893ca3dea

Observation ddb230e4-9917-410a-909a-64c124ad6dfe · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface.

Spider: Any-to-Many Multimodal LLM Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.457362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.073981Z digest=sha256:95264b71c8c677f2dad4ac9c8b0a1a394a29b6d5c9fa60229e0f77fe9834a33c

Observation 2fc207c6-e93c-4a82-aa49-24603a8472d0 · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

Spider: Any-to-Many Multimodal LLM Make-a-video: Text-to-video generation without text-video data

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.446110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.077665Z digest=sha256:f94aa03c6fa694e2854c5d392789e141a8338dfe8869388ffc449be0a93b5d5d

Observation fd3bdb4c-d298-4238-b49f-be9628a7a14e · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Spider: Any-to-Many Multimodal LLM UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T20:34:09.080817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:34:09.080817Z digest=sha256:a14687f8e9859ff8a6a2ced7f19e5c05a4cbf62ee6400be64e5944ece04d4740

Observation af46836e-fc3f-4fae-8024-d86bddb60e42 · outbound

This paper cites Lan- guage models can see: Plugging visual controls in text gen- eration.

Spider: Any-to-Many Multimodal LLM Lan- guage models can see: Plugging visual controls in text gen- eration

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.435211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.084499Z digest=sha256:9b97f5fa68698888daed6f5a41280475968b51f81e7ae78648d9131aa33a621a

Observation 334befef-3ef3-443c-84fb-7b31632802d1 · outbound

This paper cites Pandagpt: One model to instruction-follow them all.

Spider: Any-to-Many Multimodal LLM Pandagpt: One model to instruction-follow them all

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.423640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.087787Z digest=sha256:874973a5c9a6ce8fa1f1520e1123dbf9c17a087849f0821298a795c143c00674

Observation c8fb613d-3121-4800-b3c3-ee896ba2d032 · outbound

This paper cites Any-to-any generation via composable diffu- sion.

Spider: Any-to-Many Multimodal LLM Any-to-any generation via composable diffu- sion

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.412428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.091154Z digest=sha256:b7ab6207595df0e6d4b6d20ad1aca79a108b782dc1dd282abb0bcb6c5390a863

Observation 30870797-96ab-4e91-a31c-e98c2e46fbae · outbound

This paper cites Galactica: A large language model for science.

Spider: Any-to-Many Multimodal LLM Galactica: A large language model for science

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.401066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.095430Z digest=sha256:bb8ed3773de522566a0b2deabebc8c9ee1bb21f99c2186fe8aa740ed58fd550d

Observation 0551c22d-96cd-4b88-b179-a0f161af4e3d · outbound

This paper cites Gemini: a family of highly capable multimodal models.

Spider: Any-to-Many Multimodal LLM Gemini: a family of highly capable multimodal models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.389913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.098862Z digest=sha256:41c4b509a1c5bfa4e1158accf6a6f02a1b8d87c1195ad76d855dc4ef4bda2f3b

Observation 2ec57542-e34c-44eb-82f7-014367412fc1 · outbound

This paper cites Llama: Open and efficient foundation language models.

Spider: Any-to-Many Multimodal LLM Llama: Open and efficient foundation language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.379603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.102142Z digest=sha256:ad451e9fd5ddd222b86214080f5d748ce98e3d9c98f89785872d432a9c573780

Observation b3f2d3bf-d972-4d4a-8cbb-1ecb806b36d5 · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models.

Spider: Any-to-Many Multimodal LLM Llama 2: Open foundation and fine-tuned chat models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.368870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.105628Z digest=sha256:8c12c6dbc12b2122d4f5c481c41e03b43e8a5cbe307fa2c363346d4d716d34ba

Observation 2a5bec41-d5dd-4191-a617-ce771d08bee8 · outbound

This paper cites Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit.

Spider: Any-to-Many Multimodal LLM Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.358612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.108787Z digest=sha256:4e3d67123cb40bc2c5fa6a3034648800e8e96e84786e94014116c6ad655d80a4

Observation 89d33028-0227-4c98-be1f-56578b5497dd · outbound

This paper cites GIT: A generative image-to-text transformer for vision and language.

Spider: Any-to-Many Multimodal LLM GIT: A generative image-to-text transformer for vision and language

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T20:34:09.112227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:34:09.112227Z digest=sha256:e5a7e03f7d1741265504080092e7fa7c57a051c57668356f5d100136162b6d9f

Observation 461d48b8-71d4-46f7-a40f-15e24bef1b26 · outbound

This paper cites Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework.

Spider: Any-to-Many Multimodal LLM Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.341383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.115367Z digest=sha256:a87811c8bf9a07fa6237a3e248c2e93000e5daf6c4e30e5b1bceaa77529a8e4b

Observation 290ded71-e602-4753-9df8-80228250801b · outbound

This paper cites Campnet: Context-aware mask prediction for end- to-end text-based speech editing.

Spider: Any-to-Many Multimodal LLM Campnet: Context-aware mask prediction for end- to-end text-based speech editing

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.331254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.118588Z digest=sha256:51901b4a573ca6e2af0c7d4fd8199d6da7cd7414ee6f762bf3ff7425eddfcd6d

Observation 820c89e5-9c00-4678-85d9-15ba3204d9d7 · outbound

This paper cites Visual chatgpt: Talking, drawing and editing with visual foundation models.

Spider: Any-to-Many Multimodal LLM Visual chatgpt: Talking, drawing and editing with visual foundation models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.321608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.121667Z digest=sha256:3e218ea446dd6110f56f02ced0dd228075429db295d599ede6319d507f615c00

Observation f3482a93-9430-4576-9cf7-3e51f9ea277c · outbound

This paper cites Tune-a-video: One-shot tuning of im- age diffusion models for text-to-video generation.

Spider: Any-to-Many Multimodal LLM Tune-a-video: One-shot tuning of im- age diffusion models for text-to-video generation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.311833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.124788Z digest=sha256:928b4d6351355377e7a32ae4bfd595cb39ef082c38189c3762720dd676b40aa0

Observation af9eb693-1002-480f-9eb7-50fa7046290a · outbound

This paper cites Next-gpt: Any-to-any multimodal llm.

Spider: Any-to-Many Multimodal LLM Next-gpt: Any-to-any multimodal llm

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T20:34:09.128314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:34:09.128314Z digest=sha256:2b938a6e094275cedcb7b65a6f15cc6392b36062b3efa2b44c29bfef39740697

Observation 0003e81f-25dd-4ee5-aa48-79f9cea3b850 · outbound

This paper cites mplug-2: A modularized multi-modal founda- tion model across text, image and video.

Spider: Any-to-Many Multimodal LLM mplug-2: A modularized multi-modal founda- tion model across text, image and video

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.295674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.131632Z digest=sha256:fa9633eb9d99a2910cdda4284288e927a25eb002341515222f51d2dc2442b600

Observation 17f10101-2ba0-4cb5-abe9-0257bed5a4af · outbound

This paper cites MSR-VTT: A large video description dataset for bridging video and lan- guage.

Spider: Any-to-Many Multimodal LLM MSR-VTT: A large video description dataset for bridging video and lan- guage

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.286399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.134758Z digest=sha256:1ab40f67a8b849042a0b3ad00b1b1ef039764fa997bfc8dd65647cacf82b1f77

Observation 6f79a270-668f-4181-b6bc-04a9d9639e7c · outbound

This paper cites Diffsound: Discrete dif- fusion model for text-to-sound generation.IEEE ACM Trans.

Spider: Any-to-Many Multimodal LLM Diffsound: Discrete dif- fusion model for text-to-sound generation.IEEE ACM Trans

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.276314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.138050Z digest=sha256:7e15d07e42228cd329c60276338923335b1498a2d97ff84d0b6e5d42104f19e8

Observation a3796435-b038-41bd-883e-6a210aba67ed · outbound

This paper cites Speechgpt: Empow- ering large language models with intrinsic cross-modal con- versational abilities.

Spider: Any-to-Many Multimodal LLM Speechgpt: Empow- ering large language models with intrinsic cross-modal con- versational abilities

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.265980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.141144Z digest=sha256:98ad5ad61223fad2f6d6542d9ed1d81f481ffab005ab3c2a3bc29fc41fd88f3a

Observation fa1d3c69-b0cc-43c0-97c4-cdd92099050f · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding.

Spider: Any-to-Many Multimodal LLM Video-llama: An instruction-tuned audio-visual language model for video un- derstanding

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.254659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.144231Z digest=sha256:9e622aaf24ae616f159ea75a50409a768d9ef6cd3a255b41e8a123aa6533f19d

Observation cc7c4f99-a4ff-46d0-bcbe-762f9bc2bc39 · outbound

This paper cites Object relational graph with teacher-recommended learning for video captioning.

Spider: Any-to-Many Multimodal LLM Object relational graph with teacher-recommended learning for video captioning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.244681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.147389Z digest=sha256:e8eee0da473a795d177cdf05736dca08f73594dbd584a97741a59acc42e61b45

Observation e40a7b48-6890-407d-a034-9eb22b966aea · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

Spider: Any-to-Many Multimodal LLM Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.234960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.150293Z digest=sha256:051fdcf527c11d948dcdfe24c2aa7809d7108833cccdad3bc2ab498b20d49607

Observation bac3282e-acb2-486b-bdfc-b61e3a3b3a74 · outbound

This paper cites Thus, the modalities generation performances of our Spider are limited by the integrated Decoder models.

Spider: Any-to-Many Multimodal LLM Thus, the modalities generation performances of our Spider are limited by the integrated Decoder models

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.224721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.153845Z digest=sha256:d387188134e02edb8fdeee1a7c59fdb6fc21d06796f33e5c3bfd3b5bbede0099

Observation 49b10b23-3b5c-4c6f-920b-b5f861ed235f · outbound

This paper cites Image-to-Text generation on COCO-caption [34].

Spider: Any-to-Many Multimodal LLM Image-to-Text generation on COCO-caption [34]

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:34:09.214488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:34:09.157279Z digest=sha256:46f91bb7dbb65350ac817f1a68259cf77a24d69d4d63999e6205e4c45bcef287

Pith citing papers

Observation 0fbc1ba3-14af-48e3-b4d2-2f4459be40ff · inbound

A Unified Multi-Agent Framework for Universal Multimodal Understanding and Generation cites this paper.

A Unified Multi-Agent Framework for Universal Multimodal Understanding and Generation Spider: Any-to-Many Multimodal LLM

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T20:30:17.344362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T20:30:16.892900Z digest=sha256:f5d413da4d977c1f14d155f653e39a702e79e727d0f3cfa1483787406f44a242