Pith. sign in

Paper Citation Record · LEDGER

ImageBind: One Embedding Space To Bind Them All

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2305.05665.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.05665 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:37:33.205620Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

13
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 000f3bdd-248f-4333-abc4-be74ef8d2baf · inbound

PandaGPT: One Model To Instruction-Follow Them All cites this paper.

PandaGPT: One Model To Instruction-Follow Them All ImageBind: One Embedding Space To Bind Them All

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:01:37.127118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:01:37.096206Z digest=sha256:5b5558576ffae34cd9ffe1c028d434d8ee8abe36c5a1d91bfb7e1624b73b764a

Observation c08849fd-f62f-447d-9d13-85cc380c45a7 · inbound

Learning to Plan via Supervised Contrastive Learning and Strategic Interpolation: A Chess Case Study cites this paper.

Learning to Plan via Supervised Contrastive Learning and Strategic Interpolation: A Chess Case Study ImageBind: One Embedding Space To Bind Them All

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:37:33.205620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:37:33.205620Z digest=sha256:60b605feee0d79b3aa33a300934c552fe40971d22da177a3c752d3797ec09c3c

Observation 8420d55f-4461-4204-8bee-fa62af796319 · inbound

Learning from Limited and Imperfect Data cites this paper.

Learning from Limited and Imperfect Data ImageBind: One Embedding Space To Bind Them All

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T13:09:57.253731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:09:57.253731Z digest=sha256:e73f52e3f15eb51fbf6b76fdc73bd0949d5c484cd0a8660377781b101104e39d

Observation cc9105c9-c3f9-462e-85db-995cc7d6b868 · inbound

Effectively obtaining acoustic, visual and textual data from videos cites this paper.

Effectively obtaining acoustic, visual and textual data from videos ImageBind: One Embedding Space To Bind Them All

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.656278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.656278Z digest=sha256:b281154b3a81aea2fcb1bc432cef11f126b8a424eb25ff3fef4539ec07782128

Observation 09e8c540-3a19-48e0-8d72-22511bc0c912 · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation ImageBind: One Embedding Space To Bind Them All

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.782500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.782500Z digest=sha256:00ce31d7be16cb9816558a6f3ef99cfe175d65c4572ee13474beb997565634e8

Observation c72b6f62-dcd0-4c93-a0f1-c5b58f530078 · inbound

Artificial Phantasia: Emergent Mental Imagery in Large Language Models cites this paper.

Artificial Phantasia: Emergent Mental Imagery in Large Language Models ImageBind: One Embedding Space To Bind Them All

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:44:22.697166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T21:41:39.111769Z digest=sha256:1ab9fbc4e5235206c73facf7fcbb37140e4ed583c9e1ff4d7890b2ac8e3c21f2

Observation 5a106b2e-9977-4464-bba4-fa6e18ed3986 · inbound

Assessing Factual Music Comprehension in Large Audio Language Models cites this paper.

Assessing Factual Music Comprehension in Large Audio Language Models ImageBind: One Embedding Space To Bind Them All

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T00:29:46.287773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:29:46.287773Z digest=sha256:1d2fdd35b307ce73db53dcdf298b741e0e0279b5e4157edfda5a6fdf4a74c794

Observation 965d9d39-ac50-4bd5-a220-74b627f6fd4e · inbound

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models cites this paper.

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models ImageBind: One Embedding Space To Bind Them All

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:56:33.505525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T19:53:18.200223Z digest=sha256:3ee9f4c9f247c4dea66afe12eca002193d64ce448c34aa1d1d5661e01ed17471

Observation a3751b5a-98a1-45f1-a77b-efceaf80bb50 · inbound

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models cites this paper.

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models ImageBind: One Embedding Space To Bind Them All

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:02.077149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:53:51.162967Z digest=sha256:a4e0f42bf00b41de3a1370ff20acca110bb16f91b64a2569a06cb2437d087a85

Observation 4f017f60-e75c-4d55-b1b2-38c786c9510c · inbound

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models cites this paper.

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models ImageBind: One Embedding Space To Bind Them All

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:28.761245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T07:16:15.202466Z digest=sha256:c9009196314845ee765a04f954a84e38a9837581fa3ad4a4baa92f047a00b0b0

Observation 174e19d0-a978-49c4-bf2a-f34fa0a26dd1 · inbound

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models cites this paper.

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models ImageBind: One Embedding Space To Bind Them All

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:28.931907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:57:47.356373Z digest=sha256:bdc5fd9371f3b3ab8e6d11a7db31c20ab2671a1c885a3ef475f5cb7ebc029f7e

Observation a2e579a7-ca37-40b8-8b01-43f79547c9f5 · inbound

When to Align, When to Predict: A Phase Diagram for Multimodal Learning cites this paper.

When to Align, When to Predict: A Phase Diagram for Multimodal Learning ImageBind: One Embedding Space To Bind Them All

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-27T14:10:57.788301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T14:08:33.542700Z digest=sha256:9b22c09ec82e46724f6b74489700f2ee032863949df493c17279aa3b6d61ffcd

Observation 8eeb4785-dba7-400a-97bc-01df1b8d7591 · inbound

AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning cites this paper.

AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning ImageBind: One Embedding Space To Bind Them All

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T00:14:56.772592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:14:56.772592Z digest=sha256:b85e96cc123781b58247d3382fbfbbdd1d76765337db43d2ec26ba7d78738d64

Observation ca4e6d03-8add-4756-a62d-18245f5a083d · inbound

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio cites this paper.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio ImageBind: One Embedding Space To Bind Them All

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:35.520484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:35.520484Z digest=sha256:9a53870c4c0c4ecceef694431aa957501f7e55ca2892471bf26f7e937f6a7a41

Observation b7e0f27c-b379-42e9-8c91-8930a251271b · inbound

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation cites this paper.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation ImageBind: One Embedding Space To Bind Them All

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-30T10:34:59.307223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T10:34:59.307223Z digest=sha256:2e92582766bc212fc65f932dfa354f27fc78ea589fbc8b46bc4fe208e3aaf5dc

Observation 83822ea9-a08a-4d7b-b085-7fa3729229f7 · inbound

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation cites this paper.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation ImageBind: One Embedding Space To Bind Them All

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:57.995004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:57.995004Z digest=sha256:07373951747730f6cd9e97d535151195e3ce4d3e74392aded3662b5c5b376fb1