Pith. sign in

Paper Citation Record · LEDGER

ImageBind: One Embedding Space To Bind Them All

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2305.05665.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.05665 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:12:27.928826Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

13
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 000f3bdd-248f-4333-abc4-be74ef8d2baf · inbound

PandaGPT: One Model To Instruction-Follow Them All cites this paper.

PandaGPT: One Model To Instruction-Follow Them All ImageBind: One Embedding Space To Bind Them All

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:01:37.127118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T09:01:37.096206Z digest=sha256:d07cdda5b985005dfa351ea7b8fbcdceb972692987e6afa124f64b94cc357020

Observation 83c7a18b-0bf5-4d91-af60-dc6190bc2b97 · inbound

Wearable Accelerometer Foundation Models for Health via Knowledge Distillation cites this paper.

Wearable Accelerometer Foundation Models for Health via Knowledge Distillation ImageBind: One Embedding Space To Bind Them All

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:12:27.928826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:12:27.928826Z digest=sha256:cc2f0d92f03f688df8b9c0ca538abf481b61e5c04a7cec07550cde15aa16597f

Observation b37fe7cd-2232-484e-bdc8-a2fada3b1a5d · inbound

Movie2Story: A framework for understanding videos and telling stories in the form of novel text cites this paper.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text ImageBind: One Embedding Space To Bind Them All

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.158316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.158316Z digest=sha256:7970eab9327b093c2e47461ade4e4c4a1dc69acb81d587f3cbbcf9840fb42063

Observation d7ee552c-f1e1-47fe-86a6-1a4a352ef217 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey ImageBind: One Embedding Space To Bind Them All

Reference 136

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.712040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.712040Z digest=sha256:012a96068ef7d30a37ab5b6a122af35a9c0cdb89f93d44e41c535e890491910f

Observation 108a9d9e-1eab-4de0-a8f2-6909738d0dcd · inbound

GPUs, CPUs, and... NICs: Rethinking the Network's Role in Serving Complex AI Pipelines cites this paper.

GPUs, CPUs, and... NICs: Rethinking the Network's Role in Serving Complex AI Pipelines ImageBind: One Embedding Space To Bind Them All

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T16:33:23.996080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:33:23.996080Z digest=sha256:47242f453f458deff6ceaef044726f357c33d4bcd46333d08bb77fd2604c78ed

Observation c08849fd-f62f-447d-9d13-85cc380c45a7 · inbound

Learning to Plan via Supervised Contrastive Learning and Strategic Interpolation: A Chess Case Study cites this paper.

Learning to Plan via Supervised Contrastive Learning and Strategic Interpolation: A Chess Case Study ImageBind: One Embedding Space To Bind Them All

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:37:33.205620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:37:33.205620Z digest=sha256:f5845ed367b63c971aea0273222fc87a41413303cd6cde82e86a09245f69f2f6

Observation 8420d55f-4461-4204-8bee-fa62af796319 · inbound

Learning from Limited and Imperfect Data cites this paper.

Learning from Limited and Imperfect Data ImageBind: One Embedding Space To Bind Them All

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T13:09:57.253731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:09:57.253731Z digest=sha256:68f4e6540315e3134aeeee614eec7c9ba53698c61d879513bc5e707f602ed2ae

Observation cc9105c9-c3f9-462e-85db-995cc7d6b868 · inbound

Effectively obtaining acoustic, visual and textual data from videos cites this paper.

Effectively obtaining acoustic, visual and textual data from videos ImageBind: One Embedding Space To Bind Them All

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.656278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.656278Z digest=sha256:df8faed6b5ddcc064300313cacfb75f1df7374f36e26e0a887ba42b022f995bb

Observation 09e8c540-3a19-48e0-8d72-22511bc0c912 · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation ImageBind: One Embedding Space To Bind Them All

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.782500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.782500Z digest=sha256:32991d0d9198a19dba7bfe00884ecfc61211a1ca6228d32f7a91bc272365db06

Observation c72b6f62-dcd0-4c93-a0f1-c5b58f530078 · inbound

Artificial Phantasia: Emergent Mental Imagery in Large Language Models cites this paper.

Artificial Phantasia: Emergent Mental Imagery in Large Language Models ImageBind: One Embedding Space To Bind Them All

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:44:22.697166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-21T21:41:39.111769Z digest=sha256:adc0bdab51829e29af68bdc0484995f1e8bcd029262d813e9876032a5cbcd95d

Observation 5a106b2e-9977-4464-bba4-fa6e18ed3986 · inbound

Assessing Factual Music Comprehension in Large Audio Language Models cites this paper.

Assessing Factual Music Comprehension in Large Audio Language Models ImageBind: One Embedding Space To Bind Them All

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T00:29:46.287773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:29:46.287773Z digest=sha256:59e77eb92897149f4bd966c9031c265e98bd458f80d62f6368f1b8cdc2bc55a2

Observation 965d9d39-ac50-4bd5-a220-74b627f6fd4e · inbound

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models cites this paper.

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models ImageBind: One Embedding Space To Bind Them All

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:56:33.505525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T19:53:18.200223Z digest=sha256:c129c5c17927c9f93bde76993285ab661c20e66a3ade57ca283a448dd5bc3fd6

Observation a3751b5a-98a1-45f1-a77b-efceaf80bb50 · inbound

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models cites this paper.

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models ImageBind: One Embedding Space To Bind Them All

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:02.077149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T15:53:51.162967Z digest=sha256:34660bd8f2c6d1f4913a2b137fcc0599ca7b3f89b45ca43510da27047bebf53a

Observation 4f017f60-e75c-4d55-b1b2-38c786c9510c · inbound

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models cites this paper.

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models ImageBind: One Embedding Space To Bind Them All

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:28.761245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T07:16:15.202466Z digest=sha256:731f6286699fe282485c394f27e76e2e0cc8fda954f8526b032eef5aceed6507

Observation 174e19d0-a978-49c4-bf2a-f34fa0a26dd1 · inbound

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models cites this paper.

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models ImageBind: One Embedding Space To Bind Them All

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:28.931907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T13:57:47.356373Z digest=sha256:31675ce03823ede7df441f0a30d09008730db3708e5555d9ff4f9bf12707b086

Observation a2e579a7-ca37-40b8-8b01-43f79547c9f5 · inbound

When to Align, When to Predict: A Phase Diagram for Multimodal Learning cites this paper.

When to Align, When to Predict: A Phase Diagram for Multimodal Learning ImageBind: One Embedding Space To Bind Them All

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-27T14:10:57.788301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T14:08:33.542700Z digest=sha256:a3a01cc418777873d466ac8f6d7590307b70e6624673c00cbf6061ed2873cd1d

Observation 8eeb4785-dba7-400a-97bc-01df1b8d7591 · inbound

AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning cites this paper.

AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning ImageBind: One Embedding Space To Bind Them All

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T00:14:56.772592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:14:56.772592Z digest=sha256:1eb7c5fc2f46965dd265dafeb6dc424c6f36ff1d323f3af0e19c337c4238b8b4

Observation ca4e6d03-8add-4756-a62d-18245f5a083d · inbound

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio cites this paper.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio ImageBind: One Embedding Space To Bind Them All

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:35.520484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:35.520484Z digest=sha256:773addf2f39cb730bb615a37066ee308789dcc30048184adedcc700265250cc2

Observation b7e0f27c-b379-42e9-8c91-8930a251271b · inbound

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation cites this paper.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation ImageBind: One Embedding Space To Bind Them All

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-30T10:34:59.307223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T10:34:59.307223Z digest=sha256:2110eb6929620eb585c40a10c65a826e254227cc789e47f032edf35ecfb25ad4

Observation 83822ea9-a08a-4d7b-b085-7fa3729229f7 · inbound

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation cites this paper.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation ImageBind: One Embedding Space To Bind Them All

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:57.995004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:57.995004Z digest=sha256:c11ec7e178f9ebe251c726e0a5c39b471c1c4f01e03d29b980ed1df450b27bb9