Pith. sign in

Paper Citation Record · LEDGER

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio

As of 22 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 2 inbound Pith citation observations for arXiv:2607.18666.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.18666 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T14:46:41.586652Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:50:07.266379Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:13:54.329150Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved71
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57684541-261b-488e-9973-dbc2176ff3ac · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:34.810923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:34.810923Z digest=sha256:e1a68006c13104c45e5727fbb646befa7d44d26f1e6ccaf5ecd332dd8d46ba79

Observation 03c50d2d-29da-42ee-90ed-d6ba7bb606b0 · outbound

This paper cites AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:34.886995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:34.886995Z digest=sha256:dba1ceabb4d56fe9c2dc0d25b78158d7cb83ec706793e47f79adc817d0c7bf57

Observation 242164de-1c8d-46b7-a021-cd9b27f28338 · outbound

This paper cites Cross Modal Retrieval with Querybank Normalisation.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Cross Modal Retrieval with Querybank Normalisation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:34.942567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:34.942567Z digest=sha256:a6ea65de12d12a9cb54757cb5be69eb766a613faa3e0ccde739951858394cd20

Observation bf783d6b-2f12-4a0f-a51e-a82bd4abfd61 · outbound

This paper cites VGGSound: A Large-scale Audio-Visual Dataset.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio VGGSound: A Large-scale Audio-Visual Dataset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:35.018173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:35.018173Z digest=sha256:a946d7f2376026888cd6a0316bca946d54f4d1af38a22869d1056236c77c5fd9

Observation 296b60f5-0f1c-4cfb-9137-57dcd02e2971 · outbound

This paper cites Debiased Contrastive Learning.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Debiased Contrastive Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:35.110010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:35.110010Z digest=sha256:df4cfb7cd53e5d2c03ec6aa84c4559f633668d1d9dde1795354ece09eb353be4

Observation 4202d1d2-ba6c-46eb-92f2-b77b47d1e43a · outbound

This paper cites an unresolved cited work.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:35.176107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:35.176107Z digest=sha256:4d7dacee953d125b5d2bef9e00ff47697e6a2381ca740322d076b9eb848d773c

Observation c46ce0a0-409c-4b8c-bad2-cd897e7b0403 · outbound

This paper cites Scaling up masked audio encoder learning for general audio classification.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Scaling up masked audio encoder learning for general audio classification

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:35.218459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:35.218459Z digest=sha256:dff0a9d05dbe02c123a626e873b0ed88db14e3c4c6c88e1a7467af5ea2f76204

Observation 1bb70593-9ea8-46d8-870c-9f531e446cea · outbound

This paper cites Clotho: An Audio Captioning Dataset.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Clotho: An Audio Captioning Dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:35.273134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:35.273134Z digest=sha256:27d75617c50db14e83dbbfeb3658bc452d64d2baeb122cba5e84d36a505063cd

Observation ef2afe84-bda0-4101-a002-d9f0b655fa38 · outbound

This paper cites an unresolved cited work.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:35.355685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:35.355685Z digest=sha256:b931ee7a1e922a8c7b287a58ccf017c2d208cac2873b329677aa5a85ea8d293f

Observation 7fb4a74b-5154-47f9-9d28-fbcc3f8f59a8 · outbound

This paper cites FSD50K: An Open Dataset of Human-Labeled Sound Events.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio FSD50K: An Open Dataset of Human-Labeled Sound Events

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:35.435587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:35.435587Z digest=sha256:bac8448f000ed6708ce9e91ead517f4d7f09afa9f11e4e67e2aad698d84f4ee8

Observation ca4e6d03-8add-4756-a62d-18245f5a083d · outbound

This paper cites ImageBind: One Embedding Space To Bind Them All.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio ImageBind: One Embedding Space To Bind Them All

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:35.520484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:35.520484Z digest=sha256:7cae52228b8e531c84c572da66ab3d933cb5e602f6411970ccb4d228bf291a75

Observation a3493da3-643c-4812-a2c6-07c005deb742 · outbound

This paper cites Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:35.599874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:35.599874Z digest=sha256:a22bd93077ad0aed289e1f00655af9b9e72e1d65e416c93782ff42f69f000f52

Observation d2a70e26-d842-459e-8545-be39bd557b7f · outbound

This paper cites an unresolved cited work.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:35.680898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:35.680898Z digest=sha256:d1aec6811fdeddf131dec7274ae266c714d8db94da1dd2b8a828f29bcb41f11c

Observation df098999-84fb-4e24-a248-f25fdd37b765 · outbound

This paper cites an unresolved cited work.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:35.924125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:35.924125Z digest=sha256:fc8c7b3bd706b718f8ca91b491106c15766093b7b4380e726dcf96ce65e5b40f

Observation 11bb537f-801d-4aba-a5db-69bbd204d499 · outbound

This paper cites an unresolved cited work.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:36.008176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:36.008176Z digest=sha256:7697b10937f605bfbbc9d830d0adb021c133686eb866f71eabbbf2e5126bc1e9

Observation 5d425ea7-cb0e-463e-9cdc-f29fdaf1eb04 · outbound

This paper cites Matryoshka Representation Learning.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Matryoshka Representation Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:36.068844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:36.068844Z digest=sha256:ee4da45475c17ae096847985fd9d7a540fbcbaa10cfd3780bd7aab1a6aa46322

Observation 69113afc-ea52-4ef1-975b-5bf0ac9db095 · outbound

This paper cites Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:36.225816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:36.225816Z digest=sha256:09cb78013ac6f61e0123eb6139991479bd3ac0de47bfa04923a2cb4b70bf7b32

Observation 97cd49ab-0e2e-4030-b8c0-f60474ab909f · outbound

This paper cites WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:36.312770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:36.312770Z digest=sha256:d80f4e40808b0e23b4a817ad3d330f820315ceeee43fcbdc25aabba6bdc96ffa

Observation 22ea5032-9e87-416a-a7da-f8a923c320ff · outbound

This paper cites VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:36.392432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:36.392432Z digest=sha256:e4f3ce641e69e466728263cfdc8afed5f728f30491ec45f6f08fb012cf266739

Observation 1505e678-0dd3-497d-9131-d24882b32d28 · outbound

This paper cites an unresolved cited work.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:36.527677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:36.527677Z digest=sha256:768f0c07ae52c4a733b7440b17f50a86fad9cd60c50d7e0ff4f9bd404738fe4c

Observation bae42b79-20c9-403c-94f3-ca2738838678 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Learning Transferable Visual Models From Natural Language Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:36.628398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:36.628398Z digest=sha256:915f0c366e26158088786f29de6365c17bee8168244aef6f533a3a70a075e7b4

Observation 5afa41be-7b07-4e50-a3a3-3bc5122dab85 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Robust Speech Recognition via Large-Scale Weak Supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:36.719806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:36.719806Z digest=sha256:a7855ff292c9ec4883669eff3410c4a6abdb4a2941ca2b23157bd11c66877339

Observation c21cdf96-475e-4b8e-99ad-456da980eb74 · outbound

This paper cites Contrastive Learning with Hard Negative Samples.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Contrastive Learning with Hard Negative Samples

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:36.787822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:36.787822Z digest=sha256:943cbf70f4676c14d82a05c7ecbc195f24c025423cb218dcfc9d8d5ed46d3ce2

Observation 1722764a-e47b-4f7a-b2c5-3704b6d0fc5c · outbound

This paper cites Deep CORAL: Correlation Alignment for Deep Domain Adaptation.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Deep CORAL: Correlation Alignment for Deep Domain Adaptation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:36.861796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:36.861796Z digest=sha256:ff2ffbc6de632b2a2073d564a1b686724421b28fd19adcfb815d0d0d00e3e7e2

Observation 7a516517-397a-4f39-9f64-924ee457cd8b · outbound

This paper cites ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.000739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.000739Z digest=sha256:c731ada672afd2d40a8f73cbbfcfb94fca99d6d01007aefeba780b7e686ba033

Observation 16fb1421-1626-4ff0-8819-e61a07bb8cec · outbound

This paper cites FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.087060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.087060Z digest=sha256:504b391fe5f7c14328fec2a761dc9929cff657ce86edaae312af55ca8da98d1b

Observation 5c244156-7a6f-4b8b-a694-fb7c9f93d971 · outbound

This paper cites OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.144047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.144047Z digest=sha256:1080d1c5faf900ee94776fc43ecaf261a129a0579c5d106f1895b8c9e37d3780

Observation 69d92c79-17be-4057-992c-9775f701c4a2 · outbound

This paper cites LLaMA Pro: Progressive LLaMA with Block Expansion.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio LLaMA Pro: Progressive LLaMA with Block Expansion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.217436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.217436Z digest=sha256:cf255fc5b1dd6971c78e0d37c0080aaf0fd2dee06a6ca3ce5a25d65f2da2eda9

Observation 466884e3-e9a4-4cb0-8fe1-c4c759f3ddd1 · outbound

This paper cites Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.267511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.267511Z digest=sha256:9da0d32fd83f3980b1b97e4a4a78812390601cfee69436dc4da0ee70e9616c6c

Observation 6bdb6fd7-8c7f-4bb0-a7d5-6875b20e7d87 · outbound

This paper cites Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.452149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.452149Z digest=sha256:a2b593dea241f388fa3b143750df69355b0a6172aae7da6881d94e4bc2062114

Observation 9f2ca9a1-f9c7-4202-b1aa-55997cda9d8b · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.507620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.507620Z digest=sha256:84a37e438dfc4c6e0fe189f485b91f0deb438e9640551c93619f3fa26f23f241

Observation 1c5044bf-9698-4589-9779-b99ceea62a69 · outbound

This paper cites Cacophony: An Improved Contrastive Audio-Text Model.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Cacophony: An Improved Contrastive Audio-Text Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.580702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.580702Z digest=sha256:c035f0d3c44d94545de24cbf74732de673817e80f9a11f32aae79a7ecca9fb45

Observation fd0d16ba-1e07-4306-9b19-f2682e1fc1be · outbound

This paper cites Proceedings of the 38th International Conference on Machine Learning (ICML) , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Proceedings of the 38th International Conference on Machine Learning (ICML) , year =

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.635148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.635148Z digest=sha256:523722dd138b622220e1d5fa21060d18cb84be70d9c8092473e489beb5c4f1a6

Observation fa801a38-2d2c-4f50-a062-a5a306221f7a · outbound

This paper cites IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , year =

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.681871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.681871Z digest=sha256:e06cb020eb2e76b425a1342119a1b20aabebff69d639eb6200c98310281854e5

Observation 599f4eb7-17db-44f4-9c7b-faad7eaed3ea · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE/ACM Transactions on Audio, Speech, and Language Processing , year =

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.725928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.725928Z digest=sha256:ace0231a933266bda62339ba97460811013c64484b314a24b69d6d19226d9656

Observation e781a5ab-81a0-4df7-b7fb-4e3e8fc0e30b · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE/ACM Transactions on Audio, Speech, and Language Processing , year =

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.801557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.801557Z digest=sha256:5deab4f03e9aa6855cb0c307bfc95998403abd305709fbff476949b368abdc6a

Observation 2dfa1100-0031-4ca3-82c9-ecc7068e01b8 · outbound

This paper cites arXiv preprint arXiv:2503.22104 , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio arXiv preprint arXiv:2503.22104 , year =

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.855705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.855705Z digest=sha256:7e3a3a1730d638195fb0a76172b88dcd1e0336e9623c53aac705275d8332f507

Observation 061103b7-6c54-45fa-bbc7-636c46071a40 · outbound

This paper cites IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , year =

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.894574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.894574Z digest=sha256:552cda64f98ad42cede0a0fe551f18cfc9588df6ded06a8d6cb62b63c68cb50a

Observation 952714c6-b091-4fd4-8dfb-bffe40bc37f9 · outbound

This paper cites Proceedings of Interspeech , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Proceedings of Interspeech , year =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.952707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.952707Z digest=sha256:6b183ea1d3a55f5671606527653093e5db1b58f4119a358b08f71cd282903cd8

Observation cb7a9ca5-6e72-48a3-bfbd-4b9b9141d9ff · outbound

This paper cites IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:37.999121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:37.999121Z digest=sha256:0ff899cb2d6d0d10c3a89754165d4e9406533f90f7fc680a4fc3ede98cac149e

Observation 3a98c961-34bd-4f77-af8f-8f28281cc7d2 · outbound

This paper cites an unresolved cited work.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:38.057883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:38.057883Z digest=sha256:7046349b2437af705084e20e35fcfa3a3848e6eb033dc75902e12855dba4022c

Observation 2415df8d-c687-4351-aed6-2560df305a3d · outbound

This paper cites NeurIPS 2024 Workshop on Audio Imagination , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio NeurIPS 2024 Workshop on Audio Imagination , year =

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:38.105039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:38.105039Z digest=sha256:6fc07cb6336640ecf1f8dc3c675e9719a96d44350832d13399e5a1b161aa0342

Observation e1289a50-85c5-4696-8087-93709f687896 · outbound

This paper cites Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL) , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL) , year =

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:38.159072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:38.159072Z digest=sha256:56db4cef57f43339971ab4baf3a34389f9fc252765743ee96cd451aa489c53c7

Observation d312b185-47de-430b-9a07-43a04289dc71 · outbound

This paper cites arXiv preprint arXiv:2602.18010 , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio arXiv preprint arXiv:2602.18010 , year =

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:38.180937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:38.180937Z digest=sha256:2588b3a9f6d04e4a997eacc6473900707fef39416f5e5c3f5095dad3a7bbb6a2

Observation fbc09cc0-077c-4eb5-8efc-3d8d06d458b3 · outbound

This paper cites IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:38.190713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:38.190713Z digest=sha256:1dbf5463a65088b2521fd246bc85100f968041b16ecfc094c8f7b22cefa34375

Observation 80f6a33a-30a6-4d31-887f-7209a011635d · outbound

This paper cites 2026 , howpublished =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio 2026 , howpublished =

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:38.345481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:38.345481Z digest=sha256:715abb1f129f42aedb1ada274ae9633506a0a6945648eb7045ccd35bba341f14

Observation a5ff43d4-5367-4828-9ccd-f27d986b2e10 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio International Conference on Learning Representations (ICLR) , year =

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:38.581883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:38.581883Z digest=sha256:09f9f093220a1da5864dd404596bd759a398d84b50a4d431db4f6f8b82047e05

Observation 29b6ff79-97ed-4a18-be6c-0d86169c30eb · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:38.748217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:38.748217Z digest=sha256:d05eaf9bdf0d37b42c0cde99772cd8da1227bff9408c484cbdb8a3728b7a57f4

Observation 7259845c-4ffd-4294-8ecc-1215f35304e2 · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:38.841577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:38.841577Z digest=sha256:faa2e9742b289d968c050c9abdd31de54cd91956c870901de8501bd701b35376

Observation 86e71f96-af87-4432-8e3c-dc9f73061587 · outbound

This paper cites 2025 , note =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio 2025 , note =

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:38.958639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:38.958639Z digest=sha256:331e63fe5ca09bb677015482bd74bce8267aed784b1b50553da22a4a4adbb710

Observation d08b9b34-db73-45bd-83e1-18c7d7791da9 · outbound

This paper cites Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:39.076080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:39.076080Z digest=sha256:4f545a62bed4e74745e4da8010016be9be1bd5616b83f84b2849cbf224f543fd

Observation 8c1e66ad-1c0a-4d6b-99a7-3319471a8826 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Qwen2.5-Omni Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:39.196285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:39.196285Z digest=sha256:e31cd1a27c4881b25a43c4ea8a173182eb0d0ffd19f5c5ea4e37f091730ae8f6

Observation bb53797d-d1a0-482d-93c2-7f5783ddd66a · outbound

This paper cites Proceedings of the 40th International Conference on Machine Learning (ICML) , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Proceedings of the 40th International Conference on Machine Learning (ICML) , year =

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:39.387161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:39.387161Z digest=sha256:1081869435ada5734e591ec9e8820888246b45878479a6a1095bd4bda0eb97ec

Observation 7f793620-dc88-486f-bf46-aeba4f56aec4 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:39.547173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:39.547173Z digest=sha256:80c292341452e45cfb0c5784fb21af46f2ff2fd03b2fadc5672d4a5029b2dfaa

Observation 04eb5247-40af-4f8e-af5a-04d69deb0837 · outbound

This paper cites 2024 , note =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio 2024 , note =

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:39.715292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:39.715292Z digest=sha256:79e5e0782366a5e47af0482fcb93f111ac504ad5a1c3a30a20682e144b4da7a0

Observation 52ac5066-cd7d-4ce6-9123-5feb62406751 · outbound

This paper cites jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:39.850618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:39.850618Z digest=sha256:605b54d75ab18bca27a43149abe34eb5acbcd1a41afba6c51cd05641dcb531ca

Observation e9bc642b-c236-4b50-857b-a291c80e89dc · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Representation Learning with Contrastive Predictive Coding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:39.962437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:39.962437Z digest=sha256:99567f1b7bcf2074f18c11b6d26ff4ad0b29871a26c406253b90d12136081f77

Observation ebb58c2a-665a-4d20-a1d9-1552299344fc · outbound

This paper cites European Conference on Computer Vision (ECCV) Workshops , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio European Conference on Computer Vision (ECCV) Workshops , year =

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:40.118060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:40.118060Z digest=sha256:7859d5fd9b7a00b30a4877797f9978567af1fd15147fcccc1642be0f6e0ce19a

Observation fbaca78a-f0da-416d-88ef-0dbfeeaa8cdf · outbound

This paper cites Proceedings of NAACL-HLT , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Proceedings of NAACL-HLT , year =

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:40.258916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:40.258916Z digest=sha256:77aa08d742ba6358baf21c72b6f8051fed09c38adf354a4b0fd7fbfd1b785813

Observation 91b90cb9-a1f3-4862-8d5d-132ec57dfad7 · outbound

This paper cites 2025 , note =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio 2025 , note =

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:40.365360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:40.365360Z digest=sha256:6773fe731b67c5029006f1352e2f5307baf6ff0e06f9beeae5f7411d81aed5a5

Observation f89ec496-6970-4802-9775-eb2cf4c13327 · outbound

This paper cites IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , year =

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:40.527250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:40.527250Z digest=sha256:6d477a780c3348d3b438d655f4b00e4d3fc4cc7eb4edafdb4ac70adeaf426084

Observation c95594cd-9c5d-44f8-92db-ec7b4d5b23aa · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE/ACM Transactions on Audio, Speech, and Language Processing , year =

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:40.664514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:40.664514Z digest=sha256:e8593d3cafd8dff814bff8aeddac2b64d1b6644269fe00e7fc2053d0d8120879

Observation 694240ea-f39b-454b-9f16-0dc740280031 · outbound

This paper cites IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , year =

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:40.786590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:40.786590Z digest=sha256:6cfc13b4a5ce34cf0001dde6e3e15e398b50a20f3e79c44f63c0e81eaf1f93af

Observation 0ef6708a-a80b-4f44-b2d0-7c83bda471da · outbound

This paper cites Proceedings of the 23rd ACM International Conference on Multimedia , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Proceedings of the 23rd ACM International Conference on Multimedia , year =

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:40.895832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:40.895832Z digest=sha256:f33261ebcdeeec9ce41984bf8fed3112f28a9400ecdb58b1dd8dcd0f56574526

Observation 9ebc7435-2115-4e6b-bde4-3403d1f32732 · outbound

This paper cites Proceedings of Interspeech , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Proceedings of Interspeech , year =

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:41.004932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:41.004932Z digest=sha256:847d9afbad1d3e9131379d234db3c41f5526d574d50319ddf9f6437d90fda206

Observation b6359bc2-20b6-43c5-9391-5569c67d2768 · outbound

This paper cites Proceedings of Interspeech , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Proceedings of Interspeech , year =

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:41.150952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:41.150952Z digest=sha256:6da6da5e0d55c8d87aa5031c58910a0258534f1cdb36112ca96b062184fb444e

Observation 65fe257b-aab3-4299-b080-e6cddc37614b · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:41.223342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:41.223342Z digest=sha256:f7539244d10c7085e19661e75c934a3a6d032a30c36e3850e795fde8868f0432

Observation bac82362-830d-4c79-b817-39cecad1d3d9 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio International Conference on Learning Representations (ICLR) , year =

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:41.299158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:41.299158Z digest=sha256:aed2ec48ee63aba7608236ea2412549b1416762889638d4b2ad50f9f2df75050

Observation 871750c1-3fb7-4a76-9732-f124b8e0bf12 · outbound

This paper cites an unresolved cited work.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:41.381838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:41.381838Z digest=sha256:53f3348deec1d186971bca2cfe7cb265665daae3995ffc3f0d6418b4cf7e38f8

Observation e7c46df3-0d70-4dbe-89ee-082ef7a5542c · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio International Conference on Learning Representations (ICLR) , year =

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:41.476775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:41.476775Z digest=sha256:3cf3cb7e104943241ab7b81707e64909524db5c8b226d64f8452aee85497d4f6

Observation 5b437d22-4816-4edf-84e9-98204630fe69 · outbound

This paper cites International Conference on Machine Learning (ICML) , year =.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio International Conference on Machine Learning (ICML) , year =

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:41.586652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:41.586652Z digest=sha256:4ad246d7368c46bc6a3457ffae906859ec7948e9f9be1eadb478521c25909c2d

Pith citing papers

Observation b21b671d-365c-4e1d-b228-490938aed096 · inbound

Discriminative Axis, Not Data Volume: What a Contrastive Corpus Teaches an Audio Embedding cites this paper.

Discriminative Axis, Not Data Volume: What a Contrastive Corpus Teaches an Audio Embedding Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:13:54.384676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:13:53.298988Z digest=sha256:4656f7d753ba9ff924dbd3d4266e5b6a004eb4ba64bc884c73766d678ed245a5

Observation e8ce1989-4146-47c9-8c32-b3bb1f51097e · inbound

Tactus: Open-Vocabulary Object Recognition from Low-Cost Pressure Arrays cites this paper.

Tactus: Open-Vocabulary Object Recognition from Low-Cost Pressure Arrays Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:07.266379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:07.266379Z digest=sha256:79127d3c17222c63a21b7c23cba9f174282a0ee579102cf36c597e758b8d3871