Pith. sign in

Paper Citation Record · LEDGER

Open-set Cross Modal Generalization via Multimodal Unified Representation

As of 7 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2507.14935.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14935 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:48:33.963779Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:53:28.021948Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T21:53:34.084878Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact5
  • verified fuzzy45
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9a653eda-7ed8-461e-9c05-254125e640ed · outbound

This paper cites Robust cross-modal representation learning with progressive self- distillation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Robust cross-modal representation learning with progressive self- distillation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.945276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:25.548996Z digest=sha256:72f921ab7e8945c8d0eef0a681a766c61328e4c33c15a1472d54437a1a03ac5e

Observation da722cd7-14fe-4fba-9270-de4615f6ea92 · outbound

This paper cites On the effectiveness of image rotation for open set domain adaptation.

Open-set Cross Modal Generalization via Multimodal Unified Representation On the effectiveness of image rotation for open set domain adaptation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.699431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:25.670170Z digest=sha256:0bf265377b2642e6b80f33aa7a07876074bac1c7b779bf1f6b6ade36a09f904d

Observation f27956e9-7a49-4e87-ac8a-f86e3d175bdf · outbound

This paper cites Domain generalization by solving jigsaw puzzles.

Open-set Cross Modal Generalization via Multimodal Unified Representation Domain generalization by solving jigsaw puzzles

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.448429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:25.778914Z digest=sha256:dc37d0b7120688ba4fbe93953527d5ec66fafdd626598239520a49ff77175062

Observation dee88b25-f9c6-4d64-b7b1-a9ce145e3d7c · outbound

This paper cites Collecting highly paral- lel data for paraphrase evaluation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Collecting highly paral- lel data for paraphrase evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:25.886544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:25.886544Z digest=sha256:d9c6a556560bc04c991301b9aec288ee38932d8c0b684a5e1c83a4848a146e8e

Observation 396e61a4-d780-417d-a576-1fbeb402b0c1 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Open-set Cross Modal Generalization via Multimodal Unified Representation Vggsound: A large-scale audio-visual dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.213437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:26.060304Z digest=sha256:20dbaaa8103afabc9c85161b23fe8f08a282eba1ee955325e9135e099e6f309e

Observation 40a90f4a-be71-4554-a7b7-24b51ec6dc28 · outbound

This paper cites Hts-at: A hierarchi- cal token-semantic audio transformer for sound classifica- tion and detection.

Open-set Cross Modal Generalization via Multimodal Unified Representation Hts-at: A hierarchi- cal token-semantic audio transformer for sound classifica- tion and detection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.924381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:26.165853Z digest=sha256:7960f68ae88d0e43b1923f3c1772d0d699cd69c8e673841eb2670dccf42e3dfd

Observation 9eacc5aa-1288-437f-a5f3-2e665af6c317 · outbound

This paper cites VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset.

Open-set Cross Modal Generalization via Multimodal Unified Representation VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:26.258804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:26.258804Z digest=sha256:2bab738fe6c1d0a7ce06b72167bb4b5dc8fffc5721cf60cf89f33277347be9e9

Observation 4528c55e-2a93-4781-ae74-7c72c085ea0e · outbound

This paper cites Uniter: Universal image-text representation learning.

Open-set Cross Modal Generalization via Multimodal Unified Representation Uniter: Universal image-text representation learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.722400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:26.393490Z digest=sha256:a8af02f27bef60807f8338e4b573d033219f483bcd2de37668d603b9e4a31544

Observation 7406c049-2539-466b-aa97-8122938036f1 · outbound

This paper cites Sinkd: Sinkhorn distance minimization for knowledge distillation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Sinkd: Sinkhorn distance minimization for knowledge distillation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.497935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:26.524128Z digest=sha256:602d6feca68dd77aeca566e8ea489972fe633dd847247d158a7897d33a13a9fe

Observation d694a12a-1115-452f-949a-aca3c1e76ad3 · outbound

This paper cites Sinkhorn distance minimization for knowledge distilla- tion.

Open-set Cross Modal Generalization via Multimodal Unified Representation Sinkhorn distance minimization for knowledge distilla- tion

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.205864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:26.655853Z digest=sha256:681f0d576bcf014e5f7446fe531e88f7833b292728f92eeb3f52633b4b8f9972

Observation 1b229b04-2c2d-4e19-a83e-327b027cdebf · outbound

This paper cites Optical: Leveraging optimal trans- port for contribution allocation in dataset distillation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Optical: Leveraging optimal trans- port for contribution allocation in dataset distillation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.922516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:26.771225Z digest=sha256:0d912c5693178f5833906f5da4fa664489f9a6daf6301b0a7c11c698b0e792fd

Observation 2f994dd1-b7c2-4e6d-8938-a2e94fdcde70 · outbound

This paper cites Layoutenc: Leveraging enhanced layout rep- resentations for transformer-based complex scene synthesis.

Open-set Cross Modal Generalization via Multimodal Unified Representation Layoutenc: Leveraging enhanced layout rep- resentations for transformer-based complex scene synthesis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.599120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:26.885590Z digest=sha256:ce1622cb213c5c41b45ad5b2c88c5137edc5ebfa97803dd959d58d7db3b06516

Observation a3cf246e-ee9e-41e9-9ee8-07a0e685898f · outbound

This paper cites Streetsurfgs: Scalable ur- ban street surface reconstruction with planar-based gaussian splatting.

Open-set Cross Modal Generalization via Multimodal Unified Representation Streetsurfgs: Scalable ur- ban street surface reconstruction with planar-based gaussian splatting

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.384755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:27.037175Z digest=sha256:4c02a93ee80304e0e3c44b6e95384d022d8cde3dce0dab226d158524e0d805f8

Observation af806a17-5526-4171-8036-6b1e99c5917e · outbound

This paper cites Towards Multimodal Open-Set Domain Generalization and Adaptation through Self-supervision.

Open-set Cross Modal Generalization via Multimodal Unified Representation Towards Multimodal Open-Set Domain Generalization and Adaptation through Self-supervision

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:35.385262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:27.140292Z digest=sha256:e0c78d67ec067f7692c3667aa1a336e15d153a4b5f40c98cd2a9192d89a7d74b

Observation 33e87a74-1eb6-464c-b70a-86a9749d9f10 · outbound

This paper cites Simmmdg: A simple and effective framework for multi-modal domain generalization.

Open-set Cross Modal Generalization via Multimodal Unified Representation Simmmdg: A simple and effective framework for multi-modal domain generalization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.159885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:27.241584Z digest=sha256:a1f472067416ed6ee8bb4cfa0afcbab68bc8d6007f38a6991c11bc51030b1ade

Observation 507023a4-0327-44b5-b8ea-a620ee805697 · outbound

This paper cites Clotho: An audio captioning dataset.

Open-set Cross Modal Generalization via Multimodal Unified Representation Clotho: An audio captioning dataset

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.936791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:27.379391Z digest=sha256:8df95b33fcf28190fa2be80b2f83454ca628c82c6dee0784d940ac996f6213eb

Observation 09215f9d-9f71-494e-b01f-299865375760 · outbound

This paper cites Multi-modal align- ment using representation codebook.

Open-set Cross Modal Generalization via Multimodal Unified Representation Multi-modal align- ment using representation codebook

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.707248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:27.511828Z digest=sha256:6126f93df7c6b60fe8ba8efa4b3cef06cec75ece9a6c213c705c393f7b4cf514

Observation 51531b0d-6a4b-410d-be22-2e8426a6893c · outbound

This paper cites Ace: A generative cross-modal retrieval framework with coarse-to-fine semantic modeling.

Open-set Cross Modal Generalization via Multimodal Unified Representation Ace: A generative cross-modal retrieval framework with coarse-to-fine semantic modeling

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-06T15:48:35.160122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:27.648562Z digest=sha256:40e38508064b3417772b12f8a45ed3aea8d2b635711ec63b11740d76449631a2

Observation 2f814264-7211-4be5-a0d5-027e32162906 · outbound

This paper cites Slowfast networks for video recognition.

Open-set Cross Modal Generalization via Multimodal Unified Representation Slowfast networks for video recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:27.779563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:27.779563Z digest=sha256:907016c08f31e66fd9e0315abe4829420dcb4c6525cb432bbc11d04d11b6b309

Observation fd407d55-5386-4d64-9c20-ccaf26e2e10f · outbound

This paper cites Domain-adversarial training of neural networks.

Open-set Cross Modal Generalization via Multimodal Unified Representation Domain-adversarial training of neural networks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.460597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:27.947964Z digest=sha256:6b2e46780b4f5ad6b36ee2ff5a5b65934f5d6796cca4e0daf6a03f53da293020

Observation 4c64e29f-8989-4401-991e-dc5f1e2c3740 · outbound

This paper cites Imagebind: One embedding space to bind them all.

Open-set Cross Modal Generalization via Multimodal Unified Representation Imagebind: One embedding space to bind them all

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:28.115661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:28.115661Z digest=sha256:19bf2d292f58229384969f6e12dfd08baa9e02177d37f44ab0f9ad15d1bab1ee

Observation 5693eb52-ffb9-4a0e-a0b6-c3fd604f14b2 · outbound

This paper cites Enhancing Multimodal Unified Representations for Cross Modal Generalization.

Open-set Cross Modal Generalization via Multimodal Unified Representation Enhancing Multimodal Unified Representations for Cross Modal Generalization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:28.218101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:28.218101Z digest=sha256:366f82fbba2b8a5c8b97c2f8d9797c1ae9fec3471e3d12ce535c802f29e9eb18

Observation 302450ac-17aa-42a6-8c26-e1ed2f82c0bd · outbound

This paper cites Semantic residual for multimodal unified discrete representation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Semantic residual for multimodal unified discrete representation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.239672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:28.320120Z digest=sha256:ece5d963eb5d3e22bf898a95c1b49cba67cca8eaa1b4357d59dc0b5f2e8b2137

Observation bdc993b6-96dd-4fdc-b61a-fd7b27d75e34 · outbound

This paper cites Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations.

Open-set Cross Modal Generalization via Multimodal Unified Representation Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:28.447490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:28.447490Z digest=sha256:c106ce8d371e2ca5eb1fcf2abe792263c5d6535ced992edb5131de0c3534ea99

Observation b0f93d54-2536-44a3-91c0-881c016cfe4f · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

Open-set Cross Modal Generalization via Multimodal Unified Representation WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:28.600859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:28.600859Z digest=sha256:106e1fa0d845d8e1562c08ec1b97519a79e567ca5ec854b530fddd7b509571c7

Observation 807fa901-b252-4444-8f09-06bbc1d8feb9 · outbound

This paper cites Learning to generalize: Meta-learning for do- main generalization.

Open-set Cross Modal Generalization via Multimodal Unified Representation Learning to generalize: Meta-learning for do- main generalization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.994612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:28.728539Z digest=sha256:8949bede5e6f7e12e63f912f35278aaa7021f84bc2bd8c9585c70a581586432d

Observation e4e07c83-b58c-4ac9-9336-480b1ebb98c7 · outbound

This paper cites Domain generalization for med- ical imaging classification with linear-dependency regular- ization.

Open-set Cross Modal Generalization via Multimodal Unified Representation Domain generalization for med- ical imaging classification with linear-dependency regular- ization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.783890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:28.883688Z digest=sha256:1bc347a7bc3b0d96147498d44bcbe5294a036b78727db44b97780d0f99ae8f42

Observation 882ef81a-0a97-43bc-8971-19b5b257d701 · outbound

This paper cites Adjustment and alignment for unbiased open set domain adaptation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Adjustment and alignment for unbiased open set domain adaptation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.499297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:29.019728Z digest=sha256:772f188b866ce601aeb84bc191893b87fe614bf084a333d49e3db2d07524e1a0

Observation b9e429f1-b3ba-46b4-82a2-594c73c93eff · outbound

This paper cites Cross-Modal Discrete Representation Learning.

Open-set Cross Modal Generalization via Multimodal Unified Representation Cross-Modal Discrete Representation Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:29.139488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:29.139488Z digest=sha256:78c6bd09f75cf0f1aa74742afafe9b31f158912f710dbeff24856a9916492e3f

Observation 1c7a639d-2e41-4ca0-b4a7-ed46348417e7 · outbound

This paper cites Feddg: Federated domain generalization on medical image segmentation via episodic learning in continuous fre- quency space.

Open-set Cross Modal Generalization via Multimodal Unified Representation Feddg: Federated domain generalization on medical image segmentation via episodic learning in continuous fre- quency space

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.266176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:29.274536Z digest=sha256:7fa75dce9f2f957d46b8ef30ca7442a9b1e29bdb08fefda0dd1b762d8f44d36e

Observation a53e4ee0-723a-424f-9877-952fddd56894 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Open-set Cross Modal Generalization via Multimodal Unified Representation Swin transformer: Hierarchical vision transformer using shifted windows

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:29.435332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:29.435332Z digest=sha256:e52a5a01558ed5e98cd89a70a85993239c414d858ff018a59d45376dc09caef6

Observation 3d48944d-23b5-4973-ace7-88db099cecdf · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

Open-set Cross Modal Generalization via Multimodal Unified Representation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:29.595479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:29.595479Z digest=sha256:ae4cf282e75d911e3c86c8a1bebf632559384dc67907b6e3cc66b38f1bf3157c

Observation 604260d1-9db9-4e96-9617-9735f183c9ad · outbound

This paper cites Unsupervised learning of visual representations by solving jigsaw puzzles.

Open-set Cross Modal Generalization via Multimodal Unified Representation Unsupervised learning of visual representations by solving jigsaw puzzles

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.956515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:29.699219Z digest=sha256:8858b572d8eae7273065d558031fd1ba0e6626839b292461004e7e86dd2a69a8

Observation a19c92f6-8dc0-4a89-b56e-f1530b4f90c6 · outbound

This paper cites Two at once: Enhancing learning and generalization capacities via ibn-net.

Open-set Cross Modal Generalization via Multimodal Unified Representation Two at once: Enhancing learning and generalization capacities via ibn-net

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.672767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:29.830914Z digest=sha256:e0827acd8f6bc9211467148f662f160ef3014a77e3ef92eceb4e7b87dec65e42

Observation 0a5add89-0bdf-4104-b171-e11c1a236067 · outbound

This paper cites Estimating Visual Information From Audio Through Manifold Learning.

Open-set Cross Modal Generalization via Multimodal Unified Representation Estimating Visual Information From Audio Through Manifold Learning

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:34.769435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:29.953749Z digest=sha256:b79be9992ecc926b05a4b47366996e29cba8690426f291f6a6bcf4d0167267e0

Observation b771fb0b-8b4f-4812-8c64-7d035d63b104 · outbound

This paper cites Audio-visual speech recognition with a hybrid ctc/attention architecture.

Open-set Cross Modal Generalization via Multimodal Unified Representation Audio-visual speech recognition with a hybrid ctc/attention architecture

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.464962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:30.116086Z digest=sha256:2d7bdd9ad0f755618807f5dc9199a294614f58c3688286db81362bbd1d478226

Observation 0a61c64f-76dc-42b1-930c-c750e10ea19f · outbound

This paper cites Domain generalization through audio- visual relative norm alignment in first person action recog- nition.

Open-set Cross Modal Generalization via Multimodal Unified Representation Domain generalization through audio- visual relative norm alignment in first person action recog- nition

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.216157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:30.254045Z digest=sha256:68563d94f3f0fc97c65d3ee6504eec7cf4b315951633d821eaa98f4815e9e8bf

Observation eb751c58-b667-4e10-8fff-866a3022576c · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Open-set Cross Modal Generalization via Multimodal Unified Representation Learning transferable visual models from natural language supervi- sion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:30.385147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:30.385147Z digest=sha256:8b1346af9982ee8e1ef4f778a39f34a4aba49145062d2656c3022bde90cd780e

Observation 0f742929-52d9-43e4-a3c9-5c0dc224b5e2 · outbound

This paper cites Mask2anomaly: Mask transformer for uni- versal open-set segmentation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Mask2anomaly: Mask transformer for uni- versal open-set segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.949706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:30.518039Z digest=sha256:8445c79b799c72438b81decb7a69c7a4104e59f0e5c0f7b0f4e451c82718d2b2

Observation 62bdf01a-5e88-4d31-b262-8bf4980429fa · outbound

This paper cites XKD: Cross-modal Knowledge Distillation with Domain Alignment for Video Representation Learning.

Open-set Cross Modal Generalization via Multimodal Unified Representation XKD: Cross-modal Knowledge Distillation with Domain Alignment for Video Representation Learning

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:34.487611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:30.648250Z digest=sha256:3be6456660c15b2dd52954bfe382af2f91d4630a64ea50ead21f82bd7840b0b1

Observation 24c727c0-79b1-4434-95d5-26697d9abdf6 · outbound

This paper cites Open domain generalization with domain- augmented meta-learning.

Open-set Cross Modal Generalization via Multimodal Unified Representation Open domain generalization with domain- augmented meta-learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.701072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:30.804084Z digest=sha256:13f699daca8e19b443bad7df218977ced06e573e874030a9d861982542494dd1

Observation 7281a904-b633-496b-bf6b-43836baae6e1 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Open-set Cross Modal Generalization via Multimodal Unified Representation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:30.917220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:30.917220Z digest=sha256:0e8bb2ccf10ed19e1611d67250be1599f4508cd18e7cd9544755a2e4547173b5

Observation ef32786e-a737-49b6-8428-03a2dcfd6dff · outbound

This paper cites Audio-visual event localization in unconstrained videos.

Open-set Cross Modal Generalization via Multimodal Unified Representation Audio-visual event localization in unconstrained videos

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.399365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:31.035641Z digest=sha256:23cdad8da853e4ad258780ced62c435cf03d06bc08c746aaaa0a4d167fda1e8a

Observation 9651d3e8-bc5d-4b0a-bc27-b1d713ae757f · outbound

This paper cites Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing.

Open-set Cross Modal Generalization via Multimodal Unified Representation Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.163541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:31.148002Z digest=sha256:ea41f43456d4a3c5f2cd11ccf1e366d8b3b8ad9add8723a4e3830bc2b0caa3dc

Observation 5de95d9f-85c0-426e-9575-b89836e6821f · outbound

This paper cites Domain randomization for transferring deep neural networks from simulation to the real world.

Open-set Cross Modal Generalization via Multimodal Unified Representation Domain randomization for transferring deep neural networks from simulation to the real world

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.944179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:31.310142Z digest=sha256:44a0e209dfe7b1c4ae000b1487eda243841e1a49ed855c0c2d5848d11c2bee3d

Observation ec0d5cd5-d07d-4e68-a099-f64740d5f903 · outbound

This paper cites Deep Domain Confusion: Maximizing for Domain Invariance.

Open-set Cross Modal Generalization via Multimodal Unified Representation Deep Domain Confusion: Maximizing for Domain Invariance

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.427887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.427887Z digest=sha256:2c3bed6d87aeefb61350f7100115c1c2d522434e54f1845255965672aa231fa4

Observation 4f23e8e0-ffe8-4fd5-b1ec-ea7f4cc4fc3e · outbound

This paper cites IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models.

Open-set Cross Modal Generalization via Multimodal Unified Representation IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.574133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.574133Z digest=sha256:45c346f11cd13054957d189a67f5f924b439ac7284e4836bcd546d827bb87552

Observation 02e59edd-c139-4d7d-a293-980bf430387a · outbound

This paper cites Towards Transformer-Based Aligned Generation with Self-Coherence Guidance.

Open-set Cross Modal Generalization via Multimodal Unified Representation Towards Transformer-Based Aligned Generation with Self-Coherence Guidance

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:34.215630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:31.694527Z digest=sha256:f232d64703034120d6555452e79286842b028f422a4cf5bd80a1bdc1c1d861ef

Observation 449e3041-9ccf-43fd-b742-3b06a1d220cb · outbound

This paper cites Vlmixer: Unpaired vision-language pre-training via cross-modal cutmix.

Open-set Cross Modal Generalization via Multimodal Unified Representation Vlmixer: Unpaired vision-language pre-training via cross-modal cutmix

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.699412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:31.802571Z digest=sha256:1049be7f3f762b030537b66f2b27a5a49a818744a2eaaccc90dd7fccebe28cd9

Observation 4d17012a-086b-478e-98a8-7edd453a268c · outbound

This paper cites General- izable decision boundaries: Dualistic meta-learning for open set domain generalization.

Open-set Cross Modal Generalization via Multimodal Unified Representation General- izable decision boundaries: Dualistic meta-learning for open set domain generalization

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.416582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:31.971663Z digest=sha256:a458f6fb6a185656ce13e38851baf5fd375491f8510e5d2e1ef80dca41ce6df0

Observation 820ec6f5-c22a-461a-9ce0-6d02045a981c · outbound

This paper cites Achiev- ing cross modal generalization with multimodal unified rep- resentation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Achiev- ing cross modal generalization with multimodal unified rep- resentation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.140480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:32.142674Z digest=sha256:a11c10b3a9f7b7ebe454a4af6269ea7c7406a521d6d5b4e9dc0f878b06926d6e

Observation 19e3608f-e6cf-420e-aaaa-a654c73fda0a · outbound

This paper cites Hyper- spectral image classification based on unsupervised hetero- geneous domain adaptation cyclegan.

Open-set Cross Modal Generalization via Multimodal Unified Representation Hyper- spectral image classification based on unsupervised hetero- geneous domain adaptation cyclegan

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.914667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:32.246628Z digest=sha256:abe348b90a50c369883cf988932d7b042a93a51a74c80da7d3bf95b86e09291e

Observation 5d2da7bd-d3ac-45e0-a4d3-cf220f061a64 · outbound

This paper cites Class semantics modulation for open-set in- stance segmentation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Class semantics modulation for open-set in- stance segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.668841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:32.352440Z digest=sha256:700a81ad9643edadab30554e65f766a958170c2fc45ba35c0d8d5815ff977952

Observation 0c7c5007-6929-411b-8fac-ce29d6d3443e · outbound

This paper cites Learn- ing domain-invariant and discriminative features for homo- geneous unsupervised domain adaptation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Learn- ing domain-invariant and discriminative features for homo- geneous unsupervised domain adaptation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.436124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:32.477375Z digest=sha256:fdbda9ce6a01d85b2109ccb0b4aa2e2a8cb712608f0d88011bf71810719b8d15

Observation c55857b4-5572-411e-8a7d-b081986d96b4 · outbound

This paper cites A du- ality based approach for realtime tv-l 1 optical flow.

Open-set Cross Modal Generalization via Multimodal Unified Representation A du- ality based approach for realtime tv-l 1 optical flow

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.192212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:32.607661Z digest=sha256:e8a102f1f481e28297ac6bcb4232aed8d572a7bf1d84da86db235d49c681f00b

Observation 7f10b910-2872-4c0b-8b06-b8d5b1dc0448 · outbound

This paper cites mixup: Beyond empirical risk management.

Open-set Cross Modal Generalization via Multimodal Unified Representation mixup: Beyond empirical risk management

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:37.927785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:32.769497Z digest=sha256:6cc4f5f49223e560d36c41139d9484c9493f5186a6df3f688690bd3a11af7044

Observation 78b6b8c4-a9aa-4cd1-91a4-8751551b2670 · outbound

This paper cites Extending multi-modal contrastive rep- resentations.

Open-set Cross Modal Generalization via Multimodal Unified Representation Extending multi-modal contrastive rep- resentations

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:37.701085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:32.872091Z digest=sha256:0146a4bc8a993ffc99f3002958ce4f6f6417ed4298df7d539aaaecd56c26db09

Observation 92ae4c94-4bfa-47f6-af31-dc2573259f2d · outbound

This paper cites Towards effective multi-modal interchanges in zero-resource sounding object localization.

Open-set Cross Modal Generalization via Multimodal Unified Representation Towards effective multi-modal interchanges in zero-resource sounding object localization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:37.459916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:33.001801Z digest=sha256:8bb1a356753602b9d2cab8d381be98c7168d005a26a5864a5f48bb0072c0c3ea

Observation 55de53c6-1ece-419d-9a92-e46d4b6bf367 · outbound

This paper cites Positive sample propagation along the audio- visual event line.

Open-set Cross Modal Generalization via Multimodal Unified Representation Positive sample propagation along the audio- visual event line

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:37.179781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:33.180736Z digest=sha256:07cee9878a1945e0065d68c0d6bd6ece974003cc9372897e7f22c43d3b102a1b

Observation 97ddbd13-ffc0-4fa6-b81a-5493df3479fa · outbound

This paper cites Contrastive pos- itive sample propagation along the audio-visual event line.

Open-set Cross Modal Generalization via Multimodal Unified Representation Contrastive pos- itive sample propagation along the audio-visual event line

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:36.975633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:33.294913Z digest=sha256:4bc2c2c8ffc9f76fc647e8e1b8a12178bd4bc59a1310af504fcc47897212084b

Observation 9e964db3-0843-4c9d-8e06-6d5cebaacee9 · outbound

This paper cites As shown in Figure 4, applying the same mask to paired multimodal samples helps improve model perfor- mance.

Open-set Cross Modal Generalization via Multimodal Unified Representation As shown in Figure 4, applying the same mask to paired multimodal samples helps improve model perfor- mance

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:36.711555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:33.446319Z digest=sha256:0db36130526cef6976ec0ae4c0158255ef406bd2a4d8449f57448fecf79b476c

Observation ea7d8129-f526-46f0-a843-fcf27af389a9 · outbound

This paper cites As shown in Figure 5, we experimented with five different settings: 256, 400, 512, 800, and 1024.

Open-set Cross Modal Generalization via Multimodal Unified Representation As shown in Figure 5, we experimented with five different settings: 256, 400, 512, 800, and 1024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:36.441196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:33.590702Z digest=sha256:280c95e15639e93af936d5f960c8457854c0d35b293c016c7f43e9dc9f26e643

Observation de07682c-87f8-4353-a6c1-e34fea5fdacb · outbound

This paper cites Lcoarse serves as the foundation of the model, while Lf ine and Lcujp further refine the unified representation space and enhance the model’s open-domain detection capabilities.

Open-set Cross Modal Generalization via Multimodal Unified Representation Lcoarse serves as the foundation of the model, while Lf ine and Lcujp further refine the unified representation space and enhance the model’s open-domain detection capabilities

Reference 63

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T15:48:36.199359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:33.713818Z digest=sha256:970d99d8aee70727fe788c9d13ef625321f9dcf862ed609ef21ff1df464c45fe

Observation 0ed9da58-6fc8-4401-ab6b-a4ca137ccec8 · outbound

This paper cites CUJP8, despite having more split block reorder- ing, optimizes memory usage and reduces training time compared to MMJP6 [14].

Open-set Cross Modal Generalization via Multimodal Unified Representation CUJP8, despite having more split block reorder- ing, optimizes memory usage and reduces training time compared to MMJP6 [14]

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:35.956451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:33.844344Z digest=sha256:be64a5c907ce63afbd5e698e36d58f2b384d8cc0001dadfa241ae65ca37bfe3b

Observation 4d271cca-30a9-434f-9f9a-a9835f3a91a5 · outbound

This paper cites The visualization maps audio-video-text triplets from the Valor32K dataset [7] into the unified rep- resentation space (codebook).

Open-set Cross Modal Generalization via Multimodal Unified Representation The visualization maps audio-video-text triplets from the Valor32K dataset [7] into the unified rep- resentation space (codebook)

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:35.616927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:48:33.963779Z digest=sha256:08249d055c048832d451126d9ae752878747ff153acd3c7819ad7cd23a2a9761

Pith citing papers

Observation aa53c0d5-f2ba-4c21-9783-00fd73b7fa6f · inbound

TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal cites this paper.

TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal Open-set Cross Modal Generalization via Multimodal Unified Representation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:53:34.172007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T21:53:28.021948Z digest=sha256:597c7bc25e652953f2ae8f040c38d4edb1bd7ba71c2f126c46949c30e553b25d