Pith. sign in

Paper Citation Record · LEDGER

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning

As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2606.11602.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.11602 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T10:39:37.920953Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 407f5517-65c5-49dc-9f2c-0ecbacd8af5f · outbound

This paper cites Look, listen, and attend: Co-attention network for self-supervised audio-visual representa- tion learning,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Look, listen, and attend: Co-attention network for self-supervised audio-visual representa- tion learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:abe6a28cb0ab9b0ba3b5082ad592b3dda9dde6412dadf35e697f3d0294549c61

Observation 427b5d01-357a-46bc-9155-2cea6392bb8c · outbound

This paper cites Self- supervised object detection from audio-visual correspondence,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Self- supervised object detection from audio-visual correspondence,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:ed132b62367fd75939c94330497d05cc2c5769043fb32b0a1c176980080c9d52

Observation 7705b846-6f68-461e-8303-2d18d069dbd6 · outbound

This paper cites Detection of audio-video synchronization errors via event detection,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Detection of audio-video synchronization errors via event detection,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:776c6da5fcd917a7480b982a454c7700a56e42e940b8968419f5d2423dc69fdd

Observation 8a8885ba-b003-423e-a10c-cc72a3821ff8 · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Audiovisual SlowFast Networks for Video Recognition

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:57:47.661964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:59464d96b93ad8f773309f7fa8eb5037b8ee01c57cd5dc6b7af9470990557777

Observation d06aa5fc-5c46-4c6a-818b-64f8fad912a4 · outbound

This paper cites Text-to-feature diffusion for audio-visual few-shot learning,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Text-to-feature diffusion for audio-visual few-shot learning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:e5917c3c9d60166ee94fef4571373f2dcc53cfa837e7db20158f48912f749829

Observation 54404212-e78c-4650-b4f9-5933a0f61669 · outbound

This paper cites Advancing weakly- supervised audio-visual video parsing via segment-wise pseudo label- ing,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Advancing weakly- supervised audio-visual video parsing via segment-wise pseudo label- ing,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:c06d0f10a5ba4dbd279f79929b8d4ef6616d83357e0c6ed68a399b5bced482ef

Observation 5153b275-6677-40c3-a6cb-48ccbd299d6a · outbound

This paper cites Aloha: Adapting local spatio-temporal context to enhance the audio-visual semantic segmentation,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Aloha: Adapting local spatio-temporal context to enhance the audio-visual semantic segmentation,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:630077e795ab25b6dce2af7787a93628f42cd83fc6c629f71152ba2ccc726b30

Observation 7d3c7e45-ce3d-4c2c-938d-a7f0aed9b531 · outbound

This paper cites Audio-visual event localization with cross co-attention and dynamic audio-object semantic alignment,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Audio-visual event localization with cross co-attention and dynamic audio-object semantic alignment,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:8fbf8d4fc621a51906ed210cc6fb42b1f963caaabf09aa59cfc0d70d455cafc6

Observation 2959b878-b23b-4c2f-8b06-ca6442f5a6cf · outbound

This paper cites X-sta: Cross-modal spatial-temporal alignment network for unified audio-visual segmenta- tion,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning X-sta: Cross-modal spatial-temporal alignment network for unified audio-visual segmenta- tion,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:18bff5884c3560f225e8115a0f9f1687c2211234b26ee07af2678fd936bb2b91

Observation fe56af1e-d90c-4690-b982-ded29839d47e · outbound

This paper cites Coordinated joint multimodal embeddings for generalized audio-visual zero-shot classifi- cation and retrieval of videos,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Coordinated joint multimodal embeddings for generalized audio-visual zero-shot classifi- cation and retrieval of videos,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:d8e1c147c4919c454f45b9ec1dd4d9fe8cf6fc69f91f42261bcdde0cca6788a3

Observation 8b59b25f-814a-4d8c-8653-d6283f5eacd5 · outbound

This paper cites Avgzslnet: Audio-visual generalized zero-shot learning by reconstructing label fea- tures from multi-modal embeddings,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Avgzslnet: Audio-visual generalized zero-shot learning by reconstructing label fea- tures from multi-modal embeddings,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:e7a930b090f2bb673aeb494ae534dfb0c4df4b377a7bca9254522c1afc0f8f7a

Observation 16a33141-9b62-4e3b-92f1-47baa334d382 · outbound

This paper cites Audio-visual generalised zero-shot learning with cross-modal attention and language,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Audio-visual generalised zero-shot learning with cross-modal attention and language,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:2de7a0e61e071a557b8c91c06404b98a8067e36144474010431b672dc0ec2946

Observation 5a0dd993-46ba-4cd9-9439-0e3d68b41b26 · outbound

This paper cites Temporal and cross-modal attention for audio-visual zero-shot learning,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Temporal and cross-modal attention for audio-visual zero-shot learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:c131a9cf34bf8f8e1d3e5892cd464bc49e23ac4c29d6177b81674d073bd5b35f

Observation 1d9a12c0-d7d0-48e7-8569-fd45e1995514 · outbound

This paper cites Hyperbolic audio-visual zero-shot learning,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Hyperbolic audio-visual zero-shot learning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:cdf95747d08176bdd123f66b93d5547d8827e82145ed1c00216e5dec88e20da3

Observation 0fd3846c-9034-4dba-b5f0-500a7bb01383 · outbound

This paper cites Motion- decoupled spiking transformer for audio-visual zero-shot learning,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Motion- decoupled spiking transformer for audio-visual zero-shot learning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:d6b2feb64f219329b5ed4d62d2a1673b789f2820f5317c47d40cc72ba288375f

Observation 3da2b60a-7be1-4598-bb26-53eee3b2f6c1 · outbound

This paper cites A generative approach to audio- visual generalized zero-shot learning: Combining contrastive and dis- criminative techniques,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning A generative approach to audio- visual generalized zero-shot learning: Combining contrastive and dis- criminative techniques,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:1e84255654da9356e52d2c5557d09b0bae4bb381f7120f9aac81d1ae917d8cca

Observation 70b2dc0a-9254-43b3-89da-0204dc39aa1d · outbound

This paper cites Spiking tucker fusion trans- former for audio-visual zero-shot learning,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Spiking tucker fusion trans- former for audio-visual zero-shot learning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:ee96d7ec58ec89a319fe90d0d04eab1587390f4d6dd12645b8c8df06455efb6b

Observation f343b011-37cc-45a4-b789-b76f70a541bc · outbound

This paper cites Audio- visual generalized zero-shot learning using pre-trained large multi-modal models,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Audio- visual generalized zero-shot learning using pre-trained large multi-modal models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:8a5ad068a84d3bda03db0ccee0ab9d2a8989b1a23cb2cac48521cea1f23ffeb4

Observation 3b6ef146-2b3f-4e39-8d5c-9265421c2bb5 · outbound

This paper cites Audio-visual generalized zero-shot learning the easy way,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Audio-visual generalized zero-shot learning the easy way,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:3d576a09c44626c32ed2e937c5cd55cd8ddbcd9742db369f05edaf64f4df3648

Observation 58a3d039-e239-4f10-8228-e408eb613dc7 · outbound

This paper cites Discrepancy-aware attention network for enhanced audio-visual generalized zero-shot learn- ing,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Discrepancy-aware attention network for enhanced audio-visual generalized zero-shot learn- ing,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:1e88cdc4d732afc8a6651bbb61e601639c8adb0ecff600cbe55119419494ddf8

Observation b1f66517-45b1-4e3a-b415-0024d349a36d · outbound

This paper cites Fusion-regularized alignment modality-adaptive audio-visual network for audio-visual zero- shot learning,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Fusion-regularized alignment modality-adaptive audio-visual network for audio-visual zero- shot learning,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:64c1686bb3ba18c2fb1efb7a1cac9fa9b644680d983571412d4916fab0b2477e

Observation f50ba71e-18b2-47c4-8a6b-5c24760609e0 · outbound

This paper cites Z-score normalization, hubness, and few-shot learning,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Z-score normalization, hubness, and few-shot learning,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:dc2d2e84e32dca1ffc02df7d138969da40e8fd90f7363756244c69d760e1afd3

Observation b1d40317-1968-4647-b5f6-d577c1b754b8 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Learning transferable visual models from natural language supervision,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:2d6b136e7461dac11a2322d64bb2e7f7ef1666321ac54488e181f42790b16b6b

Observation 05d828a6-f434-4c1c-8347-61637631926e · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:ea7608bf9e3192f461940cc5b83ce33b3008061352de12c3daa9ca8bffc176d2

Observation 4d8135ae-73c5-4da3-84e0-ad35475fea95 · outbound

This paper cites Tackling uncertain correspondences for multi-modal entity alignment,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Tackling uncertain correspondences for multi-modal entity alignment,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:82b3967c82777f27c0161e169f794409ffe422004ea4b573ac9446f57dcef880

Observation 22b77b5c-c76d-4c97-85a9-c2d1d97125c5 · outbound

This paper cites Unialign: Scaling multimodal alignment within one unified model,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Unialign: Scaling multimodal alignment within one unified model,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:4755de7e212bd48ea95a870f8125d8d48f0e663ffb79d5b04952a292233abbc8

Observation a6312b14-7efb-4631-9251-229ff764654a · outbound

This paper cites Relational knowledge distilla- tion,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Relational knowledge distilla- tion,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:6bdd16f4b38ed8131067658eedb113c6a65984845876627f37df98395ab662b5

Observation f54605b0-cedf-401e-ba73-e45c41f2ca84 · outbound

This paper cites On Distilling the Displacement Knowledge for Few-Shot Class-Incremental Learning.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning On Distilling the Displacement Knowledge for Few-Shot Class-Incremental Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:57:47.664025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:3151ce6bd963c78cf7651f75d47335ab9ddb610c46d396bfaf91032f2bded600

Observation 8ad004f7-4a2a-40d2-beff-f16e962c6e0c · outbound

This paper cites Extremely Simple Out-of-distribution Detection for Audio-visual Generalized Zero-shot Learning.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Extremely Simple Out-of-distribution Detection for Audio-visual Generalized Zero-shot Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:57:47.661468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:33cc090997510577fb35c515ab7e8c9f9d69ce270710a70efa151f1ad0340c20

Pith citing papers

No inbound Pith citation observations are available.