Pith. sign in

Paper Citation Record · LEDGER

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition

As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2507.09323.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09323 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:03:03.579463Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 81fd07c6-b149-4500-99ae-b80d39c05c51 · outbound

This paper cites Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:04.021804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:02.905926Z digest=sha256:e73973b5ae1216c424625e2462f18358e811febf83b5da0ac5965dbfeae8ea96

Observation cefb466b-5949-46af-bcea-ebb65e9331cc · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:04.007900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:02.983484Z digest=sha256:ea8df9001603b21ea80d95a90f7b7c3f34457ab44b52d0d10107be979cd46ca7

Observation b0797321-987a-43db-877a-5e0851227593 · outbound

This paper cites Multimodal machine learning: A survey and tax- onomy.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Multimodal machine learning: A survey and tax- onomy

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.994738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.160526Z digest=sha256:fe3a3bda1b22f317ef8da458234a90d447b5bd870f391261ebf3de54dcdf92a0

Observation 8ee51abf-9a5d-485d-9cd4-b7c4d669e9d0 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Vggsound: A large-scale audio-visual dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.980373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.298243Z digest=sha256:5e8e13e6b1949565678fef3d14b799d925c3253b4cab9ac354f456a4a4a3e1c7

Observation 7537496e-dbc4-4ce6-b0a8-25fb935740eb · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition A simple framework for contrastive learning of visual representations

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.965880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.420105Z digest=sha256:6847c70df2ad876c6c2149340118153f8aa1694e0c4a9efbd754595458fe52d2

Observation 87a38546-9c25-4da7-b423-bcd7692e4e77 · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Scaling egocentric vision: The epic-kitchens dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.951067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.475163Z digest=sha256:044431a209b5c01ee6bd3a9e4d8d18f02c733b38b978c22c395b71d6f16b0b42

Observation 08276ac1-58f2-4dcf-be21-fe5fceb7038d · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition An image is worth 16x16 words: Transformers for image recognition at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.936862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.480607Z digest=sha256:7cd00b720d45a07f85c97a5303461198850fb1f5c28563c18854ee1ce8d43a62

Observation cda91782-4f55-477d-a412-4ffe5229e882 · outbound

This paper cites Large Scale Audiovisual Learning of Sounds with Weakly Labeled Data.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Large Scale Audiovisual Learning of Sounds with Weakly Labeled Data

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:03:03.685469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.485282Z digest=sha256:bdf42f3e215a4377c6e664943872264324f64a94bb260e67ba640b922f74dec2

Observation db469e9c-eab5-497f-84ec-e9d938026c04 · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Audio set: An ontology and human- labeled dataset for audio events

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.922445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.490036Z digest=sha256:2d91a6a4a460c2fcd0330d308e651173ae4310ff4275b2f680ca981db2c74529

Observation 13d93ebc-8506-4e7e-ac6c-8d0f6ba74705 · outbound

This paper cites Audiovisual masked autoencoders.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Audiovisual masked autoencoders

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.909656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.494154Z digest=sha256:1afea03b75e0c8b3fe88317c4a925c4adc32ffa5189871ed5e233f92c3e60d0d

Observation 8eeb24fa-c3e4-4bf2-bce0-6e3ca5a37666 · outbound

This paper cites AST: Audio Spectrogram Transformer.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition AST: Audio Spectrogram Transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.498534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.498534Z digest=sha256:45df32776ebf2288413abd01da454c1b733aa7d2a20a890e0be155d910b9dd87

Observation 75677a28-1a87-4e6a-8b4c-a88b97687b0e · outbound

This paper cites Uavm: Towards unifying audio and visual models.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Uavm: Towards unifying audio and visual models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.896484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.503076Z digest=sha256:f28b63407b443226902499306dd2bde38d93151c658b9d7e1e300411ae2e483c

Observation 8696b57a-3e26-4303-8f96-eccab980b3f1 · outbound

This paper cites Contrastive Audio-Visual Masked Autoencoder.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Contrastive Audio-Visual Masked Autoencoder

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.507305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.507305Z digest=sha256:f8866ee2d2499637d947e21ce4e258691584c6e4b887897f5bcb8eafcf3a3e00

Observation 43fe541e-1486-4185-b0fb-1a857cf9e9d8 · outbound

This paper cites Deep residual learning for image recognition.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Deep residual learning for image recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.511781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.511781Z digest=sha256:5cffef405a5e279aa43699601c7b87d849fc7f3b91fea55483098436bf62f75c

Observation 34a4a523-6847-4011-a690-2d8c5743b512 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.875713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.515621Z digest=sha256:5b9ddb60bce4a52415888974e350f46afa69590c9254dbf1a405cd7d3608969a

Observation 66f5017d-ab57-4e8a-9d3d-ba3f78cce6a2 · outbound

This paper cites Perceiver io: A general architecture for structured inputs & outputs.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Perceiver io: A general architecture for structured inputs & outputs

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.862821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.519693Z digest=sha256:ea84f7f24415435b2246b42381895acd54d1e23d2e82895551c4e9b9149f2e8a

Observation 9fd1149d-b653-4743-b5b1-1c9cda6cd8de · outbound

This paper cites Supervised contrastive learning.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Supervised contrastive learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.523260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.523260Z digest=sha256:4ba8a96b30eb3323e099c661b87b1f26e7fef50a50035b83db6331fae7cf1fe9

Observation 3dfc1e15-bec4-4adf-9842-a7c36837af65 · outbound

This paper cites Kipf and Max Welling.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Kipf and Max Welling

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.840950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.527353Z digest=sha256:fda50bd1aecaaee10971aaad0a04d45a47feba355549b972b65727b019b80674

Observation 91f8c0f3-0761-4cd8-88b6-c75018c4447e · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Swin transformer: Hierarchical vision transformer using shifted windows

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.531207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.531207Z digest=sha256:1d1765070568b6bcadb03e9e9bae21bf2bb62a138b320c0326c3ebe55b4be8c3

Observation 8884ea43-29e0-40e3-97c3-353b84b113d4 · outbound

This paper cites Attention bottlenecks for multimodal fusion.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Attention bottlenecks for multimodal fusion

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.817741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.534986Z digest=sha256:836abd1f53bb3ee185ce226411a34c25c396a4d5a69dc5bff0104ca5921be07b

Observation d8e499b7-5b22-433e-a139-d600351e428f · outbound

This paper cites OmniNet: A unified architecture for multi-modal multi-task learning.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition OmniNet: A unified architecture for multi-modal multi-task learning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:03:03.635985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.538749Z digest=sha256:96a85f88f3ccfd03501238e2a99abf994b6a0039cbe9cd593fe0cf898cd1fe46

Observation cbb55465-42fc-4aad-af02-d8caa9f7e5b2 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Robust speech recognition via large-scale weak supervision

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.804017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.542837Z digest=sha256:5c5f850958e1944b9800bc756570530c2fa6c0b1a9af5fdde06fd572ee3aa5ba

Observation 45308ca1-9e25-49cf-8ec7-039260699ba7 · outbound

This paper cites Multimodal fusion for audio-image and video action recognition.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Multimodal fusion for audio-image and video action recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.789208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.546642Z digest=sha256:4c649ee66034b6a4a7d32adb188619759b05f013f4017610ef8122f6006bf762

Observation 93092bbe-2d4a-4d8b-9715-0b47c7270d60 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.550744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.550744Z digest=sha256:f33fd57a2baf9de7737497927b4b4e0943b7bfff6eb6d1a6afdb873207e57b52

Observation 69247b40-9d48-4fe2-82f4-5a69517c00d3 · outbound

This paper cites Omnivec: Learn- ing robust representations with cross modal sharing.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Omnivec: Learn- ing robust representations with cross modal sharing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.776085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.555128Z digest=sha256:3c6b741d9b366894f398beb2125c66ea59cb44d6374d21b8786be5db59b34933

Observation 43b5864b-2380-435f-9838-5310b5699aad · outbound

This paper cites One-peace: Ex- ploring one general representation model toward unlimited modalities.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition One-peace: Ex- ploring one general representation model toward unlimited modalities

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.762263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.559513Z digest=sha256:7cec3bc37e8f329d81cb3bd866c63f65164455c25a7eb633f52663c051781b63

Observation 7db7b7c4-35da-46b2-9112-372c202ae259 · outbound

This paper cites What makes train- ing multi-modal classification networks hard? In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12695–12705, 2020.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition What makes train- ing multi-modal classification networks hard? In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12695–12705, 2020

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.749426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.563766Z digest=sha256:d72202ff91a5a2a15565cdf6f5e4d734a724819d4f61cfefd435d99f487f3780

Observation 1ee6d336-c67d-4573-95c9-601d530a2871 · outbound

This paper cites Multi-stream multi-class fusion of deep net- works for video classification.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Multi-stream multi-class fusion of deep net- works for video classification

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.736313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.567838Z digest=sha256:3efe297ab630873fc791690ebf1fc6f51aba60ec11d89b51c9367bc31a82663a

Observation 430b57eb-fd29-4b5e-a891-bc128515017a · outbound

This paper cites Multimodal learning with transformers: A survey.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Multimodal learning with transformers: A survey

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.571745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.571745Z digest=sha256:b57dccd5e2420cb2b92e4e5615922d3ea06cb4194601d818d6f0a5ebdc64eda4

Observation 386e133a-4d35-425b-b3ee-13d6a532b2b8 · outbound

This paper cites Peters, and Yejin Choi.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Peters, and Yejin Choi

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.712763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.575757Z digest=sha256:eb058ebf07f59bda62e041d5a1dcdbcc2612c95f00ec129813cfda0aaa2c8e91

Observation 43cce9bc-5253-4dba-a12c-ba2f1337c6fd · outbound

This paper cites Multimodal representation learning: Advances, trends and challenges.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Multimodal representation learning: Advances, trends and challenges

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.699421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:03:03.579463Z digest=sha256:df91436ecf95f19810f27960b96071475e0173413c254495c1797d296c8d6924

Pith citing papers

No inbound Pith citation observations are available.