Pith. sign in

Paper Citation Record · LEDGER

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing

As of 20 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2505.02096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02096 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T01:07:11.155102Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7dd5de0-2dab-41e4-a835-14a405e210d3 · outbound

This paper cites Unified multisensory perception: Weakly-supervised audio-visual video parsing.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Unified multisensory perception: Weakly-supervised audio-visual video parsing

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.469693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.072369Z digest=sha256:abf04c005aa66d0e357f340508f42b2952d337a24aa0e8113506437a4bbb6072

Observation 021b902c-6d23-4de5-b802-8a5810a5ba1f · outbound

This paper cites Anchor-aware Deep Metric Learning for Audio-visual Retrieval.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Anchor-aware Deep Metric Learning for Audio-visual Retrieval

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.453549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.077331Z digest=sha256:ae5347c034900242f0495f968b0e41dfd5bd0ee1778e553b4a7ff564d75aae70

Observation a3492174-f6e7-4e7d-9f35-a25b1b3c0052 · outbound

This paper cites Boosting Audio Visual Question Answer- ing via Key Semantic-Aware Cues.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Boosting Audio Visual Question Answer- ing via Key Semantic-Aware Cues

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.439219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.081688Z digest=sha256:afc64fc21c0e736dbcb88f4a313b9980a5d7ff72b0e3a54817b9e266cf6be90c

Observation add480ef-be21-4bf4-9aa9-760545df8d74 · outbound

This paper cites Open-Vocabulary Audio-Visual Semantic Segmentation.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Open-Vocabulary Audio-Visual Semantic Segmentation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.425110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.086055Z digest=sha256:3fd69bb9102d28b78d6d78cbaae77a8c8f867cf6374dc2f7e6f4d48281f4b3f0

Observation 58e56d40-8fb5-4b30-a1ea-3aa846ee1c41 · outbound

This paper cites Drcnet: Dynamic image restoration contrastive network.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Drcnet: Dynamic image restoration contrastive network

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.408437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.090531Z digest=sha256:5706f3ff0e01ccab714ad7fc625dbe61279e80ef6016f1527ef60b0cc7a0354a

Observation e0ab1285-db4f-42fb-80d3-4c5d3154ca63 · outbound

This paper cites Collecting cross-modal presence-absence evidence for weakly-supervised audio-visual event percep- tion.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Collecting cross-modal presence-absence evidence for weakly-supervised audio-visual event percep- tion

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.393321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.095159Z digest=sha256:62dff77476865a5bc69169939684d935e0ee605b216788ee998c49c699246706

Observation dae141fa-5161-48f0-ac2f-5ad330576e3f · outbound

This paper cites ColeaF: A Contrastive-Collaborative Learning Framework for Weakly Supervised Audio- Visual Video Parsing.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing ColeaF: A Contrastive-Collaborative Learning Framework for Weakly Supervised Audio- Visual Video Parsing

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.379326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.100048Z digest=sha256:777085593a2158fb0da828181d0e08e9856a8cbb5748128dea260b68742bafb7

Observation b7ea8ed5-813f-4743-a7c7-4aaa4fd7f381 · outbound

This paper cites Modality-independent teachers meet weakly-supervised audio-visual event parser.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Modality-independent teachers meet weakly-supervised audio-visual event parser

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.365057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.104350Z digest=sha256:6180a01d421bdacf88db673a67c772c18d5d31c180ce94bf991dc7da7d91dfc1

Observation 3046f933-f1a6-488f-89cc-edad067283ab · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.350667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.108434Z digest=sha256:ff4403c529deccc3a7b70f6f1d8b39678a54f0ea5904795e84ec3de3bdf0cbcd

Observation c576de82-eff6-4850-ac0a-16b5c962a0ed · outbound

This paper cites Learning transferable visual models from natural language supervision.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Learning transferable visual models from natural language supervision

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T01:07:11.113438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:07:11.113438Z digest=sha256:a6e30650b1cdb68783fb247d425798a6ce05ada81e6e0cb6bfade4c5655a0e61

Observation dc383dae-d85b-413c-9e92-fd23c178ec86 · outbound

This paper cites Label-anticipated event disentanglement for audio-visual video parsing.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Label-anticipated event disentanglement for audio-visual video parsing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.325932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.117790Z digest=sha256:74e136190bff0e2997bd217b64264d53d0e95482b7544ba1e45a18216ddfcf9e

Observation b2387bf3-1c95-4779-86e2-6a98eb59c3f4 · outbound

This paper cites Revisit weakly-supervised audio- visual video parsing from the language perspective.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Revisit weakly-supervised audio- visual video parsing from the language perspective

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.312269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.122031Z digest=sha256:dca370ac309ac2c14e84e0582a975fb8eb9dcde81b7301500fa7bfda8d8cb38e

Observation c0f97708-68f5-42ec-9bf9-e3c403287f8f · outbound

This paper cites Multi-modal grouping network for weakly- supervised audio-visual video parsing.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Multi-modal grouping network for weakly- supervised audio-visual video parsing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.297903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.126128Z digest=sha256:e0b870335c8c869a856a0bf5bc34ef00dcb22ea5260412e4f50c174593488859

Observation a522badc-0009-4868-8061-94284f83e6c2 · outbound

This paper cites CM-PIE: Cross-modal perception for interactive-enhanced audio- visual video parsing.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing CM-PIE: Cross-modal perception for interactive-enhanced audio- visual video parsing

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.283830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.130288Z digest=sha256:5af6dc7c791c1bddff1a3611d0182719a0e7ebe46999e6afc6a3f6727b828c17

Observation 205d1c3f-f0ce-4a7c-a39f-29915a5f5e2a · outbound

This paper cites Advancing Weakly- Supervised Audio-Visual Video Parsing via Segment-Wise Pseudo Labeling.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Advancing Weakly- Supervised Audio-Visual Video Parsing via Segment-Wise Pseudo Labeling

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.268622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.134351Z digest=sha256:5284ec298f91f426a02471bf91dc8980abf6906769bfce17f19b2e50ae48cf8a

Observation 6f958f28-13b4-4bd8-a62b-2ed47f291f51 · outbound

This paper cites Resisting Noise in Pseudo Labels: Audible Video Event Parsing With Evidential Learning.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Resisting Noise in Pseudo Labels: Audible Video Event Parsing With Evidential Learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.253974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.138245Z digest=sha256:ea7ec1f0b64fe38a81bec34564a69ea17f83e1f9d46cdc7ff522aafa3b4a773d

Observation 5ece17e2-c5a1-4539-b62f-01df159f5a7a · outbound

This paper cites LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-16T01:07:11.196864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.142298Z digest=sha256:af3a3a6555bbacca59ebc323d537a5ffd0a26aa6f4ab61b1a50cab3a2f2aa8fb

Observation 5270ffcd-4953-4a08-bf54-75c0b06c962e · outbound

This paper cites Multilayer perceptron (MLP).

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Multilayer perceptron (MLP)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.239969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.146746Z digest=sha256:02239e93cd08fde0caf8ef91c3dce25ad4c30b80f1d3511fca1585dbbcf7c72f

Observation ee8ed6a6-d325-4e62-8c34-e0eae99be691 · outbound

This paper cites Graph Attention Networks.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Graph Attention Networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.225269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.150922Z digest=sha256:06120c6da7cd88b84e0a4b53a64217b69ab32ac8818d0856ac8083d797e8c37f

Observation 7a03950c-6d74-458c-bc6f-f95aa5fb23ae · outbound

This paper cites Exploring heterogeneous clues for weakly-supervised audio- visual video parsing.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing Exploring heterogeneous clues for weakly-supervised audio- visual video parsing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:07:11.211167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.155102Z digest=sha256:02547ee8fe65ab0b79e6cbef4963f1cf2e33752b4467027a0803de5ebc6fea26

Pith citing papers

No inbound Pith citation observations are available.