Pith. sign in

Paper Citation Record · LEDGER

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings

As of 9 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2502.06012.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06012 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:09:18.680337Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9307a56-36dc-43b4-add9-feb4284afeb0 · outbound

This paper cites Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:19.091720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T17:09:18.540723Z digest=sha256:e24e8d59db4f816687593b550cf198542cbf0b51e8460f52cd3d85d792a6e514

Observation 47ddadcb-d087-4eb1-86d6-06434784b766 · outbound

This paper cites Combining Residual Networks with LSTMs for Lipreading,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Combining Residual Networks with LSTMs for Lipreading,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:19.074488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T17:09:18.547355Z digest=sha256:bd9f24b5b5932eab539e5f5da93b414dbb44b52a3a9bfc2ca962b517525cf344

Observation 4b4a39e7-a127-4a86-8281-b656e5e37680 · outbound

This paper cites Active Speakers in Context,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Active Speakers in Context,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:19.056810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T17:09:18.552428Z digest=sha256:fd535cd77403a1c9a955bbe082147c45ae16846a397ba3ede98693e06e98fcb3

Observation c1136810-d3a3-402f-819d-de2d8f289c8e · outbound

This paper cites ASD-Transformer: Efficient Active Speaker Detection Using Self And Multimodal Transformers,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings ASD-Transformer: Efficient Active Speaker Detection Using Self And Multimodal Transformers,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.557553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.557553Z digest=sha256:4ebcfe6b924800fa0eff6d941bed6c7071f57cd4bbb39b62b700b6bafd635c98

Observation 68a840b7-088d-4d2b-8965-d29d9c87d898 · outbound

This paper cites Hello! My name is... Buffy.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Hello! My name is... Buffy

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.564019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.564019Z digest=sha256:f01a20fffeafa83a0305d2e82c943a3dd027239dcba3c0fb5cc60ee2c37a75cb

Observation 09db0d23-3d42-46be-82b3-4652a159bc36 · outbound

This paper cites MAAS: Multi-modal Assignation for Active Speaker Detection,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings MAAS: Multi-modal Assignation for Active Speaker Detection,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.569695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.569695Z digest=sha256:df191d37806b8224d8cc893cab3f1a7814cf444eced27718d965bbfd1e0b06ce

Observation d87e82d8-2b51-441b-bd18-8c5834c0465f · outbound

This paper cites Target Active Speaker Detection with Audio-visual Cues,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Target Active Speaker Detection with Audio-visual Cues,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.576552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.576552Z digest=sha256:d4d4b39a5b3862005be043b1b2947537e14556f197e2df7b80692733b0f22530

Observation 6f99e09c-4b40-4109-b829-75bf046371a9 · outbound

This paper cites Ava Active Speaker: An Audio-Visual Dataset for Active Speaker De- tection,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Ava Active Speaker: An Audio-Visual Dataset for Active Speaker De- tection,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.581381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.581381Z digest=sha256:b056afa42759b9f818973027db17dcf87909806ea8275156c59add1448ff9a23

Observation ac4f4dfd-872a-45dc-891f-bf0e2010181a · outbound

This paper cites Improving Audiovisual Active Speaker Detection in Egocentric Recordings with the Data-Efficient Image Transformer,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Improving Audiovisual Active Speaker Detection in Egocentric Recordings with the Data-Efficient Image Transformer,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.988225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T17:09:18.585590Z digest=sha256:c6c636e1f4858b44c59c6170945981d9af2c5d1a181b42f344cbeee5dca87ce2

Observation 21ef0601-6f5b-49da-a5f2-e9ad0d680dfc · outbound

This paper cites Ego4D: Around the World in 3,000 Hours of Egocentric Video,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Ego4D: Around the World in 3,000 Hours of Egocentric Video,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.590424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.590424Z digest=sha256:959c0235cb2a2264b53f1298be6d606fbcf66f79efdcb5bc648cf8f00543c981

Observation 36daf742-d0ff-45dd-9b53-f0257df901ad · outbound

This paper cites End-to-End Active Speaker Detection, author=Juan Leon Alcazar and Moritz Cordes and Chen Zhao and Bernard Ghanem,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings End-to-End Active Speaker Detection, author=Juan Leon Alcazar and Moritz Cordes and Chen Zhao and Bernard Ghanem,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.962200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T17:09:18.594992Z digest=sha256:fd6c5f9ec399b8864c1c70d15b58554a177bd60b0d69dfe84a48453665f38001

Observation 04974023-5009-4843-8fbd-f24c2d8d0f4f · outbound

This paper cites How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the Wild,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the Wild,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.600233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.600233Z digest=sha256:497968602f2853a55bdf17c009dd69c51d0d265fb6e9935c74f8385d0af4167d

Observation eb15f594-52b5-4eb8-a1a7-cfaff91a8646 · outbound

This paper cites A Light Weight Model for Active Speaker Detection,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings A Light Weight Model for Active Speaker Detection,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.605275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.605275Z digest=sha256:b03caf9681577c2f7968462b7ced52eaf9b02f7b25de531f531ea0823b851bcb

Observation 672bc6f7-3ec5-4460-969d-ccaff54edcc4 · outbound

This paper cites LoCoNet: Long-Short Context Network for Active Speaker Detection,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings LoCoNet: Long-Short Context Network for Active Speaker Detection,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.609862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.609862Z digest=sha256:9fd533f3c42fcae306c436d99a88fec955c32d14ba7f7ba72ac6e34ea385dd99

Observation 1ba33c77-cfdf-4790-afb3-92d924c42c6a · outbound

This paper cites Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.615111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.615111Z digest=sha256:25229bd5d0846db85ffccbd2a644859724c5c9c2b309087ce3e987725f6276f7

Observation 7985d78f-7386-4bc4-bfe9-64a21de0d797 · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.619710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.619710Z digest=sha256:424fb07ebb13f4858c33abd6148ee2237a0d2af283c91ca57863247a94ade70f

Observation a9c9d624-29e9-44ae-9f97-53500061f0d5 · outbound

This paper cites Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-08T17:09:18.766401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T17:09:18.624973Z digest=sha256:64521cc98ab98fdc3db25d32a148104a443b6199e4dd1dfe5fbad93d3079ad3c

Observation 7eda9495-fd1b-4522-bbbf-c9191089d911 · outbound

This paper cites Efficient Personal V oice Activity Detection with Wake Word Reference Speech,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Efficient Personal V oice Activity Detection with Wake Word Reference Speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.904791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T17:09:18.631171Z digest=sha256:d0eb8048fe790f3a73bcdb59c4ae66f8ef4e3bf3df214f55b8f05f900fce952c

Observation 9fe9b316-a7de-4906-a1bf-4b9883193f88 · outbound

This paper cites V oxCeleb: A Large-Scale Speaker Identification Dataset,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings V oxCeleb: A Large-Scale Speaker Identification Dataset,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.636126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.636126Z digest=sha256:ef9d19adf08bea12b85b07997b8737890d28a8a86e7f1b1025b0e9a80720d422

Observation b8792db4-8364-4ab7-9d03-35e7e6946e63 · outbound

This paper cites Ring Loss: Convex Feature Normalization for Face Recognition,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Ring Loss: Convex Feature Normalization for Face Recognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.887481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T17:09:18.641060Z digest=sha256:a981a9a9512a510c7c253b053ac4e69d3b0fc1e425e8af60f190d35e5109a03b

Observation 1c34fe6a-207f-4beb-a411-a14f2c7ce986 · outbound

This paper cites ArcFace: Additive Angular Margin Loss for Deep Face Recognition,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings ArcFace: Additive Angular Margin Loss for Deep Face Recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.870149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T17:09:18.645922Z digest=sha256:f2d6db0828d3cbc3c380e0d8c1bd53565ecd140f5fe638d8dcf4443a472f7648

Observation 786bcc82-f5bf-4bc3-acaa-a9ff2a5cfed2 · outbound

This paper cites SphereFace: Deep Hypersphere Embedding for Face Recognition,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings SphereFace: Deep Hypersphere Embedding for Face Recognition,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.853270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T17:09:18.650052Z digest=sha256:3239c3cda495599000cb165d7fc5642a19414c079b475b6ff883556316eb2f60

Observation e72199d4-0acd-4d42-85b9-72cdffc84cec · outbound

This paper cites CosFace: Large Margin Cosine Loss for Deep Face Recognition,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings CosFace: Large Margin Cosine Loss for Deep Face Recognition,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.835514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T17:09:18.655038Z digest=sha256:7cb5a234584c887dc8bd4a6e3fce7af141889bb5700e20daf51fe061ff07cc32

Observation 599cbf32-1274-4f73-a393-10805583039d · outbound

This paper cites Partial FC: Training 10 Million Identities on a Single Ma- chine,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Partial FC: Training 10 Million Identities on a Single Ma- chine,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.819850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T17:09:18.660952Z digest=sha256:b999a393609500d895ff3524436d7caa3876d62d23a5910c6e4a987a05526e56

Observation 72c51122-fd47-4229-8770-65d107bced6b · outbound

This paper cites Attention is All you Need,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Attention is All you Need,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.803865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T17:09:18.665695Z digest=sha256:7f41e2d82b60012547c62f5fec4241861bf61b2603615bd679a0b62c10c00fdd

Observation 497d1872-2ede-461d-b85a-85a0398a7510 · outbound

This paper cites Robust Object Recogni- tion Through Symbiotic Deep Learning In Mobile Robots,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Robust Object Recogni- tion Through Symbiotic Deep Learning In Mobile Robots,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.670905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.670905Z digest=sha256:7a3297a4728521c0dce66315b1debd02b159883020f96b2e63ced0d9aefeb2e8

Observation a683ce62-4341-4bca-9986-5edeae446089 · outbound

This paper cites The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.675378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.675378Z digest=sha256:8b161ae65276f39f9dc50c55f14c80f3ed39fc40f1d1019e28c909d9b1f0a45a

Observation 23a06e36-1d2c-49dd-80aa-aab1d90798dc · outbound

This paper cites Technical Report for Ego4D Long Term Action Anticipation Challenge 2023.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Technical Report for Ego4D Long Term Action Anticipation Challenge 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.680337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.680337Z digest=sha256:9231760234a68f51d2bcf0cf995a947109558147c1e75b82afe97b57a35dc117

Pith citing papers

No inbound Pith citation observations are available.