Pith. sign in

Paper Citation Record · LEDGER

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models

As of 7 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2506.20361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20361 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:53:51.803073Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:53:48.450553Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:53:52.505268Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact3
  • verified fuzzy16
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cae30958-580b-4a12-898d-bb35d334a24c · outbound

This paper cites an unresolved cited work.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:53:56.440542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:48.237549Z digest=sha256:c142e3de3353207c4bc57b4d4c2d6acf5821d0dfd19123af491f205541ce0a81

Observation 24f8d8b4-f788-404d-b227-2091412a59d3 · outbound

This paper cites an unresolved cited work.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:53:56.252196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:48.651613Z digest=sha256:d2e3f35938fe0f4806b96f98aa1ed710dc3788efc0f06a141a896b9a858442e4

Observation c0134223-de03-4e42-b025-6ce6f5285cbb · outbound

This paper cites Dataset We analyzed the phonetic decodability window on two datasets separately.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Dataset We analyzed the phonetic decodability window on two datasets separately

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:56.095836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:48.823441Z digest=sha256:7defae49c56f0c6634da2587b5b819fd9686ec0eda87982edb344aed32d4a8c5

Observation 61bea0c4-4e63-4d3f-b9ac-cf73c541fca9 · outbound

This paper cites an unresolved cited work.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:53:55.892928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:49.013333Z digest=sha256:bb06010f1cc0b2d64bab4977c757e41b4adfe2347e301905cf6698e1cd345716

Observation 7eae74aa-131c-4dc9-8a0f-c363e80322ef · outbound

This paper cites We found that A V-HuBERT’s encoding of speech temporal dynam- ics is dominated by its audio input.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models We found that A V-HuBERT’s encoding of speech temporal dynam- ics is dominated by its audio input

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:55.743362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:49.188316Z digest=sha256:e31ee88d215046f7212eb8080692bf221d38bdb776a9460d32492b1e90974a39

Observation 29e31935-0a1a-4314-8027-bc8c1e01f636 · outbound

This paper cites The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:53:52.613791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:48.450553Z digest=sha256:dc82a884c939518154d8c8de7b16358fab9a5df97168aefbd407d9b426cc5a2a

Observation dc7c5188-171b-4f40-886a-d53c5020da7c · outbound

This paper cites We would like to thank Biao Zeng from University of South Wales and Hao Tang, Sharon Goldwater from ILCC, Univer- sity of Edinburgh for useful discussion.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models We would like to thank Biao Zeng from University of South Wales and Hao Tang, Sharon Goldwater from ILCC, Univer- sity of Edinburgh for useful discussion

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:55.556123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:49.377704Z digest=sha256:18097eae806ebaf59db80ebc52ff56d7e932ad4cdf925c452d9d1a67388e8557

Observation 64abd8f1-82eb-4fbe-abb1-9639bad530a5 · outbound

This paper cites Using artificial neural networks to ask ‘why’ questions of minds and brains,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Using artificial neural networks to ask ‘why’ questions of minds and brains,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:55.394127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:49.498398Z digest=sha256:27f5d4b316e34d9cf29f67195c6bdec55bff8065d831746c7b81804688bd8110

Observation b48a8210-51ec-4cc5-9f5f-21253ce644b2 · outbound

This paper cites Parallel hierarchical encoding of linguistic representations in the human auditory cortex and recur- rent automatic speech recognition systems,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Parallel hierarchical encoding of linguistic representations in the human auditory cortex and recur- rent automatic speech recognition systems,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:55.241297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:49.674523Z digest=sha256:48e58dbe54fe1d81f29f1ae9bb21a33e1ef1ee1ed175e6979b28512d5e4aacf2

Observation 8ec8c1a4-95b3-4bb8-a244-04ffd366196b · outbound

This paper cites Hearing lips and seeing voices,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Hearing lips and seeing voices,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:49.888090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:49.888090Z digest=sha256:1cddb7e0ce99c6c36ba3d7ada5d1143102d8d5382831919b61254930455edbf0

Observation 8cc38be9-dce2-4cfe-aa45-8ac763842dd8 · outbound

This paper cites Multisensory integration: current issues from the perspective of the single neuron,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Multisensory integration: current issues from the perspective of the single neuron,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:55.026324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:50.029945Z digest=sha256:57bd78c8b8893d626549b8ba601118626dab5bcb96f6429ee637078d8f954837

Observation 0059e289-68da-490d-b5e5-6050a2e55bba · outbound

This paper cites Multimodal deep learning.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Multimodal deep learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:54.856937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:50.185624Z digest=sha256:16ad4a348ff56d1dff5f7679f84a601f855065c2f2bc18b4205c7053a95133d6

Observation b6c4587a-83bb-48dc-a28a-7e80954ebd17 · outbound

This paper cites On the role of noise in audiovisual integration: Evidence from artificial neural networks that exhibit the McGurk effect,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models On the role of noise in audiovisual integration: Evidence from artificial neural networks that exhibit the McGurk effect,

Reference 13

Resolution
verified exact
raw_fallback, observed 2026-08-06T22:53:52.389153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:50.297411Z digest=sha256:42a0b8fbfc4f0fa4212ef4d8969fc9a786b38405382b88ef3ad395350fb82373

Observation fd50aea4-84af-4bc9-8022-d12eb1547d04 · outbound

This paper cites Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:50.391701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:50.391701Z digest=sha256:e015d141b67043b380a16fded786e88e9149495e6a295d97c9585803145a3627

Observation 28fdc799-899e-4acf-b4c6-6a96d44263e6 · outbound

This paper cites The natural statistics of audiovisual speech,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models The natural statistics of audiovisual speech,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:54.662843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:50.483547Z digest=sha256:c42954aad06cde72b08cee9fd6f645f45bfd79c6610c0b8a26a87f0377cb4a37

Observation 71710634-2011-40e9-a40e-26b6ceaf68bf · outbound

This paper cites Bimodal speech: early suppressive visual effects in human auditory cor- tex,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Bimodal speech: early suppressive visual effects in human auditory cor- tex,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:54.377216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:50.593870Z digest=sha256:a95086ca41d4223d341aaa64d4aab4de3abca75ae490852d4c5b128425285341

Observation 7fd2bde6-cde9-4a31-bc63-52dfa385c0da · outbound

This paper cites Visual speech speeds up the neural processing of auditory speech,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Visual speech speeds up the neural processing of auditory speech,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:54.175482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:50.668168Z digest=sha256:9ffd948b683c78c275f09133e893e205e981a1a178cb6d412bb2510672013029

Observation d470de32-c3e2-4f2a-9d10-d283a1e640c9 · outbound

This paper cites Asynchronicity between visual and auditory information in audiovisual speech: Evidence from four types of consonant- words/b/,/t/,/k/and/g,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Asynchronicity between visual and auditory information in audiovisual speech: Evidence from four types of consonant- words/b/,/t/,/k/and/g,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:53.994140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:50.739263Z digest=sha256:45102b3162526404a35700e8128bae49ac4b26977d0a8c9105c285781653116c

Observation e307496b-1ff3-4266-ad1e-6a2c3f165244 · outbound

This paper cites A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:53:52.083800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:50.840056Z digest=sha256:4f53106668cca8b7f03d09badeabc136cbfa2157b81450d7c0e3bb93b27e5477

Observation 38ebd5c8-3751-4a3f-8bd7-7411414dde71 · outbound

This paper cites Neural dynamics of phoneme sequences reveal position-invariant code for content and order,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Neural dynamics of phoneme sequences reveal position-invariant code for content and order,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:53.724658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:50.949106Z digest=sha256:b413cf631890c25d7304d96656e8880956425fa04da8e6185f209d45d1935333

Observation 123861fc-64e1-40ad-9645-3e6dc382826e · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:51.030814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:51.030814Z digest=sha256:08a05e988213376d51968f12c39dd75c995ccb183498efeac4134286387e955a

Observation d15b4bb8-4347-45bc-9b1a-8b7614b2d6f7 · outbound

This paper cites Deep residual learning for image recognition,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Deep residual learning for image recognition,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:51.105709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:51.105709Z digest=sha256:5ea20efb4d131b1f5ecb5514afc9b1ca0063eb7406ee5aadbcb0fd44673c32db

Observation 22583bb2-f733-4bba-b601-d0c22e9abf22 · outbound

This paper cites Dynamic encoding of acoustic features in neural responses to continuous speech,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Dynamic encoding of acoustic features in neural responses to continuous speech,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:53.448903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:51.187054Z digest=sha256:0ab7799a9e33ef713ec1f1df18377f248791ad797633c9a5d01fab60d7d4f8ab

Observation e5c540e0-211b-4e72-91e0-98f4ac694089 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Representation Learning with Contrastive Predictive Coding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:51.262919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:51.262919Z digest=sha256:32efcfc1890fa32f11390e287e74bd335f0ec49d09988ebaf798d1b182a6a0bb

Observation f4d1bdad-7f83-49d2-89a1-2787d06aa1af · outbound

This paper cites Montreal forced aligner: Trainable text-speech align- ment using Kaldi.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Montreal forced aligner: Trainable text-speech align- ment using Kaldi

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:53.205875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:51.369421Z digest=sha256:9a37f6ed8715b6ac1bd74c7f4af0a7028a04bf37b4fa8082156e9190fedbdf05

Observation 1fb8ac9e-8336-43ee-b842-f7e5a04fb035 · outbound

This paper cites Prosodylab-aligner: A tool for forced alignment of laboratory speech,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Prosodylab-aligner: A tool for forced alignment of laboratory speech,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:52.994338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:51.485127Z digest=sha256:e31039a0bbacd03b01670118b9440195e1921eea7d52998268a54e5df6c75503

Observation 216414d2-5e51-46c8-94e4-9dcd976e5bcb · outbound

This paper cites An audio- visual corpus for speech perception and automatic speech recog- nition,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models An audio- visual corpus for speech perception and automatic speech recog- nition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:52.804935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:51.637293Z digest=sha256:f32db47776120986b68556e67b0703e80fedbb45e228c8da25cbc377c008210f

Observation a70d74ae-21d5-4051-92ac-96d592074ef8 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models LRS3-TED: a large-scale dataset for visual speech recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:51.803073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:51.803073Z digest=sha256:55cc81d2feea87ebef9da718aa134c47b5f41107ee620318d28eeaa58c6abedc

Pith citing papers

Observation 29e31935-0a1a-4314-8027-bc8c1e01f636 · inbound

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models cites this paper.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:53:52.613791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:53:48.450553Z digest=sha256:dc82a884c939518154d8c8de7b16358fab9a5df97168aefbd407d9b426cc5a2a