Pith. sign in

Paper Citation Record · LEDGER

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness

As of 13 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2607.08112.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08112 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T12:53:22.119074Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact5
  • verified fuzzy24
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77885fbd-0945-44e7-86bc-be2bcc157ec9 · outbound

This paper cites LipNet: End-to-End Sentence-level Lipreading.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness LipNet: End-to-End Sentence-level Lipreading

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.609021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:4e6a52307e6ef49137d3f8e6ea571484b8ad05898515aa5cd777aef647083c52

Observation b74b966e-d759-4cf7-ad9e-5132fcc1e64d · outbound

This paper cites Lip reading in the wild.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Lip reading in the wild

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.798938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:6c0369cbd21c75d381a5d9e7c265a4c3bd5e37c0ac098ec3a35dbfa1ce538264

Observation 529f548e-fc6d-4dd8-999c-290708cc5386 · outbound

This paper cites Lip reading sentences in the wild.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Lip reading sentences in the wild

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.766741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:c62f7fe010cb81b0dee5e2efd72409d273f4b874a414e095287738cab0ee7e41

Observation 42e8e4b9-097d-493d-9bdf-5f18b0fc8fd5 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness LRS3-TED: a large-scale dataset for visual speech recognition

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.605649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:4489ab6702429f1854d51c59a3bad1beceb03b8eb4951c6867ca3ae83d42f1d2

Observation 79c96a54-7300-4bee-b00f-38a6e5f6d1d9 · outbound

This paper cites End-to-end audio- visual speech recognition,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness End-to-end audio- visual speech recognition,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.793434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:54df1617fbdff4a3eeb7f0a571b40e276e810ad7e54b78abff38a8c361737816

Observation 7bcee6f3-6794-470b-9331-e67a34ba5fbc · outbound

This paper cites Lipreading using temporal convolutional networks,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Lipreading using temporal convolutional networks,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.788000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:aa582d22bd83f053cf09bc95bddaf72b0a985bc79af27077fd61d59e0986ab79

Observation bd7ccf15-1558-4d13-aac9-b69e475421bb · outbound

This paper cites Training strategies for improved lip- reading,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Training strategies for improved lip- reading,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.795076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:d7908cafbc4a310067bc50cca3382a7031244e170fc012d05d0be002610ec01d

Observation 7c97ae26-d06e-4bd5-a39d-547fc6051e06 · outbound

This paper cites A multimodal german dataset for automatic lip reading systems and transfer learning,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness A multimodal german dataset for automatic lip reading systems and transfer learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.779253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:7bec1d9e4657f886f247676afb949fb3b32da79346055ceba979ea465ad3bbdd

Observation 8432314e-b67e-41bc-8f0f-eb825ce1a1cc · outbound

This paper cites Visual Speech Recognition for Multiple Languages in the Wild,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Visual Speech Recognition for Multiple Languages in the Wild,

Reference 9

Resolution
verified exact
doi, observed 2026-07-10T12:57:07.523625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:783393b56d8acea986b468f388b3fde68cc9d5d9dffd65f4aeb92a4106ed3130

Observation e6ff0fab-f406-4880-8956-bc472252b2e4 · outbound

This paper cites Scaling multilingual visual speech recognition,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Scaling multilingual visual speech recognition,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.786234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:86bc6925fc7c49da0be16fa16676cf74edc9ccd9d2e81f80c1460599d1ff8293

Observation abd173d5-ed40-43b3-a6b6-279c9b085549 · outbound

This paper cites Visual speech recognition for languages with limited labeled data using automatic labels from whisper,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Visual speech recognition for languages with limited labeled data using automatic labels from whisper,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.777454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:01fddc6b7181210db199b17b4033d732e71ffcfc3f45ac4babd5d4db7d43aaa9

Observation ceffa10a-dd0f-4d26-8f47-471fbe63cd37 · outbound

This paper cites Lrro: a lip reading data set for the under-resourced romanian language,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Lrro: a lip reading data set for the under-resourced romanian language,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.759395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:44bcac037e13d65ec963ec88c3b8482e9fa57c04d997529e4332bdede68724fa

Observation f4714316-4d92-419d-a0f8-f6eff61246ad · outbound

This paper cites Toward language-independent lip reading: A transfer learning approach,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Toward language-independent lip reading: A transfer learning approach,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.800840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:f68e855e1cfe13583c2000f33493fc98994f589a08203be3c324ad30e27939e8

Observation 4bfc1d75-93ba-4c70-93e5-5097e924385b · outbound

This paper cites End-to-end lip reading in romanian with cross-lingual domain adaptation and lateral inhibition,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness End-to-end lip reading in romanian with cross-lingual domain adaptation and lateral inhibition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.768627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:7925de107c864a5f2f523b5bc395abdb29f6ab797f99b11e20fc02a3e1ebca36

Observation 3c8376d8-ff51-41ea-b8b4-869f1e6e7bef · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Robust speech recognition via large-scale weak supervision,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.772266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:3ba4053b254bb04061888d5caf2f44a9a035f8f091d9f1f71f727091e1910251

Observation 0de941e6-3fcc-438a-9517-7541fe6a68d5 · outbound

This paper cites An audio-visual corpus for speech percep- tion and automatic speech recognition,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness An audio-visual corpus for speech percep- tion and automatic speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.775662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:1ea7203a00715e81a3775a9ac43989cce1388f918157b684b894da3475c6a0b5

Observation fbcadae0-5547-446a-ae41-80573d54ffd5 · outbound

This paper cites Lrw-1000: A naturally-distributed large-scale benchmark for lip reading in the wild,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Lrw-1000: A naturally-distributed large-scale benchmark for lip reading in the wild,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.763500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:1f86484c910c0a499d4c322608a487594e447efae180c3ba53003fd9a9c73cb1

Observation ef0bb2c7-fc3e-47d3-9fea-b41cce30ebe5 · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Auto-avsr: Audio-visual speech recognition with automatic labels,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.781003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:69f1ed8f18034f34c6f55128e7c44d5d63e974059849fa5246062b037638d119

Observation e5bfc115-1220-4bb7-ad90-530d55611d8e · outbound

This paper cites The Multilingual TEDx Corpus for Speech Recognition and Translation,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness The Multilingual TEDx Corpus for Speech Recognition and Translation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.797065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:b057bb50bc47bec317ad6a1e472791b6ff53959e25af68647f6003d24aa0f1bc

Observation 9f1b07c5-e432-4a41-bc6a-073758491283 · outbound

This paper cites V oxCeleb2: Deep Speaker Recognition,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness V oxCeleb2: Deep Speaker Recognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.791707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:e32c566d0579299aa92a8b97efaf9516811caa18c9ff923093ec03ab45d0b24a

Observation 433e90eb-7505-44ee-851b-54890b9b5d12 · outbound

This paper cites Looking to listen at the cocktail party: A speaker- independent audio-visual model for speech separation,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Looking to listen at the cocktail party: A speaker- independent audio-visual model for speech separation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.789885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:c351fb24db68e1c6129547ede7266ed96cf738a3063189cd9d3e4481492d7b8b

Observation 7bd304dc-626b-4817-9e4b-c6bc2c805a04 · outbound

This paper cites PySceneDetect: Python and OpenCV-based scene cut/transition detection program & library,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness PySceneDetect: Python and OpenCV-based scene cut/transition detection program & library,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.770542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:02ac60a2c91141ee441652279946de6c3ab8e70aa02749d78f810bed4e752733

Observation 3733d10e-7a35-456f-add9-85a6c13c4bc6 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Arcface: Additive angular margin loss for deep face recognition,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.782691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:c9f7c145d475ce3adb517bebdd003df1c2182640bfda8763dfca318e2d84321b

Observation 5c936dd5-2659-4f4d-8793-6f7bd49a0df8 · outbound

This paper cites Sample and computation redistribution for efficient face detection,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Sample and computation redistribution for efficient face detection,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.784394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:2995b07083264caf8bbd68ad86917986feabbc2bcc0141b385786835f6872a14

Observation fb8fe471-e162-4516-8aaf-6382762f8d82 · outbound

This paper cites Pyannote. audio: neural building blocks for speaker diarization.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Pyannote. audio: neural building blocks for speaker diarization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.761708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:7869d41fceb25cddcfeeea53b0761b7b3f41f57d3b7d6bb577572fc58455090f

Observation 262cb23a-0737-47be-a251-081c5cd31191 · outbound

This paper cites S3fd: Single shot scale-invariant face detector,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness S3fd: Single shot scale-invariant face detector,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.773945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:7cc399165ff219744c8bde2d8c24882eb212177e38dcba0c14b470c6946e2bcc

Observation 23260569-3503-49f7-b3be-40addb677030 · outbound

This paper cites Out of time: automated lip sync in the wild,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Out of time: automated lip sync in the wild,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.765175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:f703226abfa69e93828f5e1eb5333d3c9179b77184cad1f4bd0ee3602df13738

Observation 58a25b60-d11e-427c-ae01-2bb11e063b23 · outbound

This paper cites Ro-n3ws: Enhancing generalization in low-resource asr with diverse romanian speech benchmarks,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Ro-n3ws: Enhancing generalization in low-resource asr with diverse romanian speech benchmarks,

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-10T12:57:07.602615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:f5f4021479ed14e5c1bc2f5647b340b3cd127d9a9eac3bb417514d5c2ea4171d

Observation 5bc25880-485f-4339-9d06-dcbf39c95e59 · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness MUSAN: A Music, Speech, and Noise Corpus

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.599591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:12b0235795d714a9fcdcbb83d87f884c70d43272eacd0a010239fc51f01932bc

Pith citing papers

No inbound Pith citation observations are available.