Pith. sign in

Paper Citation Record · LEDGER

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness

As of 13 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2607.08112.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08112 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T12:53:22.119074Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact5
  • verified fuzzy24
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77885fbd-0945-44e7-86bc-be2bcc157ec9 · outbound

This paper cites LipNet: End-to-End Sentence-level Lipreading.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness LipNet: End-to-End Sentence-level Lipreading

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.609021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:29a56ad97e54a0a29938e265b8c9ca4ef259a58e845ce934e9e2c36dbf55deff

Observation b74b966e-d759-4cf7-ad9e-5132fcc1e64d · outbound

This paper cites Lip reading in the wild.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Lip reading in the wild

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.798938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:220b939b374421c3702ab9443250a9ba2c8f30d9a108fbdd29d4e860a8ce8600

Observation 529f548e-fc6d-4dd8-999c-290708cc5386 · outbound

This paper cites Lip reading sentences in the wild.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Lip reading sentences in the wild

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.766741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:93df79ee7b9aa38ce0518a4dab31b53ee9b5ddd17b16c94d35249aa09371154d

Observation 42e8e4b9-097d-493d-9bdf-5f18b0fc8fd5 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness LRS3-TED: a large-scale dataset for visual speech recognition

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.605649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:adc8c99683349afa031c655369a6b9ba28eff1117d46c2445dda8776dd52f19e

Observation 79c96a54-7300-4bee-b00f-38a6e5f6d1d9 · outbound

This paper cites End-to-end audio- visual speech recognition,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness End-to-end audio- visual speech recognition,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.793434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:e15cf2e9a92bee732f5974d2329736af9fba5a829909e1c87e01d2b180732288

Observation 7bcee6f3-6794-470b-9331-e67a34ba5fbc · outbound

This paper cites Lipreading using temporal convolutional networks,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Lipreading using temporal convolutional networks,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.788000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:95aa7f66dabf5cbd1592ed200086ef774056ebc2d7627b5f5138997073646f7b

Observation bd7ccf15-1558-4d13-aac9-b69e475421bb · outbound

This paper cites Training strategies for improved lip- reading,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Training strategies for improved lip- reading,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.795076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:5b7b0cd54c8dd45391ed9798f398d45bede30caea7e58eac7775f59edc7b7dff

Observation 7c97ae26-d06e-4bd5-a39d-547fc6051e06 · outbound

This paper cites A multimodal german dataset for automatic lip reading systems and transfer learning,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness A multimodal german dataset for automatic lip reading systems and transfer learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.779253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:73127d15e5d0404f88c5e97992bed103c0bacde2e5c1a07f775b660b1507a276

Observation 8432314e-b67e-41bc-8f0f-eb825ce1a1cc · outbound

This paper cites Visual Speech Recognition for Multiple Languages in the Wild,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Visual Speech Recognition for Multiple Languages in the Wild,

Reference 9

Resolution
verified exact
doi, observed 2026-07-10T12:57:07.523625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:e2413a4ea2faca4630646fcdbd240d40460d04a80c1cfb591257fe8b878bf0c3

Observation e6ff0fab-f406-4880-8956-bc472252b2e4 · outbound

This paper cites Scaling multilingual visual speech recognition,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Scaling multilingual visual speech recognition,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.786234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:fafc227ecfbf9e85ed7705f901197d404536cbfcaee70e8f62fe407b8517ca98

Observation abd173d5-ed40-43b3-a6b6-279c9b085549 · outbound

This paper cites Visual speech recognition for languages with limited labeled data using automatic labels from whisper,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Visual speech recognition for languages with limited labeled data using automatic labels from whisper,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.777454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:ed17cfba295a6257ee5a03d08fd1e09dcffe0dc9f4bf459c86863aab9104d4a2

Observation ceffa10a-dd0f-4d26-8f47-471fbe63cd37 · outbound

This paper cites Lrro: a lip reading data set for the under-resourced romanian language,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Lrro: a lip reading data set for the under-resourced romanian language,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.759395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:bba8d870a402e8fad66896d385357b8a5063ef851cd6389e7c47cd532e6a095b

Observation f4714316-4d92-419d-a0f8-f6eff61246ad · outbound

This paper cites Toward language-independent lip reading: A transfer learning approach,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Toward language-independent lip reading: A transfer learning approach,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.800840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:0edbdeee8774d08a6cd79d04e5b8ea08804186f2037ab131eb701fc6b90c13e6

Observation 4bfc1d75-93ba-4c70-93e5-5097e924385b · outbound

This paper cites End-to-end lip reading in romanian with cross-lingual domain adaptation and lateral inhibition,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness End-to-end lip reading in romanian with cross-lingual domain adaptation and lateral inhibition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.768627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:f7a98e14e92457982f347a74b99e454eb27d573b64fea461aca261a0d55f1426

Observation 3c8376d8-ff51-41ea-b8b4-869f1e6e7bef · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Robust speech recognition via large-scale weak supervision,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.772266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:689231722b5c8259cde2e60927dd5d81dda3ce646a835509e999e339943caaa7

Observation 0de941e6-3fcc-438a-9517-7541fe6a68d5 · outbound

This paper cites An audio-visual corpus for speech percep- tion and automatic speech recognition,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness An audio-visual corpus for speech percep- tion and automatic speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.775662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:245ae2c5fa65b5112c4d453ff6c763ed103c34c7b2bd18a34e2f6759762698d1

Observation fbcadae0-5547-446a-ae41-80573d54ffd5 · outbound

This paper cites Lrw-1000: A naturally-distributed large-scale benchmark for lip reading in the wild,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Lrw-1000: A naturally-distributed large-scale benchmark for lip reading in the wild,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.763500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:80ca34ff825d8c3dc947a7ba13bc613cf85095288a9666a385d6f94c9b57c1d6

Observation ef0bb2c7-fc3e-47d3-9fea-b41cce30ebe5 · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Auto-avsr: Audio-visual speech recognition with automatic labels,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.781003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:5cb74d2dfabe472793e1e00ccb4b6db949093e2a9827a5588b3d684db122db73

Observation e5bfc115-1220-4bb7-ad90-530d55611d8e · outbound

This paper cites The Multilingual TEDx Corpus for Speech Recognition and Translation,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness The Multilingual TEDx Corpus for Speech Recognition and Translation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.797065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:d86e9686faa93d7975152032f9fc1678cd3d92f9366e330b0d9442c605d340af

Observation 9f1b07c5-e432-4a41-bc6a-073758491283 · outbound

This paper cites V oxCeleb2: Deep Speaker Recognition,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness V oxCeleb2: Deep Speaker Recognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.791707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:bcee7a766ecab9de12b2df3c8b6618d1f613a635b52edbea4ee191b2ef4eed04

Observation 433e90eb-7505-44ee-851b-54890b9b5d12 · outbound

This paper cites Looking to listen at the cocktail party: A speaker- independent audio-visual model for speech separation,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Looking to listen at the cocktail party: A speaker- independent audio-visual model for speech separation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.789885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:ea59fa57afd3178583e0bb29e5d823369f8bf53174fa3c1528e7bed468a85245

Observation 7bd304dc-626b-4817-9e4b-c6bc2c805a04 · outbound

This paper cites PySceneDetect: Python and OpenCV-based scene cut/transition detection program & library,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness PySceneDetect: Python and OpenCV-based scene cut/transition detection program & library,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.770542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:d9d7cb1889f667fe1044f5124910c2d20cc59bd366a2efc71537e715f3d4141f

Observation 3733d10e-7a35-456f-add9-85a6c13c4bc6 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Arcface: Additive angular margin loss for deep face recognition,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.782691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:b18a52f47f7d4cec044ecda2f8b36701607927521d233b4f4476a0449e77ae47

Observation 5c936dd5-2659-4f4d-8793-6f7bd49a0df8 · outbound

This paper cites Sample and computation redistribution for efficient face detection,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Sample and computation redistribution for efficient face detection,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.784394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:73b0f05f5297d3e4789b6c488eeb3287e296feba8088dc646f0513a2bd629027

Observation fb8fe471-e162-4516-8aaf-6382762f8d82 · outbound

This paper cites Pyannote. audio: neural building blocks for speaker diarization.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Pyannote. audio: neural building blocks for speaker diarization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.761708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:2d6ed3d56b22a53fa2c045b1549e83e214abf4954fa74b10e635742706026464

Observation 262cb23a-0737-47be-a251-081c5cd31191 · outbound

This paper cites S3fd: Single shot scale-invariant face detector,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness S3fd: Single shot scale-invariant face detector,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.773945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:34bf7de1ddc2b49c62ec658d0e3e1bbb121f4cae87529feebf1be6c082dd11cf

Observation 23260569-3503-49f7-b3be-40addb677030 · outbound

This paper cites Out of time: automated lip sync in the wild,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Out of time: automated lip sync in the wild,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T12:57:07.765175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:f82c431d852918a3ccfc881eba3ee08d4d82e2de3645e657528fb2d27ae9791f

Observation 58a25b60-d11e-427c-ae01-2bb11e063b23 · outbound

This paper cites Ro-n3ws: Enhancing generalization in low-resource asr with diverse romanian speech benchmarks,.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness Ro-n3ws: Enhancing generalization in low-resource asr with diverse romanian speech benchmarks,

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-10T12:57:07.602615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:5a51134bc08d1d1eef71ebf6fc3c8445b2aa25b63b4e732eb79d7d16035cd986

Observation 5bc25880-485f-4339-9d06-dcbf39c95e59 · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness MUSAN: A Music, Speech, and Noise Corpus

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:57:07.599591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T12:53:22.119074Z digest=sha256:01c3c666f2004dfda52ce4956b8676c00558902bf2b1b05015afff882871f4be

Pith citing papers

No inbound Pith citation observations are available.