Pith. sign in

Paper Citation Record · LEDGER

Multi-Grained Spatio-temporal Modeling for Lip-reading

As of 18 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:1908.11618.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.11618 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T10:13:16.605795Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1fb123a-b994-4db9-a1cf-6d279826eff5 · outbound

This paper cites Improved speaker independent lipreading using speaker adaptive training and deep neural networks.

Multi-Grained Spatio-temporal Modeling for Lip-reading Improved speaker independent lipreading using speaker adaptive training and deep neural networks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.342047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.330178Z digest=sha256:7f23495e550a0e5c4f4eb8fefce392b5cb031325afca9bb1f655d39e8f179ff8

Observation d10a5ed7-5bf3-4b64-be67-3f884ee3769a · outbound

This paper cites LipNet: End-to-End Sentence-level Lipreading.

Multi-Grained Spatio-temporal Modeling for Lip-reading LipNet: End-to-End Sentence-level Lipreading

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T10:13:16.335704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:13:16.335704Z digest=sha256:8cd558dc371cb54de5df0b77d811d895e8dbb2eaa8fd688844e049fc2351c6ee

Observation 2c8c86ec-206a-47fb-bc35-314d54824ab0 · outbound

This paper cites The natural statistics of audiovisual speech.

Multi-Grained Spatio-temporal Modeling for Lip-reading The natural statistics of audiovisual speech

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.322733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.344748Z digest=sha256:af887bc0a3cf99b0e3cb7310fe08fb401bb6a1e5a3827a24eb18052d4f94be08

Observation de53de35-197b-4c63-8a7d-5751cf81757a · outbound

This paper cites Lipreading from color video.

Multi-Grained Spatio-temporal Modeling for Lip-reading Lipreading from color video

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.296347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.350281Z digest=sha256:301278d65ea80bcf400707a4d23141fed735b661d2c878a3bf149c4e3e0c2340

Observation eb7f2d26-0ca4-4fd3-8b24-4795f649d1c4 · outbound

This paper cites Lip reading in the wild.

Multi-Grained Spatio-temporal Modeling for Lip-reading Lip reading in the wild

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.274433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.356484Z digest=sha256:27294fb2940cdd373f0658242ae57b17056c6de1b3d78b2f7f4224ee48a959cd

Observation 9000b87c-7555-479f-968c-6449428a50ca · outbound

This paper cites Learning to lip read words by watching videos.

Multi-Grained Spatio-temporal Modeling for Lip-reading Learning to lip read words by watching videos

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.257495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.364587Z digest=sha256:0cbc1e24394c3cdcfaa34f65e65494ea4c70f0e486c21eefae53c9261dac11dc

Observation 0c9c0936-ef40-4e83-bb45-22f82e3bfaa0 · outbound

This paper cites Lip reading sentences in the wild.

Multi-Grained Spatio-temporal Modeling for Lip-reading Lip reading sentences in the wild

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.235879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.379468Z digest=sha256:b67f05b82bfbf28bc838539de13bbe4be78cdaed7c2d8789502074377582a88c

Observation 2ddffe38-9e90-4bc1-a26d-c7ac4cb618c1 · outbound

This paper cites Toward movement-invariant automatic lip-reading and speech recognition.

Multi-Grained Spatio-temporal Modeling for Lip-reading Toward movement-invariant automatic lip-reading and speech recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.214393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.389161Z digest=sha256:f38585ef14da486dde056af00dcca858f49c55b8440302ffb3f7e2ecefa7f3df

Observation 8b9c4c94-bd7f-4e58-b953-fa399ad559db · outbound

This paper cites Long short-term memory.

Multi-Grained Spatio-temporal Modeling for Lip-reading Long short-term memory

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.195559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.400542Z digest=sha256:11fae8b625fa2c59e30e32fba0f9d9189050a9f3dcdb2c8f11642b502f62ebad

Observation 9b155bce-615d-4481-b1ab-5d229fe6e8d1 · outbound

This paper cites Videolstm convolves, attends and flows for action recognition.

Multi-Grained Spatio-temporal Modeling for Lip-reading Videolstm convolves, attends and flows for action recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.175556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.410492Z digest=sha256:bc4707fbedadcb2db13ab19bca6d8b5438a01779f04eebcc26b93a5c62d5a076

Observation d19605dc-49ba-43c6-b0b8-7fbccaffdc84 · outbound

This paper cites Audio-visual speech recognition using deep learning.

Multi-Grained Spatio-temporal Modeling for Lip-reading Audio-visual speech recognition using deep learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.152173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.428574Z digest=sha256:df9627fb0495949132031965c323db15eeeb74503a3be88e2e4a5ed340bc7f2a

Observation 430d9fb0-6be7-45d2-b500-6afb1d0bcfdf · outbound

This paper cites End-to-end audiovisual speech recognition.

Multi-Grained Spatio-temporal Modeling for Lip-reading End-to-end audiovisual speech recognition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.126325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.442214Z digest=sha256:7a53ac8c7687773f3d276f051f1a6b0f97272db7b631ee27f6e0631d46262c4a

Observation 3cca9545-f18b-461f-950d-d45d147dd8f0 · outbound

This paper cites An image transform ap- proach for hmm based automatic lipreading.

Multi-Grained Spatio-temporal Modeling for Lip-reading An image transform ap- proach for hmm based automatic lipreading

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.100644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.454167Z digest=sha256:90b330b7806209f1fa6f38035bbc0744a5a30cfbc6b694363f87c53a0077382c

Observation 55ce54b8-4f7a-4136-9cd7-ace573e3b85f · outbound

This paper cites Recent advances in the automatic recognition of audiovisual speech.

Multi-Grained Spatio-temporal Modeling for Lip-reading Recent advances in the automatic recognition of audiovisual speech

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.076930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.470890Z digest=sha256:fbe912d5f917a9eebe1987b915706ece022fd23c5668ff25f820cfdc79928c81

Observation f1760879-f233-41ba-aa87-02b2c93af2cb · outbound

This paper cites Lip reading using optical flow and support vector machines.

Multi-Grained Spatio-temporal Modeling for Lip-reading Lip reading using optical flow and support vector machines

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.044818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.479365Z digest=sha256:73ef3ec6869ffc1c8c4f099291a924361eac8020ec57ad9c21ef4af887f7dcdd

Observation f94d5605-5501-4d93-882f-a11482410d74 · outbound

This paper cites Convolutional lstm network: A machine learning approach for pre- cipitation nowcasting.

Multi-Grained Spatio-temporal Modeling for Lip-reading Convolutional lstm network: A machine learning approach for pre- cipitation nowcasting

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.016990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.490705Z digest=sha256:98773b291002da5c037bc534ce0258d149f44a47e222e2d052bd327eedc42180

Observation 49763019-c3b6-4b02-892b-9d8d8c2749d1 · outbound

This paper cites Two-stream convolutional networks for ac- tion recognition in videos.

Multi-Grained Spatio-temporal Modeling for Lip-reading Two-stream convolutional networks for ac- tion recognition in videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T10:13:16.505615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:13:16.505615Z digest=sha256:e432944dce0a71789c999aeac8e395e48a482d524d3e85225c2d4c9bf222f6fb

Observation 693dd69f-c76b-4e51-832c-a69fa585f26f · outbound

This paper cites Combining Residual Networks with LSTMs for Lipreading.

Multi-Grained Spatio-temporal Modeling for Lip-reading Combining Residual Networks with LSTMs for Lipreading

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T10:13:16.519177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:13:16.519177Z digest=sha256:486bd4c9ec246cc6900c67b2fb5663dc68726dd971b4e86893897ca8d04e4871

Observation bd126973-8fec-48e2-b45e-918140376801 · outbound

This paper cites Convolutional long short-term memory networks for recognizing first person interactions.

Multi-Grained Spatio-temporal Modeling for Lip-reading Convolutional long short-term memory networks for recognizing first person interactions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.961031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.526459Z digest=sha256:aff231e12042fa3fb3a61ff6d77739f4e635fb2e7d36e5321c2099aa4d11fcb8

Observation 994ea864-7e38-4e17-9fcb-6afa9de28a89 · outbound

This paper cites Improving lip-reading performance for robust audiovisual speech recognition using dnns.

Multi-Grained Spatio-temporal Modeling for Lip-reading Improving lip-reading performance for robust audiovisual speech recognition using dnns

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.937044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.536979Z digest=sha256:3c0bd58be4bb25434508877b3fb4da7c89934950fca705eb2b194c0789d35cd9

Observation 604defd1-f3de-4f81-a1c6-cf891aee9a19 · outbound

This paper cites Hu- man action recognition by learning spatio-temporal features with deep neural networks.

Multi-Grained Spatio-temporal Modeling for Lip-reading Hu- man action recognition by learning spatio-temporal features with deep neural networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.905106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.542441Z digest=sha256:8a64d53ca5e420f165e7891506c5b6849da8c744ab1fc9fb5888484471be7f6f

Observation 20992fce-8a4c-495f-bc42-f9a5cf4453e1 · outbound

This paper cites Pre- drnn: Recurrent neural networks for predictive learning using spatiotemporal lstms.

Multi-Grained Spatio-temporal Modeling for Lip-reading Pre- drnn: Recurrent neural networks for predictive learning using spatiotemporal lstms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.876930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.552535Z digest=sha256:31081bd8f84aa38924e349479dca47b5a0025153c87fa3949b34e156bedde79e

Observation 44855688-7a4f-4c2b-9651-06b751e2e4f1 · outbound

This paper cites Lrw-1000: A naturally-distributed large-scale benchmark for lip reading in the wild.

Multi-Grained Spatio-temporal Modeling for Lip-reading Lrw-1000: A naturally-distributed large-scale benchmark for lip reading in the wild

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.856778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.564273Z digest=sha256:ad0c9499f077f7ab9055ee8ec058114300a19656fa106a83eb211b02546c0fda

Observation a62bb90f-2ddb-45aa-a245-6e0022ff9150 · outbound

This paper cites Learning spatiotemporal features using 3dcnn and convolutional lstm for gesture recognition.

Multi-Grained Spatio-temporal Modeling for Lip-reading Learning spatiotemporal features using 3dcnn and convolutional lstm for gesture recognition

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.839284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.576389Z digest=sha256:fb347a2154a7d5f81fa224bbdf740316fdf9e8b6358a2416b148876810c8cc90

Observation c3836fe0-acb8-4f34-9fd1-e27c4d14bfdd · outbound

This paper cites Adding attentiveness to the neurons in recurrent neural networks.

Multi-Grained Spatio-temporal Modeling for Lip-reading Adding attentiveness to the neurons in recurrent neural networks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.812779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.584649Z digest=sha256:ace33ab14cf1c7ea3e194c30ea43c048c5635b3c6b44bd23025c100c72c60755

Observation 7f9febc2-6774-4a98-8fde-8869dee38b18 · outbound

This paper cites Lipreading with local spatiotem- poral descriptors.

Multi-Grained Spatio-temporal Modeling for Lip-reading Lipreading with local spatiotem- poral descriptors

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.781772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.593574Z digest=sha256:83fbb6c21295e6b6f3557ba18aa0d93b8902774912176eb6d9266c759982456e

Observation 3b50c32d-886b-4b2b-bd68-df5501fe7d0f · outbound

This paper cites Multimodal gesture recog- nition using 3-d convolution and convolutional lstm.IEEE Access, 5:4517–4524, 2017.

Multi-Grained Spatio-temporal Modeling for Lip-reading Multimodal gesture recog- nition using 3-d convolution and convolutional lstm.IEEE Access, 5:4517–4524, 2017

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.753235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T10:13:16.605795Z digest=sha256:14e52d539316ce95bcbb19326d9f57f2ccfb707ef20725c5b76592005a8e1606

Pith citing papers

No inbound Pith citation observations are available.