Pith. sign in

Paper Citation Record · LEDGER

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech

As of 11 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2412.11409.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11409 v3

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:01:33.912139Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:34:07.907336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T04:34:07.970750Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c705524e-f197-4a1d-939c-0c3fa5483dc6 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.739038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.739038Z digest=sha256:53ed0f8c64099cf1bc6afc19dd8c66b74c46fc7fe21a99f091451af61d21f73b

Observation 9e631196-c032-4965-bc13-04bbe0c34f52 · outbound

This paper cites write newline.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.744102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.744102Z digest=sha256:02bc014f1efc6ca8d7b5a762871cdec3afc387c7911f58e57a22f8fa05a78418

Observation 5317d04d-bc0d-425d-83a1-b2e1a780c7e3 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.462181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.749585Z digest=sha256:575e40c0a72d8fb2396a6825d85ba7ee1e42cbe5426bc02f540986728e2b13ca

Observation ffb4a67a-d664-4158-8575-be662b556a44 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.446298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.754047Z digest=sha256:cc902882274ac82d121b8652485d9df4d664a38a32c178875963a3dd6e185cbc

Observation acb71e47-d6c8-4f24-bbca-d13450814309 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.431945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.759161Z digest=sha256:79fc7a8e57c3e0282230eb3563039a76de91c483a10e1c3fe90de32274b16876

Observation a79ebfa7-0d48-4c24-a165-d8486079fca1 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.416672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.763741Z digest=sha256:a5a932359ca4aa4c3e1d53615249bfcb5c36031aed2974f1fdcad5dd10a07dfc

Observation 67654fb8-28d2-4741-824e-0d5c02c4aead · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.399707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.767991Z digest=sha256:d1ea4b5c673d02ea5f8daa876dc7457107e624878c54cfc6d91e0695aec0de82

Observation daace53f-7531-49e7-a731-8ac77f4de9ac · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.384054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.772569Z digest=sha256:f3700c2df41e31e66606b8587a3b5722baf900c0cce0f82fa0f930a020d74b97

Observation 6afcc842-ea5b-4c2d-8d9f-943ec20cf967 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.776974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.776974Z digest=sha256:28dc62ad927721bf73b80120a89f65bc84c11f87645d77bb280a697a168aff54

Observation 1b34d21b-ea50-411c-9aee-745b07d2dab8 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.780829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.780829Z digest=sha256:4ef67fc3201e1d5f9de0e1b9574e68ca6f1b5d824aed5d1573e34c9e92208ad5

Observation e9b42dad-0124-4f16-8954-0d31983c5599 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.358773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.786306Z digest=sha256:1848272467d22dde66db14c4b0eb3c24f4881a1169bb2080b6fd582d830b7c67

Observation 9ff317ad-9929-4402-a93f-f0be088005fe · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.791200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.791200Z digest=sha256:e096d34ce31d76f4b7cd8ff2486bed18d36e073e3fa0a60803497c94ce373ce1

Observation 5ff458eb-6d03-4358-9d4f-977c5d726824 · outbound

This paper cites Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.795919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.795919Z digest=sha256:858499ba9e69ef4093b24a2c5e400d4ffa99dfbf882670a00588d6a959ce8eb4

Observation 3947395a-d4de-42a0-bf82-50dbe3d93928 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.334532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.800794Z digest=sha256:1b9ed95e1ba127893b176470455605b1ede7bd7e54976b0524e1209d3bf63770

Observation 2001f691-e00a-460b-b8fe-fb14d7720ec7 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.320257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.805342Z digest=sha256:436e9f0fac5816bbf47ab83fa2c2d6f0a6dd222cae22ca1d9bfb9df14b0569a3

Observation 1be592c6-f85d-4895-89d2-5d38ccae6684 · outbound

This paper cites I Want to Figure Things Out.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech I Want to Figure Things Out

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:01:34.306395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.810025Z digest=sha256:2c9248bf690a3b9f8c3eabb3852f9bf674d1c73dc9eaad7973bbfef03cf9c753

Observation 9574b814-ab1e-4dec-b20f-7b5d18629c11 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.292349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.814503Z digest=sha256:c7042a604d68a7a987a684300e313b215bffbd2cb615da006e0df15393c036e1

Observation dd355004-bb8f-4cbb-8192-a963e173c590 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.279446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.819181Z digest=sha256:c61e427b3cdffab663dda893f660da83e474c02c76655906190c3cc5bfae1146

Observation 40ae4745-caec-445a-a88f-6ae1dd587d2c · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.823743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.823743Z digest=sha256:66ca2d36a440a5d013ad9270055b39cb3f6c62aad15bcd80f1091d4bcd559b0a

Observation c0978bca-366d-49ac-801b-55b80b065418 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.266073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.829016Z digest=sha256:bc2a54086f0fe167a3b6e045b3574b2c4e7b5ad2c5636ea569f35bccf28a93ea

Observation 70e70d2c-3577-409f-8edb-c6d83fec4a91 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.252060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.833955Z digest=sha256:2eff1a22ce7188dc146f7bbf72261ae6488596f01e15e9b2e9ec625c551f0d34

Observation 69526ea0-acd6-4c1a-b90b-a65febcd5a91 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.237380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.838309Z digest=sha256:9326b5ce9a639c475b863f56e6bd1c0ecd7bd167380499a79973e753d51f461b

Observation 7a359cd9-a89b-4151-b592-2b6075e3e73b · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.221868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.842440Z digest=sha256:d693da6b361b814c7ec7d78fdd6c32e63d223e5e951428f14c883bb65403b3d3

Observation 95b7c3d1-10b7-4e3d-9774-a5ee7c04d747 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.206779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.846787Z digest=sha256:cfe0a97f01718aea76ba240d31ba4bb52237b891a9a38cc1be1a2a89bbf55c33

Observation 55ec6a08-cbb9-4afb-bae2-853d5c06997c · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.191700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.851111Z digest=sha256:ab0a4dd264d9118f9ae60fb6ec2b4f4d1c6c20daa775b4781e1d5b11e1c72a6b

Observation 0b52ec4d-6e79-4e64-b88e-82f8156726f0 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.176506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.855330Z digest=sha256:e7e5829800f693bfd9577085a1e14189bf9bccbc22631a45c63d93ef5ec768f1

Observation 4ed2f487-db09-40f8-856e-fd0699c39ed2 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.162005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.859348Z digest=sha256:37e088a199fe86078d8c4260a40318deeb4c5f6c89e18209472126aa08ef9cb4

Observation b48ec891-ff5a-4114-a4ff-f89f1b5c907f · outbound

This paper cites W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.864813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.864813Z digest=sha256:1aa03ea06d3c4e1d65bd8cfa1b1407b2340cd6d7dc8bc460a89aa616706084c3

Observation 1e94e569-b79c-4f41-b772-23178a59a521 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.138379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.868936Z digest=sha256:2f6f09ed0da0907d8354840772ccb469a7ffa3279078134d97a15e06a16d2aaf

Observation 640690fe-27a4-4af8-98ac-b369067a71bd · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.125191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.872844Z digest=sha256:01460cbcebc5d72b27a251498b83f1c2732382a7ba25e83a2da2eb8c08b65091

Observation 79c5ca32-baf5-4c51-a7dd-bb5b881743ab · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.112109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.878274Z digest=sha256:f5541663d8dfe067b7214005b4edf84ad3c287fceb730de95b591bfff8f257e5

Observation 0c14f8f4-2d1c-4f34-80e8-7f37a6a47d4e · outbound

This paper cites N.; Tran, S.; Yao, B.; Chilimbi, T.; and Shah, M.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech N.; Tran, S.; Yao, B.; Chilimbi, T.; and Shah, M

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:01:34.097133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.882704Z digest=sha256:d5315fcfd345e771301aacb3aea702cd996f0e01d19e2817d66c2bbae21a18bf

Observation 38206b27-7f94-472a-9891-5a092783cb75 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.081238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.887301Z digest=sha256:7b9857788b5ae07324f9714383f049d9f056312d89ec2882979ba19e0015e6d2

Observation 1b240567-7cb5-4136-b312-0db0ea272faa · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.891600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.891600Z digest=sha256:35616ed60529055f9c9c3591a540381139e6f4df6299863026b69daf10e27d1d

Observation 17aeb1c9-3e11-4734-853c-aad21ad18bd7 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.066593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.897141Z digest=sha256:9ccf26633a7f6a6306d56b426c44648c59150ecf7e88cdb2e57c722d46ae2d0d

Observation 761eacb8-d295-4c79-bb62-0f9c58bee78e · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.052267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.902340Z digest=sha256:377581e2b20ab164f21e64b373382e830b7d642675ae86497de2198c873d1503

Observation bd7e2edf-23ba-465b-9cbc-7bd953f8cfcb · outbound

This paper cites C.; and YAN, S.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech C.; and YAN, S

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:01:34.036847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.907180Z digest=sha256:ab53f105fb561bcea52a5c2adbc3915fdc2c418c19294717cd23d698b6055f1f

Observation 01a873bd-32f7-49ba-9806-04828d57aec4 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.020850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.912139Z digest=sha256:6e47a7efa8350ff939fa271730c6716d19122d4c8dac5717fecbb440712c5a85

Pith citing papers

Observation 6ece83f0-cdae-48d8-b707-a9faaa591b34 · inbound

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction cites this paper.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:07.977433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:34:07.907336Z digest=sha256:717863789327bd7b5546d5f70b401ae1370b94aef687fc514e3b363e0c418fa4