Pith. sign in

Paper Citation Record · LEDGER

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues

As of 22 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2607.10130.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.10130 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T14:04:19.786361Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c49d803a-166a-4aab-9ee8-05a883b6cbef · outbound

This paper cites Qwen2.5-VL Technical Report.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:42e30f7c999891f3a713f3f1230652c73e7d0ae1a4ddd2904ed69c7b4d6cd063

Observation dac3c4ff-f9f0-433c-aa03-151599c2cb42 · outbound

This paper cites Towards seamless interaction: Causal turn-level modeling of inter- active 3d conversational head dynamics.arXiv preprint arXiv:2512.15340,.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues Towards seamless interaction: Causal turn-level modeling of inter- active 3d conversational head dynamics.arXiv preprint arXiv:2512.15340,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:03da92d390dd6b668e93730f013c9c07afff899a6b7793b521d6160ea6c199c3

Observation 46ceed83-4a6f-4a54-af25-67775e1a8f9a · outbound

This paper cites What Does BERT Look At? An Analysis of BERT's Attention.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues What Does BERT Look At? An Analysis of BERT's Attention

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:1be191f59cf2fd1a56cc36196f4830009358bdefef30eaaf50d1c79a09e5e9e8

Observation dbac9967-66f6-44e4-91dd-7d8d5e9f3aa5 · outbound

This paper cites The Llama 3 Herd of Models.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:5cb758a625e4da3d24feac92713d31ffa5351ca8a15014f5cf501011615ffa81

Observation 17b60ef7-b2e9-4608-b22f-6d40e6a2db4b · outbound

This paper cites Ma-bench: Towards fine-grained micro-action understanding.arXiv preprint arXiv:2603.26586,.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues Ma-bench: Towards fine-grained micro-action understanding.arXiv preprint arXiv:2603.26586,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:210747735fa6213048458186c4e64d3205c821356e5c79265f89c689bb600f81

Observation aa8a4dec-f17a-400b-9b61-43fbf92eab9e · outbound

This paper cites Decoupled Weight Decay Regularization.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues Decoupled Weight Decay Regularization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:ff8a13d594975241457bca92b1886e35f57f84fd72070cc75a8efe336a36e5a1

Observation d443c431-9873-4406-98fd-c00c631786f4 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues DINOv2: Learning Robust Visual Features without Supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:3ef96108da553e61178a07a647f4bc8f29b09f480ed8ab837121994c7e1dff79

Observation de2e4452-50d5-4b5e-863f-d90f77253f8a · outbound

This paper cites AI Governance and Accountability: An Analysis of Anthropic's Claude.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues AI Governance and Accountability: An Analysis of Anthropic's Claude

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:db4fba2572a86a7f483f0ebd0e9de3cbdf9b5505771963aa139edc05fb144109

Observation 18993df2-4581-4a83-9d36-02e750f91b0b · outbound

This paper cites Sharingan: A transformer architecture for multi-person gaze following.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues Sharingan: A transformer architecture for multi-person gaze following

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:a4eb3974c46cc28c74c15b084fbfaad652edc02a9e8949a1b8f9b2f00f9e5bb8

Observation 2d67d03b-2c92-4c5a-85e7-63e13160a39f · outbound

This paper cites Multi- modal across domains gaze target detection.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues Multi- modal across domains gaze target detection

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:26de6b83a600e791e05c8bb2e9dbc368c9d53296e53ecb0146f984c8db2825d1

Observation faf481f8-29bf-4153-a006-b5f95a7a99aa · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues Emu3: Next-Token Prediction is All You Need

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:c94a8180a2d9d57def9396029f62eac01610caf2c1a6e512a707ee6cd5c2b7c8

Observation af2e02a9-1adf-40f2-8bcb-c694c5958f86 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:635265b5edb01d07cec8af5bf278eac443651dde7b072d92c8f49fde5e488daa

Observation bbbd632a-a2ca-4ad0-86d9-7785970b71fe · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:15eb6a4b128f8a693d41b2d100ef7fd5ceb4024a38ae67af05f1e8b266adebe2

Pith citing papers

No inbound Pith citation observations are available.