Pith. sign in

Paper Citation Record · LEDGER

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues

As of 5 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2607.10130.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.10130 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T14:04:19.786361Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c49d803a-166a-4aab-9ee8-05a883b6cbef · outbound

This paper cites Qwen2.5-VL Technical Report.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:1c34a134d059b8778965c98ca8edf0b0b2c22637af38832a76b99e79d3bf12b4

Observation dac3c4ff-f9f0-433c-aa03-151599c2cb42 · outbound

This paper cites Towards seamless interaction: Causal turn-level modeling of inter- active 3d conversational head dynamics.arXiv preprint arXiv:2512.15340,.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues Towards seamless interaction: Causal turn-level modeling of inter- active 3d conversational head dynamics.arXiv preprint arXiv:2512.15340,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:a2891e3a2fc2ecafcb64a81b48edc9778441f481658d8e59ed0e53f7eab49366

Observation 46ceed83-4a6f-4a54-af25-67775e1a8f9a · outbound

This paper cites What Does BERT Look At? An Analysis of BERT's Attention.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues What Does BERT Look At? An Analysis of BERT's Attention

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:2ff76ada7beb8e9dc2c904123a8f3e0cf35cce22ff1293731afea65b2a6e96a9

Observation dbac9967-66f6-44e4-91dd-7d8d5e9f3aa5 · outbound

This paper cites The Llama 3 Herd of Models.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:15ef193a2fa9633cac7bc542e0f3ea1639f33aca643ed8e0b291a582a05d9b44

Observation 17b60ef7-b2e9-4608-b22f-6d40e6a2db4b · outbound

This paper cites Ma-bench: Towards fine-grained micro-action understanding.arXiv preprint arXiv:2603.26586,.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues Ma-bench: Towards fine-grained micro-action understanding.arXiv preprint arXiv:2603.26586,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:da71c596f62cfd32cca354865b711091c54e52a3dca72a958bf7081f0475c1cb

Observation aa8a4dec-f17a-400b-9b61-43fbf92eab9e · outbound

This paper cites Decoupled Weight Decay Regularization.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues Decoupled Weight Decay Regularization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:83e8c3a3f93f76e5ddbe28b6a517839e1b8489e482ec01400365527df1247913

Observation d443c431-9873-4406-98fd-c00c631786f4 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues DINOv2: Learning Robust Visual Features without Supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:421b6ae7961a8dc86adadcc6aac8640fa36e2c87d4543344f0d41ef23437408f

Observation de2e4452-50d5-4b5e-863f-d90f77253f8a · outbound

This paper cites AI Governance and Accountability: An Analysis of Anthropic's Claude.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues AI Governance and Accountability: An Analysis of Anthropic's Claude

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:f04259aa69a0bd194924959658d2f2b88b57b96dfcdc3043d9e5d5b70ee5b353

Observation 18993df2-4581-4a83-9d36-02e750f91b0b · outbound

This paper cites Sharingan: A transformer architecture for multi-person gaze following.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues Sharingan: A transformer architecture for multi-person gaze following

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:76b3a1a3c39e940206b9a841116b940497e67e4d9fd8c9f754ad3831870539a4

Observation 2d67d03b-2c92-4c5a-85e7-63e13160a39f · outbound

This paper cites Multi- modal across domains gaze target detection.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues Multi- modal across domains gaze target detection

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:187b9f8392ded65456dd9312a1b46499d4faa7552c056c80ecba7f9025db2d59

Observation faf481f8-29bf-4153-a006-b5f95a7a99aa · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues Emu3: Next-Token Prediction is All You Need

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:0ebe3462a9c7b35c995fece0609f57385a0c208429c1b859bc75165c307311e3

Observation af2e02a9-1adf-40f2-8bcb-c694c5958f86 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:921937b170921291f6a1ed81c491b603d4847cfbfae8ae403d19d68132ce8af5

Observation bbbd632a-a2ca-4ad0-86d9-7785970b71fe · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T14:04:19.786361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:04:19.786361Z digest=sha256:27fa932a86f36437154abadfa57bdeb52e9a9113ab3157e87c03f9bc11609267

Pith citing papers

No inbound Pith citation observations are available.