Pith. sign in

Paper Citation Record · LEDGER

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images?

As of 17 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2608.12515.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.12515 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:10:47.838698Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact3
  • verified fuzzy3
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f84693a-3983-4bb9-9493-1176a072a200 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:47.751344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:47.751344Z digest=sha256:c08b872cc65373cec1800cbd72098f851c9ec52bf92ac8bc59f0fa0ab984b5c7

Observation 91bb55f4-c132-40f3-9631-9805c1725261 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:47.756376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:47.756376Z digest=sha256:78e6c14bb94503ffb7ea80c8302ca811ae57d131cb8960649bf54c1dc63f72b7

Observation fd01c464-49ad-4740-b74e-c6d19ef6e493 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:47.761130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:47.761130Z digest=sha256:3f59be310190b29beacc1e35a66430b598bb96cb9207963664dc29f13ba3b0c7

Observation 56792bab-2105-450a-9192-bd4bc74d45e2 · outbound

This paper cites Micromachines12(2), 193 (2021).https://doi.org/10.3390/mi120201932.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Micromachines12(2), 193 (2021).https://doi.org/10.3390/mi120201932

Reference 4

Resolution
verified exact
doi, observed 2026-08-16T00:10:47.896706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:10:47.765749Z digest=sha256:620800b4b4e8c26cd90d1dc70c51cb7103a7cc3d7800ec042dc9b027558b6622

Observation 7f747b71-e98f-4b92-ad52-c258711fc215 · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? QLoRA: Efficient Finetuning of Quantized LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:47.770583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:47.770583Z digest=sha256:31b62a0419310b42b027f2a9a6fd88f22ac85f7f43e8bcc2ad96ee46ebb5f951

Observation 6734bad1-9772-485d-9013-3dc3e20aa0bb · outbound

This paper cites an unresolved cited work.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:10:48.423761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:10:47.775140Z digest=sha256:a726a943298a805f74ed58e6d653945a6b8d3d869f1ca48a544f641342915c77

Observation fe7216b3-dd71-4290-9565-75ffa964314f · outbound

This paper cites Chain of Thought Prompt Tuning in Vision Language Models.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Chain of Thought Prompt Tuning in Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:47.779785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:47.779785Z digest=sha256:20169c6e395a11dc82a2f6b563a0d717bc64e751decfbb0fecbc0d1d597cec23

Observation e893d4e0-8735-4a38-8f4d-80af8b076572 · outbound

This paper cites Doubleday (1966) 2, 4.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Doubleday (1966) 2, 4

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:48.409495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:10:47.784109Z digest=sha256:1bfc34d5a64651a3fae21081cba1269a4ab747d4278a737b2aa1f1af2411699d

Observation 06da91a4-b78d-46f4-b201-3cb33b33a71f · outbound

This paper cites co / HuggingFaceTB / SmolVLM-Instruct4.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? co / HuggingFaceTB / SmolVLM-Instruct4

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:48.394935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:10:47.788329Z digest=sha256:1d69f361bd13343db4c02ff3f5fa38ca0a1213a2dbc494fb0546e309d49064d2

Observation 7cbcc615-b1cb-4391-8001-e1b58afe864f · outbound

This paper cites Rudas and D.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Rudas and D

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:47.792298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:47.792298Z digest=sha256:95f6c11e9191a1be621524ea585c6cc7c41de27b88a6e473f0f066018c0a729e

Observation 1c836cc8-6a03-450a-a1d8-b74e4773046f · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Large Language Models are Zero-Shot Reasoners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:47.796759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:47.796759Z digest=sha256:6d639a114e4e2392a537ec78e7293f5610ce0a37103812a7105cd602bf35b708

Observation 69ef42a8-4a63-4abd-9fda-ab2b9c25cc26 · outbound

This paper cites The Format Tax.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? The Format Tax

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:47.801092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:47.801092Z digest=sha256:9c2ac0817949338da9cb7b96a721a3f9ba3c8b5ac22f1ea68814b6749caa8e57

Observation 763ee66d-bd2e-466d-a9b0-01512c4223f6 · outbound

This paper cites Visual Instruction Tuning.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Visual Instruction Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:47.805686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:47.805686Z digest=sha256:5bdf05cf6d05e9491265e4a9bf62f69198ee15d3a00daaf310979444fa9704c4

Observation 1c4ce6ef-8b05-4221-9717-c919f1d6cd52 · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence45(6), 6748–6765 (2023).https://doi.org/10.1109/TPAMI.2021.30705432, 3.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? IEEE Transactions on Pattern Analysis and Machine Intelligence45(6), 6748–6765 (2023).https://doi.org/10.1109/TPAMI.2021.30705432, 3

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:47.810156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:47.810156Z digest=sha256:9603cab954fd0eef0bd027a33a5be08841b07a83d167d892161284f58f307498

Observation d632bebc-06db-4f3f-9d35-07029c2f224c · outbound

This paper cites International Journal of Social Robotics12, 267–280 (2020).https://doi.org/10.1007/s12369-019-00560-92.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? International Journal of Social Robotics12, 267–280 (2020).https://doi.org/10.1007/s12369-019-00560-92

Reference 15

Resolution
verified exact
doi, observed 2026-08-16T00:10:47.873169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:10:47.814199Z digest=sha256:c6cba0b9e0eff8acd23a3c54e8c5404f181411d2f4aa06c4ac5a7bac511dc66d

Observation 1d32d408-ff84-473d-b97a-4b303ef6eeee · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Learning Transferable Visual Models From Natural Language Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:47.818214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:47.818214Z digest=sha256:b214740ef442cd0185d8ff2614c5d7e5d1cbc6d385629630a297659a0e809164

Observation 1b89d68b-4d61-4d7a-9fde-01b25e62ebb1 · outbound

This paper cites Electronics11(16), 2490 (2022).https://doi.org/10.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Electronics11(16), 2490 (2022).https://doi.org/10

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:48.380520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:10:47.822534Z digest=sha256:2fa1d71001564c8357520bc1d84f7658b15a21ca2c98383d1f7332ec27d11d24

Observation 56860c06-b79d-455b-bdfb-7ea7ad657f8f · outbound

This paper cites Attention Is All You Need.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Attention Is All You Need

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:47.826444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:47.826444Z digest=sha256:39021376229b9dddd12dfa3415eb42677c4d2995abdcb111554d485012ba712e

Observation ea39fe7c-823d-4ce6-a1ba-301e10558c9a · outbound

This paper cites an unresolved cited work.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Unresolved cited work

Reference 19

Resolution
verified exact
raw_fallback, observed 2026-08-16T00:10:48.140983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:10:47.830747Z digest=sha256:4f3d55c5289efcd4af5453b32acd2d8b822a691eb25b141a47759daee8398d22

Observation bcc04e46-f68f-4661-97cf-7e15880793a3 · outbound

This paper cites an unresolved cited work.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:47.834568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:47.834568Z digest=sha256:96a620b538ac48687c2052fef58017c128aec43e70ed14def77ef90f111fcefc

Observation e64bcc9b-1d90-4126-a0d2-2e9e9b24709e · outbound

This paper cites an unresolved cited work.

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:47.838698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:47.838698Z digest=sha256:8c459ec98da0760600871ea794f264952ce255ff73709b67a07c32dd4e42af76

Pith citing papers

No inbound Pith citation observations are available.