Pith. sign in

Paper Citation Record · LEDGER

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions

As of 21 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2505.22627.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22627 v2

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:09:57.563270Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db4840ac-d4df-4a80-9360-edf369c6070e · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:52.757076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:52.757076Z digest=sha256:556f64550813efa6ea6206efe25abd1c3ad40b4c61d6738cc6f6f9ee075756bb

Observation 97e4b629-3eb7-46e6-8365-936fd6bbc686 · outbound

This paper cites ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:52.818163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:52.818163Z digest=sha256:205d54a7bc7174dfd9b0b1f96bb03ddc8ae2361fc8f5c2aea3e40539ae46e4be

Observation e9165a33-7006-40ca-a21b-2993fc6c247c · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:02.855876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:52.944425Z digest=sha256:440113ec7f5d9e28b269cde9866b6a69fa0266ef25c7b9cbd86efc3c932082fb

Observation 5e55dc10-fec7-44d4-b425-c6ba8f86bc6d · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:02.509109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:53.144290Z digest=sha256:7687e8ccd8a0ddfb259ff0a51864fede0e42754e85a1a946789e6d638b4966b4

Observation 294c1e09-1a70-44de-a8aa-74f1604c70ff · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:02.239948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:53.278925Z digest=sha256:3fadce75195da429733dee24f5a4ef2e9c168bb1a614ad1201d7b03122126e07

Observation c45a80dd-5fcd-41b1-809d-994b4910fbe8 · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:01.914402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:53.384267Z digest=sha256:8037baccb9ab1b202bacacab08238c8510a3922d284770adbc037c834c1180d3

Observation e5361a22-6d2b-4a9f-ae62-405c7108ffd6 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:53.467380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:53.467380Z digest=sha256:6b1581c6e96b9b82c799279e5c8562ec0b6c00c8953b0e192b7b431399430342

Observation 7f564ca4-6239-4fff-a60b-700fc1d57c29 · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:01.691535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:53.565570Z digest=sha256:40b3d2a195bcfcbb736acfa734654edaef2b84d84c6a45b997e47e60bb00cadc

Observation 8fb090b0-55be-43d3-81e2-cbf6b6fc61b4 · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:01.464413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:53.733098Z digest=sha256:b299066333a2d0bcc4969a2f59fb998c1470e55ebe63507cd5295345e92cca43

Observation c90c1a2d-35b9-47ea-82fe-1f63f5f1a413 · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:01.007856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:53.885125Z digest=sha256:eff0459b657fca316b87c2f534969308e1112e710afa6c36eec92ad85d5f9b51

Observation 8b674199-294c-4990-b51c-573c57143605 · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:00.614688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:54.030836Z digest=sha256:1f705344ec2d6e534f87cfd8b325926cc5a2a9b5745f1d1c81ef51de45c8c146

Observation 8a139a40-81c4-4ad0-b1d6-3cf9b8b275a5 · outbound

This paper cites Grignetti.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Grignetti

Reference 12

Resolution
verified exact
doi, observed 2026-08-07T13:09:57.909706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:54.187477Z digest=sha256:354f7078ace3d49ee8d7b9329876794f43ea811caba2ed1ea771c98b18739be3

Observation 5d9c37af-94c3-40a3-9450-7fe2478b83fd · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:54.348201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:54.348201Z digest=sha256:978ef1e814696ab3e726b4f310288ced1bb3e11aad1879df1544b70787c4d2a1

Observation 1ec0e489-5c86-4700-a630-1850abaa7d0f · outbound

This paper cites FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:54.477202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:54.477202Z digest=sha256:4001c3aacc35fee5628a28cd78a57a0c050da4c38cb1fbea8f7425f518edd34e

Observation 4ea8d0c7-8626-4ed6-97ff-c88b562bda53 · outbound

This paper cites RSGPT: A Remote Sensing Vision Language Model and Benchmark.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:54.615881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:54.615881Z digest=sha256:c4fe4f534d185dd1bae781d0233d11462962cedc7a1d37d767a445a279464de4

Observation b5dd3067-6217-457f-afee-8abeab642646 · outbound

This paper cites FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:09:58.535733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:54.720587Z digest=sha256:240ddbe1f4d7cc04963f9465bbac125439b634ada69f4e4dfbbac78283a92dbb

Observation 9e83ccb6-a7d0-458f-a039-26a84fa137d3 · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:00.245727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:54.855117Z digest=sha256:183e31b2c7018431bdcfbaa16283a7b477ebb95ae32ecbb2d4151be90dbc1e63

Observation 67223df3-9754-443b-b0e3-7ec53a4cce9b · outbound

This paper cites LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:55.054280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:55.054280Z digest=sha256:e6e092692a18b635039ad00be2932381acb467340546f25267a0eaeaef7f1b2f

Observation 8d7d38e8-3cbb-4df8-af98-df39a8366de6 · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:00.014721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:55.225085Z digest=sha256:4a835c82d0d728dd7d3a9d47755944ebd3d76cac8d3e99f965bca45e638f496b

Observation 5e0d2798-950b-4a36-8727-641072e1a430 · outbound

This paper cites Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:55.361772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:55.361772Z digest=sha256:fa48c3d572ac8be3812b06c63575b07e251190a8d9e903549317b6367f7a0007

Observation 746f77a6-0a68-42e1-a895-63cc1e808e60 · outbound

This paper cites Improving Multimodal Datasets with Image Captioning.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Improving Multimodal Datasets with Image Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:55.512739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:55.512739Z digest=sha256:03fd20f2149a1bd5b6da0a485f6189c45ce1736a262871fedc90d684499878d6

Observation 2a229a72-0d73-45f5-876a-07f8928da016 · outbound

This paper cites DOCCI: Descriptions of Connected and Contrasting Images.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions DOCCI: Descriptions of Connected and Contrasting Images

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:55.647399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:55.647399Z digest=sha256:b6ad13890923070906e9576475bfb9889e3e85809ca3c7fe38f66c5465c22a30

Observation af11f5c5-d302-4dc7-a9f3-23d3a0e51e42 · outbound

This paper cites GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:55.854741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:55.854741Z digest=sha256:e2b6c0280a1ddf368fb987f57a0ff3de475ca890dcccfce9c3e9566125e9c755

Observation fc42d7c7-346d-4616-82ed-346ebea65ab6 · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 24

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T13:09:58.254451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:56.008221Z digest=sha256:d27b63d04adf7eb81df6f34818653d0752d645dd818289b390b613532089cbcc

Observation b64bbed4-834a-400a-a9ec-46eeddedddc8 · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:56.139372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:56.139372Z digest=sha256:a215e15f1026b2fbd85684f0bc3bcbeb8194a031b78eabf534a6bac0eeefc133

Observation ae3e95de-5c85-46d1-95ad-5458b48c5e49 · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:09:59.766832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:56.229648Z digest=sha256:6b4e6cec64776e80475234651079d3968131263f050f3f4ab97324c21b559381

Observation 3736107b-6ebb-480c-a423-a90e243e1f81 · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:56.376558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:56.376558Z digest=sha256:55b3ac41307d32205601bfec6fe229dac672641d74a7c26a78a64c7105bf306e

Observation 71f60c78-f1bd-4e6b-b158-01b0ed7e9a87 · outbound

This paper cites Wobbrock, Kenny Liou, Andrew Ng, and James Landay.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Wobbrock, Kenny Liou, Andrew Ng, and James Landay

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:59.557122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:56.555797Z digest=sha256:b6b38cafba169c74a6a2c03ab48c3624aed23065973466d90289b95ac6c7791b

Observation f64f7796-fdf6-48cf-abc3-728656b2098c · outbound

This paper cites GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:56.728661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:56.728661Z digest=sha256:851fcd450a1d15ba2ddc6a8f2c1dad900a13e3876442a49c1488b6d4f5455dd1

Observation 8f4a2872-484e-4da9-85af-e0d56c56e405 · outbound

This paper cites From Pixels to Prose: A Large Dataset of Dense Image Captions.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:56.862819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:56.862819Z digest=sha256:f2fcf322bf67da4bbcdd2c4fa5053ef4eca336768ea27fc77d2dfae05cd9dae6

Observation 2f2e8bdd-5db6-4499-91b8-1cd5ccd3827c · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:09:59.441744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:56.996995Z digest=sha256:e9cc734ba8d4881cac995ca62e658ac91282ff3da1ab8a6fdb8206c88254f0a0

Observation 61093f1f-7136-4b60-848b-25e774f40996 · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:09:59.180572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:57.107507Z digest=sha256:6c36295d1428ea8fbac115f4a915248ce610b826ac536e61495e023e8c311f99

Observation bbf75e13-d033-40b7-bdb9-46981705fc0e · outbound

This paper cites an unresolved cited work.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:09:58.973364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:09:57.301352Z digest=sha256:de2222fc15fb75e8054f923b6d4d4ffa871b123732d0c3b08f19e51abd86ec16

Observation 8f35ec5a-8402-4b1d-aa55-51f17c26e017 · outbound

This paper cites online" 'onlinestring :=.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions online" 'onlinestring :=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:57.434166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:57.434166Z digest=sha256:e557a175f13524c0b13e156a0b286c5165d97af046f77830a307a2ecb79c295e

Observation 5412e347-c575-4fbb-8019-8e836dc7bcd7 · outbound

This paper cites write newline.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions write newline

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:57.563270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:57.563270Z digest=sha256:418f6d115e87adc793ecfce44ff121f0e8aa3863c4f66acdbf82a1aaa1e4dde8

Pith citing papers

No inbound Pith citation observations are available.