Pith. sign in

Paper Citation Record · LEDGER

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent

As of 18 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2412.11025.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11025 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:24:45.510794Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9f798f5a-8b1f-48b6-b20c-405ec84d3731 · outbound

This paper cites Good news, everyone! context driven entity-aware captioning for news images.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Good news, everyone! context driven entity-aware captioning for news images

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.792115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:24:45.422684Z digest=sha256:a2e616da6c74d64effcc35d13b42225bbf3843bfe92600a723faec8a520e9e4e

Observation dcc01e6e-35ec-4c0a-afe4-84776b8f068e · outbound

This paper cites Bge m3- embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation, 2024.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Bge m3- embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.427686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.427686Z digest=sha256:8ee515c4be37a5da2f72bc692de28a9d3a1774971f11f5e864d05c7ecec3078c

Observation 89011de0-0139-4ba5-b6df-fd8d491c500a · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.432294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.432294Z digest=sha256:508e15563c7d81066dcda98b38851255bd365242c505124b77a68ccf4894a0f3

Observation 62cf5feb-ad2f-4ad4-97c4-a9c25d693aac · outbound

This paper cites Benchmarking and Improving Detail Image Caption.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Benchmarking and Improving Detail Image Caption

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.437278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.437278Z digest=sha256:9ec567cf138ab3323ec597eeb0c5765534df094d7baa1ff9d3cf59cc1e7a8f7d

Observation 3ce52d7b-34ff-4845-8e1d-06e17d618095 · outbound

This paper cites Flex- cap: Generating rich, localized, and flexible captions in images, 2024.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Flex- cap: Generating rich, localized, and flexible captions in images, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.768616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:24:45.442433Z digest=sha256:5de680abf17294ba3c9e9af0671837d9fa152f6d257f23280d5441a3ac695db9

Observation 557222db-bd1b-4235-b7f7-8c1808cf946d · outbound

This paper cites ImageInWords: Unlocking Hyper-Detailed Image Descriptions.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.447605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.447605Z digest=sha256:9d914b190c31eae86b667e68178edd427dc46a3bd953a528a3e0ed310cc00497

Observation e0842e1a-3e7c-4ca2-a756-87da781f5ce0 · outbound

This paper cites Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.452767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.452767Z digest=sha256:47c6d9ef07815b248db5a0550ffa86b8d72d33facd01dc7eea327e092cdf0c95

Observation d799bd77-51f6-4de9-bda4-eeda26f96dd2 · outbound

This paper cites Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, and Alan Hayes.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, and Alan Hayes

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.755390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:24:45.457139Z digest=sha256:3b930643f2e93aa11f8d14ada569b49cc1b38a6e3fd36d5a767016a262a4eb22

Observation fa20db53-90ea-4395-b2ee-7b17ab03abdf · outbound

This paper cites Rap: Retrieval-augmented planning with contextual memory for multimodal llm agents, 2024.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Rap: Retrieval-augmented planning with contextual memory for multimodal llm agents, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.742260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:24:45.463875Z digest=sha256:556ea28f7b056f7e059d517b6881d8aba1bac86e652db829a37bd0ed951922a0

Observation 295be078-6f95-4d4a-8532-579b63485f6a · outbound

This paper cites St Edward's Crown.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent St Edward's Crown

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.728357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:24:45.467821Z digest=sha256:1fef178d4ba5981c526c0e2a2079f1c3e6142d20c1d3f7c83c3ce629b7cf66b8

Observation 746a2b86-0a58-44b0-a11c-64a66b273267 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection, 2024.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Grounding dino: Marrying dino with grounded pre-training for open-set object detection, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.472879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.472879Z digest=sha256:bd5eb8b7fdb626dde51acedad219cb2a48ec46651611c8c29cb23b80ce88354b

Observation c678db3f-7e05-418c-8ed4-12bb848d409a · outbound

This paper cites Senticap: Generating image descriptions with sentiments.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Senticap: Generating image descriptions with sentiments

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.706579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:24:45.476760Z digest=sha256:29396ba97f3dc685b6fb2f125fd8d564d9ab561c37e886dfa0f4c01a33698af1

Observation 18b39670-52ec-483d-a12f-b4458503fee8 · outbound

This paper cites DOCCI: Descriptions of Connected and Contrasting Images.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent DOCCI: Descriptions of Connected and Contrasting Images

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.480790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.480790Z digest=sha256:d7f98157466f14f7685e0ce11901e2f91a3f6d6543a22fa27673ed160c00dd07

Observation df434336-d4e5-4e5d-bb17-09a8496acd31 · outbound

This paper cites MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.485184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.485184Z digest=sha256:d197332c4fcfabf29dcf18fa9ce5875d8d85c0252474cdc5f47080bc8afa5cb2

Observation 1b935a20-74e6-4d84-94be-26e0e78ba214 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.489457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.489457Z digest=sha256:5be31c2e2d170c6a895016bcc11a9e8beb6566a0712352a855aad1d9c49df54b

Observation 1d414e3d-5147-4484-a82a-3bf687482106 · outbound

This paper cites Transform and tell: Entity-aware news image captioning.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Transform and tell: Entity-aware news image captioning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.684563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:24:45.494118Z digest=sha256:f3599ee17f284f8bb5eebfa068026cce914e0b9f375c898dde2d34aa2403cbcd

Observation 2dbf42ab-2fa8-46d2-b1fc-6c1aaebda271 · outbound

This paper cites Caption anything: Interactive image description with diverse multimodal controls, 2023.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Caption anything: Interactive image description with diverse multimodal controls, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.669385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:24:45.497871Z digest=sha256:e2fb6795914455951af90c6c964900c65c195a21b0ae404ba802d90dce5e543c

Observation f3af1ea4-6029-4021-b093-4c892a614eaa · outbound

This paper cites Benchmarking Complex Instruction-Following with Multiple Constraints Composition.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Benchmarking Complex Instruction-Following with Multiple Constraints Composition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.502117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.502117Z digest=sha256:aca8b39b9a181b1ee97f14654229f98b98443f0bddd9a98805399932fe237a4b

Observation c552cfe7-5b28-4fea-bb6d-ccd3e0bd8b40 · outbound

This paper cites Depth anything v2, 2024.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Depth anything v2, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.650809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:24:45.506396Z digest=sha256:ef4ebe4692aa52663ca4c9bae6b8980ad16376fba9ce8e17b64b3c87b5fdc178

Observation ef7dd429-1e6c-4b22-bbe1-abb71cc03fa6 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent ReAct: Synergizing Reasoning and Acting in Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.510794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.510794Z digest=sha256:f5402a7632ad636554e11a76c4d6620a63bda78d6782f11c4d70012d46d6c6ba

Pith citing papers

No inbound Pith citation observations are available.