Pith. sign in

Paper Citation Record · LEDGER

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent

As of 23 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2412.11025.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11025 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:24:45.510794Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9f798f5a-8b1f-48b6-b20c-405ec84d3731 · outbound

This paper cites Good news, everyone! context driven entity-aware captioning for news images.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Good news, everyone! context driven entity-aware captioning for news images

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.792115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T15:24:45.422684Z digest=sha256:c0887445521f9ce95760b348895497e843d8a586b7b5cbfceb227e55086cde9d

Observation dcc01e6e-35ec-4c0a-afe4-84776b8f068e · outbound

This paper cites Bge m3- embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation, 2024.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Bge m3- embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.427686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.427686Z digest=sha256:210b2a842113dadcb01a1e36db92bc925fdb562f709e45b81ec4bc902d5729a4

Observation 89011de0-0139-4ba5-b6df-fd8d491c500a · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.432294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.432294Z digest=sha256:ab170f6403a457a793eed3600696d901e2c574bb394dbfc1d2334ed2682f0b64

Observation 62cf5feb-ad2f-4ad4-97c4-a9c25d693aac · outbound

This paper cites Benchmarking and Improving Detail Image Caption.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Benchmarking and Improving Detail Image Caption

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.437278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.437278Z digest=sha256:dc13ede1bdadb5ac7fd9a07bb598a0c2a8f5be7a69330b346dd4d22aa99a7d5c

Observation 3ce52d7b-34ff-4845-8e1d-06e17d618095 · outbound

This paper cites Flex- cap: Generating rich, localized, and flexible captions in images, 2024.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Flex- cap: Generating rich, localized, and flexible captions in images, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.768616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T15:24:45.442433Z digest=sha256:e66791ae1464afcb4cff003d0e26df7163ac2e88e5a64a73b3b150f48ab9ad2c

Observation 557222db-bd1b-4235-b7f7-8c1808cf946d · outbound

This paper cites ImageInWords: Unlocking Hyper-Detailed Image Descriptions.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.447605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.447605Z digest=sha256:07fba240ce25dbae6e72c42297e7afcabe29829a2cf79b22f88105bb501888fb

Observation e0842e1a-3e7c-4ca2-a756-87da781f5ce0 · outbound

This paper cites Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.452767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.452767Z digest=sha256:e4550c5b0cd19c312b206a1dc70cc03f16f0ffa709869bef5f2ab306d2080751

Observation d799bd77-51f6-4de9-bda4-eeda26f96dd2 · outbound

This paper cites Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, and Alan Hayes.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, and Alan Hayes

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.755390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T15:24:45.457139Z digest=sha256:5f71e50b17322d02d0fd2805bcef313b4f13c07d1a844f965fee59e21e32ea79

Observation fa20db53-90ea-4395-b2ee-7b17ab03abdf · outbound

This paper cites Rap: Retrieval-augmented planning with contextual memory for multimodal llm agents, 2024.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Rap: Retrieval-augmented planning with contextual memory for multimodal llm agents, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.742260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T15:24:45.463875Z digest=sha256:d940b28729625a4e461c192f4bee56d2bebf2371a2b9ed9ac622a0d4f9130112

Observation 295be078-6f95-4d4a-8532-579b63485f6a · outbound

This paper cites St Edward's Crown.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent St Edward's Crown

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.728357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T15:24:45.467821Z digest=sha256:8c68c94ed5502f36b4109da6cf664e137fe450d5a1fc6e0eae0792894c1ce41f

Observation 746a2b86-0a58-44b0-a11c-64a66b273267 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection, 2024.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Grounding dino: Marrying dino with grounded pre-training for open-set object detection, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.472879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.472879Z digest=sha256:d4ac6e1fb998773f1a25bfe47b7451364109db420b33a5c3d1b2bb0b6b3c23d3

Observation c678db3f-7e05-418c-8ed4-12bb848d409a · outbound

This paper cites Senticap: Generating image descriptions with sentiments.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Senticap: Generating image descriptions with sentiments

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.706579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T15:24:45.476760Z digest=sha256:56ecfd96b50cd37018cb5312c5e90539ded00c32a56712635b6e788eb2f7667d

Observation 18b39670-52ec-483d-a12f-b4458503fee8 · outbound

This paper cites DOCCI: Descriptions of Connected and Contrasting Images.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent DOCCI: Descriptions of Connected and Contrasting Images

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.480790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.480790Z digest=sha256:14ef3b07106415b86e952366ce6548eb97e36399916eacdad5205e844953a2ed

Observation df434336-d4e5-4e5d-bb17-09a8496acd31 · outbound

This paper cites MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.485184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.485184Z digest=sha256:0afbbbc865ef56065bb87e7d3d202b490ddfb4f24557d0824b1a07dfa82a6319

Observation 1b935a20-74e6-4d84-94be-26e0e78ba214 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.489457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.489457Z digest=sha256:84324c9f5643ba797ffa369848f0d9f2846f34f4942611dd6fb14c56d0d9f33e

Observation 1d414e3d-5147-4484-a82a-3bf687482106 · outbound

This paper cites Transform and tell: Entity-aware news image captioning.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Transform and tell: Entity-aware news image captioning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.684563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T15:24:45.494118Z digest=sha256:147554da79cea4f59973649105ba843d1dc17979dcd68f64d64354021942706f

Observation 2dbf42ab-2fa8-46d2-b1fc-6c1aaebda271 · outbound

This paper cites Caption anything: Interactive image description with diverse multimodal controls, 2023.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Caption anything: Interactive image description with diverse multimodal controls, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.669385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T15:24:45.497871Z digest=sha256:4906f7f6ae41b16260821b908b2ed588e0d5548997017effc42541f6db47e62f

Observation f3af1ea4-6029-4021-b093-4c892a614eaa · outbound

This paper cites Benchmarking Complex Instruction-Following with Multiple Constraints Composition.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Benchmarking Complex Instruction-Following with Multiple Constraints Composition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.502117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.502117Z digest=sha256:b8e6baacbcd80d46d2d7823dd3672e39d58c407521ae9cf8a5d94414735ea99e

Observation c552cfe7-5b28-4fea-bb6d-ccd3e0bd8b40 · outbound

This paper cites Depth anything v2, 2024.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Depth anything v2, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:24:45.650809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T15:24:45.506396Z digest=sha256:d0f106c8c5027b124424047d88439006d74df2d8eb1b50ae5128530e069a49d0

Observation ef7dd429-1e6c-4b22-bbe1-abb71cc03fa6 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent ReAct: Synergizing Reasoning and Acting in Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.510794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.510794Z digest=sha256:a9eb0341b2f167f56058b14b0adca874f4366b9e0e1399e6dcf875620bf73d88

Pith citing papers

No inbound Pith citation observations are available.