Pith. sign in

Paper Citation Record · LEDGER

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning

As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2507.01409.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01409 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:57:43.814957Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact2
  • verified fuzzy29
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 23c193f0-fcbf-409e-9afb-1c3f6380cf0f · outbound

This paper cites nocaps: novel object caption- ing at scale.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning nocaps: novel object caption- ing at scale

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:50.377461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:40.863410Z digest=sha256:797d248db59a866fa76cb12cb5fcd9eca4460aa20cab8e2aba9da356bf72fce6

Observation 0c028d1e-0a60-42c1-92f7-fe72baa042e9 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:41.010599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:57:41.010599Z digest=sha256:1f0b8822093f009750ffdf69ca495d6f6c4269ae2e08ed7cf33555317dfc1594

Observation 98a3d402-8f0e-484a-b3e2-1e931e31acf5 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:41.063951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:57:41.063951Z digest=sha256:627536ef91f062cd93682114a5487a43518984b6c73c44a7227384eb98d503a6

Observation 4bbe6334-ef0b-4e2e-bd90-19b83e20b05b · outbound

This paper cites CLAIR: Evaluating Image Captions with Large Language Models.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning CLAIR: Evaluating Image Captions with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:41.148599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:57:41.148599Z digest=sha256:57a34357724af672076c141c6c2eec894c0124db437640a826674099f662381e

Observation 4bc130e2-e184-44c5-9f86-021db4f19c83 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Sharegpt4v: Improving large multi-modal models with better captions

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:50.361544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:41.200346Z digest=sha256:d54d744333064f2dd2ca89a49b954e79c2a1720afda02496de3c9e75f99cbede

Observation c3e50c8b-5a6d-4037-8224-16f6d5b40117 · outbound

This paper cites Learning distinct and representative modes for image captioning.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Learning distinct and representative modes for image captioning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:50.256312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:41.272463Z digest=sha256:a17206c1cdbdbfe92554a86c9ad578554c6f30d1b628c96a1c29b4243e07c0c2

Observation 039945e0-ed40-4bbc-81e1-adb3b38bcef7 · outbound

This paper cites Say as you wish: Fine-grained control of image caption generation with abstract scene graphs.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Say as you wish: Fine-grained control of image caption generation with abstract scene graphs

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:50.130384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:41.323774Z digest=sha256:00cc38f2e8d24c5f287175dbfe98b8bcc04c90677858a9e83eb95bc70b373f6d

Observation d22aa95c-0deb-4224-b9d0-4c60011d7d24 · outbound

This paper cites Length- controllable image captioning.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Length- controllable image captioning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:49.947292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:41.429597Z digest=sha256:29f8a17cad09a73779b093470e7334eb1355bb43b3071f655dde6d71bdcc380f

Observation 6cf09b56-11aa-473c-8825-2501cac2aac5 · outbound

This paper cites Flexcap: Describe anything in images in controllable detail.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Flexcap: Describe anything in images in controllable detail

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:49.711358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:41.519290Z digest=sha256:06bffe619dfa9a66538d17a6a2785a902e108dbce10090a308ee8060c5454147

Observation 9e47b379-494d-45b2-b272-ed203ffb7691 · outbound

This paper cites Captioning images taken by people who are blind.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Captioning images taken by people who are blind

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:49.458330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:41.604022Z digest=sha256:9099697bbca00ff322c243fe669aa57b0650e8f2583bbe3dbb6f8b4c5408863d

Observation 808936fa-ad98-4e25-ae44-08d21972e8a3 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:41.666447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:57:41.666447Z digest=sha256:b95feb5b0c485647e49d97c25ddb717cbb8ac40bd2fa44440e10b0d28e520eae

Observation a578ebce-6207-4a05-81cf-045317bd1fcb · outbound

This paper cites From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:57:44.319196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:41.739561Z digest=sha256:341adf292717b25529f7821445c01216ce13d6c1625c2ad66447040129544f91

Observation 704ae124-6352-48a4-b68a-a73072c524a5 · outbound

This paper cites spaCy 2: Natural lan- guage understanding with Bloom embeddings, convolutional neural networks and incremental parsing.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning spaCy 2: Natural lan- guage understanding with Bloom embeddings, convolutional neural networks and incremental parsing

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:49.291660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:41.792838Z digest=sha256:63e149a91e61dd20484051f591c24a646dd04ac69033d57290c4887c5afc6dde

Observation 240ab681-5ec4-44e4-8537-a8f1de854a0d · outbound

This paper cites Scaling up vision-language pre-training for image captioning.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Scaling up vision-language pre-training for image captioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:49.186220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:41.867976Z digest=sha256:b162cfe9ef0aeb1745399cae5990b0e819731a7bdccd0f274f8d5190a5aa7412

Observation 734de0ac-2c70-48d9-af20-4d474fbdc132 · outbound

This paper cites Noise-aware learning from web-crawled image-text data for image captioning.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Noise-aware learning from web-crawled image-text data for image captioning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:49.030536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:41.936237Z digest=sha256:7f4004b7b205d85ebb7121ff74a2ecc1b55ee42080e46660bf921261dfba3ffe

Observation 78d3db17-31d7-492a-843b-7380dd714b6f · outbound

This paper cites Imageability-and length-controllable image caption- ing.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Imageability-and length-controllable image caption- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:48.874016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:42.044551Z digest=sha256:0f9c49ba2584c569bd79fe0d09b14acfcb04461bbd718273dead49c30d451829

Observation 71e3f1bb-f1b5-49e5-9a59-50eb91f8f388 · outbound

This paper cites Novel dataset for fine-grained image categorization: Stanford dogs.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Novel dataset for fine-grained image categorization: Stanford dogs

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:48.723129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:42.109552Z digest=sha256:015a802a3a9a34ce364bf94b2e85a6d88455e351059b58e9bc12cf9fd5580da8

Observation e99e17d9-f1f5-4a99-8950-db1539daf23d · outbound

This paper cites 3d object representations for fine-grained categoriza- tion.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning 3d object representations for fine-grained categoriza- tion

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:48.556463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:42.170472Z digest=sha256:b7b3020628362870b76ec3f0aacc18f3f70ed4b913125336632710b960cc8303

Observation 542a7d66-22b8-443b-ab42-d781021d9d9c · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:48.420916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:42.257722Z digest=sha256:8cf432f64ef4989385ad4ccc494afafed38857cd2c7486040052ab5aa16a2287

Observation 9f90f9c3-c286-4526-a596-672d18b9fefb · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:48.285100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:42.330718Z digest=sha256:94ba8ab66aa4ae2cab6a1395df184e54f2dfda06297741ade9633c85703e0dad

Observation aa863290-8ad4-4580-9f3a-a8b08ebbd7ab · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Rouge: A package for automatic evaluation of summaries

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:42.442960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:57:42.442960Z digest=sha256:7821719653ca4648940a83d6054b9205cd4a1cf5f49886720177d6c5a1f03542

Observation 7b9ddc2a-5de5-4bc2-8679-2a349ffbc6c7 · outbound

This paper cites Microsoft coco: Common objects in context.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Microsoft coco: Common objects in context

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:48.124268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:42.488770Z digest=sha256:907323dd68dbbf88dd8010143daa4ec6d4a70b1dc74ba6a38cbd2343fb7520ed

Observation 1eaa68be-ff7f-4a48-a2a1-d1c81cf2cf1b · outbound

This paper cites Improved baselines with visual instruction tuning.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Improved baselines with visual instruction tuning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:47.990037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:42.544533Z digest=sha256:bd6329343ac2a0911bfd128cc978c743c70a9f63f19ed9847469ecd46b686c56

Observation d468e8f9-ca17-4a9e-af92-4dc22603aff2 · outbound

This paper cites Visual instruction tuning.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Visual instruction tuning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:47.830214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:42.619570Z digest=sha256:f257e9ac2cb629721f02c3754d954c43b1b4aa85f72f59aa033f2862da0fa1e6

Observation 97266667-6336-4c80-b398-c6c3d72c1260 · outbound

This paper cites Quark: Controllable text generation with reinforced unlearn- ing.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Quark: Controllable text generation with reinforced unlearn- ing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:47.636736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:42.665570Z digest=sha256:913f5e906c986de2271a694ad453acd6b8f16fa54654aca02b6aa7e019bb7c19

Observation 4b44490d-1b22-4409-9e77-2e74025b1cb4 · outbound

This paper cites Analysis of diversity-accuracy tradeoff in image captioning.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Analysis of diversity-accuracy tradeoff in image captioning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:57:44.081239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:42.727730Z digest=sha256:e1b45a81483e90cd0e39f8c1b127296ec4bf7f0a8933072f152e804494d1795f

Observation 8136eee2-7055-4fe8-9a3a-7ec44d51c14f · outbound

This paper cites Wordnet: a lexical database for english.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Wordnet: a lexical database for english

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:47.493190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:42.793847Z digest=sha256:8178c5d582ed0c551368d2c71e4e008a7f7e89d7b4d19825c869017318eb71bb

Observation 0016f794-ffe0-4ba3-8610-bad21f19fe48 · outbound

This paper cites Docci: De- scriptions of connected and contrasting images.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Docci: De- scriptions of connected and contrasting images

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:47.295276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:42.839594Z digest=sha256:f6f861a5014adc70eb0b9a7fd56bf2f0b759f667be03636a34a897239060f479

Observation fb2e6639-b104-4cf7-bd81-bd65ec352477 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Bleu: a method for automatic evaluation of machine translation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:42.880815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:57:42.880815Z digest=sha256:d669f1d11ad83f311b1c1e42fda6a161e75fc7e3dfa64a5785666a7532262300

Observation eed8de9d-100e-448d-a8d2-3ae512c021eb · outbound

This paper cites AdapterFusion: Non-Destructive Task Composition for Transfer Learning.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning AdapterFusion: Non-Destructive Task Composition for Transfer Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:42.933340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:57:42.933340Z digest=sha256:d02fa1faa741fcecdf08aab8ac4681971b2864014d3b27ece955b5fe92157728

Observation 5316de52-961b-40c0-b7db-852e7a0fc40a · outbound

This paper cites Connecting vision and lan- guage with localized narratives.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Connecting vision and lan- guage with localized narratives

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:47.149194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:42.998228Z digest=sha256:5428ff3308660d9abc85256cd3ef888e5ebac087d682e62168c5d72ae7d8959d

Observation 5618f520-0610-4196-a40a-fb661c7de7a5 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Learn- ing transferable visual models from natural language super- vision

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:46.905067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:43.052984Z digest=sha256:efa5fca371215d5ce0501320ae621184377954fd2206255e2ec7ab9ae34ca0ba

Observation 88721885-833b-4233-befa-73387514739f · outbound

This paper cites Laion coco: 600m syn- thetic captions from laion2b-en, 2022.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Laion coco: 600m syn- thetic captions from laion2b-en, 2022

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:46.655352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:43.118604Z digest=sha256:4c226262de264ff09eb6b37b634605189200653aaaffe42414059d977e9ae78c

Observation 4de8f8cd-bf19-4567-b6f5-90febde4c629 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:43.172375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:57:43.172375Z digest=sha256:3edec3087e5463e255995b80026f930dfb32daac629906f9ce23112fac9c234f

Observation 93576c59-dea6-48f1-82c6-a92beb03628d · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Cider: Consensus-based image description evalua- tion

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:46.295126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:43.233214Z digest=sha256:8decb2e7817b3ca23dac858dba5e09ff5b51307148347799388d9f29c625f222

Observation d4112cec-91a4-4d86-b55a-f3fb3cce6801 · outbound

This paper cites SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:43.319076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:57:43.319076Z digest=sha256:4d4db8cbd59345f9272ea2b04c7f468cb7022019eb4bd89e15243d4cd2fd93d7

Observation 371328cd-aa87-47ea-abd7-b37cfe608627 · outbound

This paper cites an unresolved cited work.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:57:46.011132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:43.370694Z digest=sha256:c0a060f13a33f608e9457f1afac3da7860900bdd50ac8e13a15443538e73056d

Observation bd45be43-a0ea-4c92-bf7e-bf57b20515ce · outbound

This paper cites Controllable image captioning via prompting.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Controllable image captioning via prompting

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:45.742872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:43.418905Z digest=sha256:30c871f5ef7c0eeeef558584123a73628c511367927e2dc5e98e4acf89440c8e

Observation cb589c28-6c8f-4190-9d92-a3ea9585cc13 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:43.481086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:57:43.481086Z digest=sha256:d516243973e992cc2dd981a52556c258b0cefb1c66097be7c77a84b1822f94cd

Observation 1f7af472-dad7-43a5-98e0-e99d3b1234bc · outbound

This paper cites On diversity in image captioning: Metrics and methods.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning On diversity in image captioning: Metrics and methods

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:45.467822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:43.527334Z digest=sha256:6dfe4c6e7f1f208a72d181c7dc1c4d34d89be096b80d693bc3e2b2256b78f33f

Observation c0d0ba97-855c-44fb-a359-3db2b8f9ff63 · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing in- ference time.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing in- ference time

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:45.173175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:43.582119Z digest=sha256:e849c3981056644274df15d6c936a70cd5e243f5c3b412d6ea1607eff9785fbd

Observation a3a58286-6d6d-4192-9d4d-cba0c6cb14ff · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning xgen-mm (blip-3): A family of open large multimodal models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:43.643606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:57:43.643606Z digest=sha256:62a709245a472e5e773c0b3c1a2d26733d26d69b3a0ce8192728736409bfbce4

Observation f27bd746-0aee-4691-8b65-77d1e0726134 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:43.704897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:57:43.704897Z digest=sha256:ed91fc08c55017cc723324fb5047ed874dd6f643a5ae1f649227007391797e4d

Observation 54e0c28a-1b44-4e36-bf23-211828c73dde · outbound

This paper cites Regularization and variable se- lection via the elastic net.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning Regularization and variable se- lection via the elastic net

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:57:44.936216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:43.754324Z digest=sha256:846c6b982cd21d3f0e34aeb780ac0bbd80b178ac90c5c113a4d9c454b129ba82

Observation db19b62a-88e5-44e8-955f-06ab7ca7ac7c · outbound

This paper cites image”, “side.

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning image”, “side

Reference 2005

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T20:57:44.637445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:57:43.814957Z digest=sha256:1c3706282d451313e3bb9ea59088f33a144f8491161ca07b2de9fea52e2ff441

Pith citing papers

No inbound Pith citation observations are available.