Pith. sign in

Paper Citation Record · LEDGER

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training

As of 23 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2412.19997.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.19997 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:43:47.791790Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 271d822e-5f7e-4d90-96c0-c01a9efb836d · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:43:47.679099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:43:47.679099Z digest=sha256:e2b000ce8566b852fd40a7ab341aad52db6483405cab78ea7e1648d6e9bfd091

Observation 84196178-abc1-4a18-af50-b40547dc5e9d · outbound

This paper cites Vl-bert: Pre-training of generic visual-linguistic representations,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Vl-bert: Pre-training of generic visual-linguistic representations,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:48.227240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.685686Z digest=sha256:f3d2a0f8d79a538fe7ccc704f4f739e8d7e17a0ebc174d98b7313d96ddeb0acd

Observation 50083edc-7a03-4a81-88c1-9d16d088e523 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Oscar: Object-semantics aligned pre-training for vision-language tasks,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:48.208643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.691441Z digest=sha256:4924211e904bcf6a232d5228046d49b590cc8645025c4866c8122271235293ae

Observation a24a72c2-37e8-4397-ae85-02a35a4cc823 · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Vilt: Vision-and-language transformer without convolution or region supervision,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:48.192041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.696588Z digest=sha256:a2679f6af781471d97b975eb5bbb6fdda972c7b2182494e681102f36ea6997f8

Observation db1da44e-93f7-4d37-878d-b1f1f9a19176 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:48.175575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.702281Z digest=sha256:14ff238b796034946e667dda3b87705802ec84a55e248ae453c87161f5174394

Observation 2bf928b7-c73c-4027-8c32-048e63917beb · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Align before fuse: Vision and language representation learning with momentum distillation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:48.155998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.707376Z digest=sha256:1acfb9edd4fb578ae9971e953fb9d2273cc6b5f9b47e0927c9817e479479dced

Observation 7d98765a-5b95-4bca-9f8a-200cefcec455 · outbound

This paper cites Self-distilled dynamic fusion network for language-based fashion retrieval,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Self-distilled dynamic fusion network for language-based fashion retrieval,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:48.134694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.713183Z digest=sha256:8ab55d208cfe9ebe38d06fa26368fbf8e83a24caa02b0baa8cc021dd0e7fcac2

Observation 801820f7-f940-4f7d-8545-f4e0ffbf88a4 · outbound

This paper cites Fashionvil: Fashion-focused vision-and-language representation learning,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Fashionvil: Fashion-focused vision-and-language representation learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:48.116953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.718192Z digest=sha256:ca11b5f1573151b8a2577cf52c92695a6c43bbd2d14889621782c9e855df1020

Observation e8665559-d64f-4865-859c-3dd2f6c8e402 · outbound

This paper cites Fash- ionsap: Symbols and attributes prompt for fine-grained fashion vision- language pre-training,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Fash- ionsap: Symbols and attributes prompt for fine-grained fashion vision- language pre-training,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:48.100411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.722798Z digest=sha256:c0e5547fcb2be96890bb0da2959386bd54ee086a4050c1404605966b0f4652d8

Observation 4337c9df-c258-418b-b96d-522d6a6510c0 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:48.084629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.727838Z digest=sha256:78fead54d02f74d97604e50bb17bf67e454bff0fc34e40e1605e9a847363a8c8

Observation 580fb0c4-045d-406b-a96f-cece8fadede6 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:48.067234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.732922Z digest=sha256:0157f5f89be0cc9884aec93ae2b569e7f9d8c188240d69e6fe84a881b26049a7

Observation 00009df5-9eda-4b04-9ea9-94b08822592b · outbound

This paper cites Fashion-Gen: The Generative Fashion Dataset and Challenge.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Fashion-Gen: The Generative Fashion Dataset and Challenge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:43:47.738356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:43:47.738356Z digest=sha256:52c877fd94eb4f53580a4f6d5004a5fcc4e37b7fcf63787e4e7c185dcb78bf0d

Observation c084c952-1b71-4d6c-9b99-256d95f9be59 · outbound

This paper cites Mmf: A multimodal framework for vision and language research,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Mmf: A multimodal framework for vision and language research,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:48.047908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.743766Z digest=sha256:c94df2192f6ec5b190f576302bbdddd8778f466632f28c2690ee11405194c546

Observation 09303349-15de-459c-be10-03f04ed19acc · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Pytorch: An imperative style, high-performance deep learning library,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:43:47.748097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:43:47.748097Z digest=sha256:8475c9db03dd9f40b0306f62da4d383a4f40ee9afdf7809f7ff0c203fa4f1267

Observation f1cb5497-9af9-45ba-9bde-2c75b574168d · outbound

This paper cites Fashionbert: Text and image matching with adaptive loss for cross- modal retrieval,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Fashionbert: Text and image matching with adaptive loss for cross- modal retrieval,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:48.018928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.752286Z digest=sha256:32b369130181e08ab4adeb767a5fc29c4a77b8109ef7af76cea76a867b77c211

Observation 32ae1a74-f30d-435b-9a2e-7bcd1c3b16b4 · outbound

This paper cites Kaleido-bert: Vision-language pre-training on fashion domain,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Kaleido-bert: Vision-language pre-training on fashion domain,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:47.998159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.756655Z digest=sha256:a046d84670b42f3dd928d6c767b41a105da57d1367080448a81c0d0ca6f93791

Observation dfcc346e-bbbb-451f-a195-94e2431bfce0 · outbound

This paper cites Commercemm: Large-scale commerce multimodal represen- tation learning with omni retrieval,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Commercemm: Large-scale commerce multimodal represen- tation learning with omni retrieval,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:47.978398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.761690Z digest=sha256:f0e11210652acb757a85a5493af369cf6a2eb35a5fa050f47b4e901ae936fd11

Observation 32b8c058-5e14-450c-a943-8c19cbdb8816 · outbound

This paper cites Ei-clip: Entity-aware interventional contrastive learning for e-commerce cross-modal retrieval,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Ei-clip: Entity-aware interventional contrastive learning for e-commerce cross-modal retrieval,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:47.959384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.766440Z digest=sha256:005ac090c6662c2e5fa38271d6330ba5fceac5bda53d2ff07fc8c80a3c3a0d21

Observation 7dcb3607-e542-459c-8687-249d6d16986e · outbound

This paper cites Masked vision-language transformer in fashion,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Masked vision-language transformer in fashion,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:47.939754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.771060Z digest=sha256:e7a8596349b33cf9245be040edad89ed2bdfb08e14c8f3de9f75e69c5b9d2244

Observation 1b27ed18-7edd-4a84-bf01-f8dab45fa7a0 · outbound

This paper cites Fashionklip: Enhancing e-commerce image-text retrieval with fashion multi-modal conceptual knowledge graph,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Fashionklip: Enhancing e-commerce image-text retrieval with fashion multi-modal conceptual knowledge graph,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:47.923492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.775476Z digest=sha256:cda90cf8d0be9f45645a52815d8e507cda6ad0b616f502cdcdbe64f88b34e22e

Observation 2ef245f4-9280-4682-b13d-f3a30b9924fb · outbound

This paper cites Fad-vlp: Fashion vision-and-language pre-training towards unified retrieval and captioning,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Fad-vlp: Fashion vision-and-language pre-training towards unified retrieval and captioning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:47.906818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.781379Z digest=sha256:31b004f61784e1235e655d85678da045abcbba69b26eaddb8d7011929be85395

Observation db0a5a0f-3224-4349-a82b-59db8bff87b7 · outbound

This paper cites Fame-vil: Multi-tasking vision-language model for heterogeneous fashion tasks,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Fame-vil: Multi-tasking vision-language model for heterogeneous fashion tasks,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:47.888796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.786069Z digest=sha256:da6169ab1f73f1b3983a6b796aadd7a3ba4c5b0174ebaf8b963a70638d11e42c

Observation 26506c2f-03b7-4cb1-ade7-90472856ca40 · outbound

This paper cites Syncmask: Synchronized attentional masking for fashion-centric vision-language pretraining,.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training Syncmask: Synchronized attentional masking for fashion-centric vision-language pretraining,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:43:47.869989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:43:47.791790Z digest=sha256:5e226ec09f392c4fdd75f0d312a85438845811d77489dbac64c83ce836af6527

Pith citing papers

No inbound Pith citation observations are available.