Pith. sign in

Paper Citation Record · LEDGER

GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2406.01210.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.01210 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:27:04.711058Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T19:38:10.424227Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fec1c287-c859-48cc-b4a2-c5cfe564494a · inbound

DCIRNet: Depth Completion with Iterative Refinement for Dexterous Grasping of Transparent and Reflective Objects cites this paper.

DCIRNet: Depth Completion with Iterative Refinement for Dexterous Grasping of Transparent and Reflective Objects GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:51:53.625363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:51:53.625363Z digest=sha256:2871aa5142aba6593182b043e0b5c1a7245564e8178d080a9c872409af4bb815

Observation 9c749121-d6e3-46f9-9c3c-d51efdf61bae · inbound

Foundation Models for Clean Energy Forecasting: A Comprehensive Review cites this paper.

Foundation Models for Clean Energy Forecasting: A Comprehensive Review GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:14.069951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:14.069951Z digest=sha256:e0e8800135821ac43204142a762a2608592a48be0b65803242aea2d96cc37d58

Observation ce3ce5fa-4e7b-41c1-b94b-627a907ea971 · inbound

Efficient Segment Anything with Depth-Aware Fusion and Limited Training Data cites this paper.

Efficient Segment Anything with Depth-Aware Fusion and Limited Training Data GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:11.519487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:11.519487Z digest=sha256:1bac5bd1284fcec638656a333fa489efa1c52b20f8f0adee2588a74c3683fd69

Observation 7ee17abe-2af8-4fda-84e3-bf2a74b05ff3 · inbound

CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation cites this paper.

CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:38:10.426281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T19:37:18.613746Z digest=sha256:c98e88dbc4500f65922c210f0a9c962497d515ae4fb0eaeff1d540db0b79c3a7

Observation b120b414-fc30-4265-9788-aab1b94c593b · inbound

RSGMamba: Reliability-Aware Self-Gated State Space Model for Multimodal Semantic Segmentation cites this paper.

RSGMamba: Reliability-Aware Self-Gated State Space Model for Multimodal Semantic Segmentation GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:00.617399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:46:10.742651Z digest=sha256:cac6f67a9bfad5d2bdd81977c6ef8b4a6b6c4019cfc8a1328977fc888c442a6f

Observation eb0f07e3-7fe5-479a-b366-0fe27808807b · inbound

Structure-Semantic Decoupled Modulation of Global Geospatial Embeddings for High-Resolution Remote Sensing Mapping cites this paper.

Structure-Semantic Decoupled Modulation of Global Geospatial Embeddings for High-Resolution Remote Sensing Mapping GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:16:06.024520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T02:06:05.390586Z digest=sha256:a9ec754ec0880d54a0e513c2c9310d5dab3079f9c2951bacb1789cdbdeced6a1

Observation 2066fe50-3937-4e4a-b811-fdbd31971b33 · inbound

Weaving Light and Time: Unified Harmonic-Geometric Representation Learning for Dense RGB-Event Parsing cites this paper.

Weaving Light and Time: Unified Harmonic-Geometric Representation Learning for Dense RGB-Event Parsing GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T05:07:01.933850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:07:01.933850Z digest=sha256:b4abb8f2191e419dfd1d522a63df3bc1d2c43bdc1685bfe115afb5b81e48e6c7

Observation cee53e8f-356b-4e9a-9d74-aef95705a5c4 · inbound

URNet: A Unified Reparameterized Network for Efficient RGB-D Semantic Segmentation cites this paper.

URNet: A Unified Reparameterized Network for Efficient RGB-D Semantic Segmentation GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:04.711058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:04.711058Z digest=sha256:f6e33d180ac3f543494a633774b5c1ea9ecd5de8e4b979913cab627d9b0ba78d