Pith. sign in

Paper Citation Record · LEDGER

GeoMM: On Geodesic Perspective for Multi-modal Learning

As of 16 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 0 inbound Pith citation observations for arXiv:2505.11216.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11216 v1

Coverage vector

measured 100 of 111 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:01:38.878311Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 111 outbound references displayed

  • verified exact3
  • verified fuzzy44
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 66cfa46f-a0bd-4a78-a9c2-e6ee31dbede3 · outbound

This paper cites Geometry of oblique projections.

GeoMM: On Geodesic Perspective for Multi-modal Learning Geometry of oblique projections

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:01:39.347568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.413972Z digest=sha256:73841f2aa5cd58ce2ab8b91e9d830b7bebd03c60d0ad7d7d83b2f237a6d58a62

Observation 8659d41e-7dc6-4ffd-93b0-1cb8e5a0b684 · outbound

This paper cites Vqa: Visual question an- swering.

GeoMM: On Geodesic Perspective for Multi-modal Learning Vqa: Visual question an- swering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.420078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.420078Z digest=sha256:c87b696856a60bea6f57f77d1429a54f77476e84c40fb6ef881682a492bb5d1c

Observation 55c1edf9-c5ee-4ba4-9e2f-9b139486cc82 · outbound

This paper cites Geodesic matting: A framework for fast interactive image and video seg- mentation and matting.

GeoMM: On Geodesic Perspective for Multi-modal Learning Geodesic matting: A framework for fast interactive image and video seg- mentation and matting

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.424841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.424841Z digest=sha256:78b9e40a30f7ec8e96d9a33821e9938fca6969c8d84faf3afac2748c862fa2cf

Observation b7dec225-807d-4d7d-82d5-77a26f2fff1c · outbound

This paper cites Vlmo: Uni- fied vision-language pre-training with mixture-of- modality-experts.

GeoMM: On Geodesic Perspective for Multi-modal Learning Vlmo: Uni- fied vision-language pre-training with mixture-of- modality-experts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.429348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.429348Z digest=sha256:f68ede0524c8101233e1723bf415f0c98a36692ba4f219845f30b34d5f023a91

Observation 21f0c57b-7496-4784-8aa4-7068f74eefbe · outbound

This paper cites Grit-vlp: Grouped mini-batch sam- pling for efficient vision and language pre-training.

GeoMM: On Geodesic Perspective for Multi-modal Learning Grit-vlp: Grouped mini-batch sam- pling for efficient vision and language pre-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.434050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.434050Z digest=sha256:060137a2bef182bc15f14afee58d82869b05d1cd452fa7123422a1cc5616207e

Observation 7a3ce4a2-f859-4cbe-af60-3c16ac592db2 · outbound

This paper cites Mafa: Managing false negatives for vision-language pre-training.

GeoMM: On Geodesic Perspective for Multi-modal Learning Mafa: Managing false negatives for vision-language pre-training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.438839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.438839Z digest=sha256:09efac11b678650d87b7d51b69aeae06f505e6269e48636dbe0e027e7d79eb75

Observation a48256e5-9822-428d-9acf-0f65b344a5e5 · outbound

This paper cites End-to-end object detection with trans- formers.

GeoMM: On Geodesic Perspective for Multi-modal Learning End-to-end object detection with trans- formers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.443858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.443858Z digest=sha256:66cc9c04dedc406f8833109c9700c86771046aca5ad4edde19dde5434c0ef1db

Observation db3622f6-0593-4b8f-87af-fff6fcf727bb · outbound

This paper cites Un- supervised learning of visual features by contrasting cluster assignments.

GeoMM: On Geodesic Perspective for Multi-modal Learning Un- supervised learning of visual features by contrasting cluster assignments

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.448550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.448550Z digest=sha256:c5f5efa2f6d176a768a37fb335c70299a0292d1817346ccad14e2adeea72643a

Observation fb5ad8ad-31d2-4fa2-a7ab-c9bb930b2de1 · outbound

This paper cites Emerging properties in self-supervised vi- sion transformers.

GeoMM: On Geodesic Perspective for Multi-modal Learning Emerging properties in self-supervised vi- sion transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.452997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.452997Z digest=sha256:e873892269953358ff6a6c1f84f750bf9d5a90b599d245116099fb99c274b1f7

Observation 99d13011-6d78-4073-b0c4-48eab8b4d2d9 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

GeoMM: On Geodesic Perspective for Multi-modal Learning Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.457508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.457508Z digest=sha256:8e8702028b4adb6e5383b561a7213310cbf44b959c7aef19599bc6cb39073982

Observation 274db597-f706-466c-94af-bff1aab1ab82 · outbound

This paper cites STAIR: Learning Sparse Text and Image Representation in Grounded Tokens.

GeoMM: On Geodesic Perspective for Multi-modal Learning STAIR: Learning Sparse Text and Image Representation in Grounded Tokens

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.462143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.462143Z digest=sha256:f1424d4056994d7574155b5ced53d98e63611aa1150a6e07e4ba7d6c864009eb

Observation 65cd4779-0669-4769-8fb5-cf7d80842a14 · outbound

This paper cites Vlp: A survey on vision-language pre-training.

GeoMM: On Geodesic Perspective for Multi-modal Learning Vlp: A survey on vision-language pre-training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.466898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.466898Z digest=sha256:aa2adf8ef8b81ff62cf1aa2d621eb57ded548e01baf453c5b2aacaa5796f0ba5

Observation 2fe0d320-69d6-4761-914f-af8965a929be · outbound

This paper cites A simple framework for con- trastive learning of visual representations.

GeoMM: On Geodesic Perspective for Multi-modal Learning A simple framework for con- trastive learning of visual representations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.471735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.471735Z digest=sha256:2f706dff2c348afe5348d42df5fcb198e34b65158cdd382d6182e0ab69be4d7f

Observation 004bccc6-a759-477a-b537-6ce78ea73233 · outbound

This paper cites Exploring simple siamese representation learning.

GeoMM: On Geodesic Perspective for Multi-modal Learning Exploring simple siamese representation learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.475985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.475985Z digest=sha256:893b58fa5fada94f7186930656414521d56c10ce48ee3903e7257daaf0d23146

Observation 4196cc9e-1fce-4152-ae6c-f046f859e2b4 · outbound

This paper cites Improved Baselines with Momentum Contrastive Learning.

GeoMM: On Geodesic Perspective for Multi-modal Learning Improved Baselines with Momentum Contrastive Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.480255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.480255Z digest=sha256:0c0451e7f37209a8022d8f561bfcf0dbe4f82a8ff2039b4855493ae15c9e114b

Observation 8adae875-cc7c-4b7c-959d-a92eeedff320 · outbound

This paper cites X-volution: On the unification of convolution and self-attention.

GeoMM: On Geodesic Perspective for Multi-modal Learning X-volution: On the unification of convolution and self-attention

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:01:39.294667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.484805Z digest=sha256:d100f794f7bee88abffbe23200a811e726839080a2d1a4e7e3e46c0ecdd0adc6

Observation ad3267df-c538-4ee7-b467-0337bcb55ec6 · outbound

This paper cites Uniter: Universal image-text represen- tation learning.

GeoMM: On Geodesic Perspective for Multi-modal Learning Uniter: Universal image-text represen- tation learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.489680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.489680Z digest=sha256:102d7ef71dc1758aabb623b93e77affbfbc69d785aee317235438b02cbeba33d

Observation b5b4d7b4-0948-4db1-bff5-bb6998010a34 · outbound

This paper cites Unsupervised Opinion Summarization Using Approximate Geodesics.

GeoMM: On Geodesic Perspective for Multi-modal Learning Unsupervised Opinion Summarization Using Approximate Geodesics

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:01:39.267081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.494361Z digest=sha256:809a24a3d84244ddaaac8d2e260fd6f0880fae3b3a386caa739006c9dbcc46e1

Observation 54f20193-659f-4878-8e05-b9d15bf37185 · outbound

This paper cites Geodesics in heat: A new approach to computing distance based on heat flow.

GeoMM: On Geodesic Perspective for Multi-modal Learning Geodesics in heat: A new approach to computing distance based on heat flow

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.499121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.499121Z digest=sha256:f28ac49073caa2bc41c0150faaf2745fdd1eb1bca8984d77444a02848e2f08b3

Observation 3478c4dc-1800-4b8a-8fa1-150a8fac611b · outbound

This paper cites Imagenet: A large-scale hierar- chical image database.

GeoMM: On Geodesic Perspective for Multi-modal Learning Imagenet: A large-scale hierar- chical image database

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.503671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.503671Z digest=sha256:0cd6d09aeb6a10195be3428c2590470177dd2d89dfb239eed3f3b260b217c49c

Observation b6a0e3ef-a335-403a-afe0-dea09844afcb · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

GeoMM: On Geodesic Perspective for Multi-modal Learning BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.508186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.508186Z digest=sha256:467ed69ea5603513049c2dcb5bb17459fce5c4f0578288c818048b61d41fb7cc

Observation 563e7ed6-9e23-49c7-8249-556a7a4e2d39 · outbound

This paper cites Similarity reasoning and filtration for image-text matching.

GeoMM: On Geodesic Perspective for Multi-modal Learning Similarity reasoning and filtration for image-text matching

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.512748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.512748Z digest=sha256:d860ec7b6cf772bde883d11a656b516e20429b7692ae42973be3bbb5131eb4cd

Observation 8ede5920-4d59-4b00-996e-06fa677e62b4 · outbound

This paper cites A note on two problems in con- nexion with graphs.

GeoMM: On Geodesic Perspective for Multi-modal Learning A note on two problems in con- nexion with graphs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.517378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.517378Z digest=sha256:cd10afb87018e264c55b01b6adf6a83698f6b92c119f50fda5570749ae12cdf3

Observation 2bf51a31-6b6b-42d0-a3bf-253e0f9ca610 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

GeoMM: On Geodesic Perspective for Multi-modal Learning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.522500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.522500Z digest=sha256:0319e9452fc20417bdc2382ac1fefabc9f4060050380a12fdc8a5fedc858f2e8

Observation 31a91c94-7b3c-49b9-b09a-29f0d3911a19 · outbound

This paper cites Algorithm 97: shortest path.

GeoMM: On Geodesic Perspective for Multi-modal Learning Algorithm 97: shortest path

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.528050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.528050Z digest=sha256:2c3e346cc03d57dab0f0f7aa8a21b591b224c59ed56da3a1c6e40c8b88f3bc68

Observation 3ef1901e-d388-47be-9797-5288ce17d716 · outbound

This paper cites Large-scale adversar- ial training for vision-and-language representation learning.

GeoMM: On Geodesic Perspective for Multi-modal Learning Large-scale adversar- ial training for vision-and-language representation learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.532702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.532702Z digest=sha256:d22476ebbee09126c4c1f93cc3caf33e6ea0429d5de5fda660b0a7186d5fc4e1

Observation 94e9c6b0-76e3-4cfc-825b-3ce545cc4019 · outbound

This paper cites Imagebind: One embedding space to bind them all.

GeoMM: On Geodesic Perspective for Multi-modal Learning Imagebind: One embedding space to bind them all

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.537300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.537300Z digest=sha256:70c93ddadac8d6bf07f764651bd6101f4dd487493d1b8c675001d7b6c8b61242

Observation 53a7fab5-efb6-4ff8-bbec-c2242f7cbfc4 · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

GeoMM: On Geodesic Perspective for Multi-modal Learning Momentum contrast for unsupervised visual representation learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.541805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.541805Z digest=sha256:83decf3d6e5207c740fadde2d80d5fb995acbcfbaa112173da8e27a112f690ab

Observation a3f9074b-fbab-445c-a06a-d73412d91706 · outbound

This paper cites Masked autoen- coders are scalable vision learners.

GeoMM: On Geodesic Perspective for Multi-modal Learning Masked autoen- coders are scalable vision learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.546457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.546457Z digest=sha256:f564193fde295d5a1fde59d0aa9a3e159e588727bec1891d277599fbc542ae67

Observation fc0aedf2-1406-4814-93fc-e27dac959f1f · outbound

This paper cites Geonet: Deep geodesic networks for point cloud analysis.

GeoMM: On Geodesic Perspective for Multi-modal Learning Geonet: Deep geodesic networks for point cloud analysis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.551765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.551765Z digest=sha256:0894f8b6f926e5641ef49e28f50ef5420555eab51f930993f13aba1c3473de86

Observation 276fe121-e048-4455-a1f0-4be44f6352e2 · outbound

This paper cites Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers.

GeoMM: On Geodesic Perspective for Multi-modal Learning Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.561792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.561792Z digest=sha256:514456fa366d7d8e1172436c97dab116db16461df7a2ba024d9f05364d4e0150

Observation 8ab911b9-f86b-4331-9caf-a74d29e77a59 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

GeoMM: On Geodesic Perspective for Multi-modal Learning Scaling up visual and vision-language representation learning with noisy text supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.566170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.566170Z digest=sha256:4c94b40619479894ee0b497254684eebb76f5187baae9994272dd0ec8cfb4892

Observation 4da22129-6ec5-4438-9d0f-38cac98f3077 · outbound

This paper cites Vilt: Vision-and-language transformer without convolu- tion or region supervision.

GeoMM: On Geodesic Perspective for Multi-modal Learning Vilt: Vision-and-language transformer without convolu- tion or region supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.570688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.570688Z digest=sha256:77733bcc403d2b7aa79d5e2ede914549a557f844da899b1f154021023cb0de21

Observation 43dd0a4b-2483-4364-90ad-ed7f8b51290b · outbound

This paper cites Computing geodesic paths on manifolds.

GeoMM: On Geodesic Perspective for Multi-modal Learning Computing geodesic paths on manifolds

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.574725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.574725Z digest=sha256:c09b97a389065a770c7c088e545ebc5fdd83626a3901729f577873b98e1867d8

Observation 27c0734b-23d0-43c1-9612-5e6207f46290 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annota- tions.

GeoMM: On Geodesic Perspective for Multi-modal Learning Visual genome: Connecting language and vision using crowdsourced dense image annota- tions

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.579261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.579261Z digest=sha256:4cb7546d4378298e8ddc9df293cb5adf1f1cc6da6c783a2687cf8f9e39c13362

Observation 7e43bedb-4f44-468c-bcf5-b3c469f4d285 · outbound

This paper cites Numba: a llvm-based python JIT compiler.

GeoMM: On Geodesic Perspective for Multi-modal Learning Numba: a llvm-based python JIT compiler

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.583852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.583852Z digest=sha256:d86eb21c7e19157094f1ff656ad09a892dae98466d09e44a80c548887cf42af4

Observation aa33debf-79a4-4188-b5a2-4d393a1c1ee8 · outbound

This paper cites Le, Vu Nguyen, Chen-Ping Yu, and Dimitris Samaras.

GeoMM: On Geodesic Perspective for Multi-modal Learning Le, Vu Nguyen, Chen-Ping Yu, and Dimitris Samaras

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.588469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.588469Z digest=sha256:747e05b56381a6ba378ae181d0353a390e8ae0764fcd6198a5d0a6aa1de9eea0

Observation 46ea7ca2-464f-42cf-bda9-8f1d75c8caee · outbound

This paper cites Stacked cross attention for image- text matching.

GeoMM: On Geodesic Perspective for Multi-modal Learning Stacked cross attention for image- text matching

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.592895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.592895Z digest=sha256:99d73ec0e9e3008d82e94c54d1516824e5c03dd6b30c3286518041481c1f0901

Observation 23f519ed-53c6-45ae-a6e6-f7b0681e5b9d · outbound

This paper cites Multimodal Foundation Models: From Specialists to General-Purpose Assistants.

GeoMM: On Geodesic Perspective for Multi-modal Learning Multimodal Foundation Models: From Specialists to General-Purpose Assistants

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.597292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.597292Z digest=sha256:002e65cbbcdd2ccf20bc62220ffc6c1bc601a1bf56ef30362b05bd01e959f884

Observation 71eddf66-8905-4e23-8355-6cd0d419bc8e · outbound

This paper cites Align before fuse: Vision and lan- guage representation learning with momentum dis- tillation.

GeoMM: On Geodesic Perspective for Multi-modal Learning Align before fuse: Vision and lan- guage representation learning with momentum dis- tillation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.602455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.602455Z digest=sha256:4a77b6a9160b65bae7765c90b80b64a0c0ccce46c507670e9e762629d05302b7

Observation f3f7b385-b0a0-4dbb-88c2-1a9fd92507e0 · outbound

This paper cites Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation.

GeoMM: On Geodesic Perspective for Multi-modal Learning Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.191284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.607159Z digest=sha256:83b8ab2badd7c6e944db3ad40ff931a3cbb65d2beb243d11dc1cf3814e906dfa

Observation 365af228-59ce-46e0-ae1c-7e91497c43c6 · outbound

This paper cites HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training.

GeoMM: On Geodesic Perspective for Multi-modal Learning HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.611677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.611677Z digest=sha256:be354fc1a1dae8755af61b61f121feddecd36a36ccf8b3b3051eb97220bcc865

Observation 4cb98fbf-d047-4e6d-bfcd-bb8854344691 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

GeoMM: On Geodesic Perspective for Multi-modal Learning VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.616515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.616515Z digest=sha256:1d39d3d485ba40873b5ad1f79613a6150bf54daf3f1ab2ad28f35b264f76cc0a

Observation 4d8b67ea-7b08-4c7a-89fb-c7d019169cfe · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

GeoMM: On Geodesic Perspective for Multi-modal Learning Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.177109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.621318Z digest=sha256:219cd13c5ed19390c5f61af6c10da7129376dfb1b5b975dba04913dbbf9a70c0

Observation b18e0303-3dcd-4727-b2b8-c843df5c473e · outbound

This paper cites Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm.

GeoMM: On Geodesic Perspective for Multi-modal Learning Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.625976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.625976Z digest=sha256:96221ef2cfd45cc7c41d0ecc4d52b36278cdbfe8300c2280ca64e94620803394

Observation c4cc866d-c672-4f5d-9e80-c0fa2e1bc243 · outbound

This paper cites Scaling language- image pre-training via masking.

GeoMM: On Geodesic Perspective for Multi-modal Learning Scaling language- image pre-training via masking

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.163177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.630755Z digest=sha256:f513e2ac3d462baa512465a0822cc0ef761f6c1755ee7591f0e71e632e8e96b6

Observation 9254e8d4-3858-4246-a97a-f3fa3252d5e7 · outbound

This paper cites Geodesic self- attention for 3d point clouds.

GeoMM: On Geodesic Perspective for Multi-modal Learning Geodesic self- attention for 3d point clouds

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.148944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.635022Z digest=sha256:e50a20e9d233d0c35338b35eb36986963979bb2c8ca610aaa03035db99290653

Observation 06178822-4c25-4d71-a750-3cdcaa1e655d · outbound

This paper cites Microsoft coco: Common objects in context.

GeoMM: On Geodesic Perspective for Multi-modal Learning Microsoft coco: Common objects in context

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.134318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.639316Z digest=sha256:f2246e03f42825702f2d29ae6f05fc8aeba81a81fd837cda139d80d7615d7b91

Observation 60a9f332-f09a-4b03-abb3-693a13838ba4 · outbound

This paper cites an unresolved cited work.

GeoMM: On Geodesic Perspective for Multi-modal Learning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:01:40.118132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.644092Z digest=sha256:987c2c9a72c87d5cd7d17f8bb8a5563b862cd008fe33c63aafdf63fd34d95c12

Observation 39a9358a-fce3-4c28-a839-3c9deeedb799 · outbound

This paper cites Adap- tive reconstruction network for weakly supervised re- ferring expression grounding.

GeoMM: On Geodesic Perspective for Multi-modal Learning Adap- tive reconstruction network for weakly supervised re- ferring expression grounding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.102936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.648566Z digest=sha256:9a0568f8fcd3eb12bd1e50f05399d6681a212718308317a21ae99f6445c7c16c

Observation 2351d9c4-6ef2-45fb-8709-8d500a8fbb70 · outbound

This paper cites Algorithm as 136: A k-means clustering algorithm.

GeoMM: On Geodesic Perspective for Multi-modal Learning Algorithm as 136: A k-means clustering algorithm

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.088456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.652851Z digest=sha256:96be3684cc7f4fac0f8885acb7cdc9702a31782fd087281002f84d2b7847c3ef

Observation 4d014c8a-5557-4533-8c80-b64489b474eb · outbound

This paper cites Decoupled weight decay regularization.

GeoMM: On Geodesic Perspective for Multi-modal Learning Decoupled weight decay regularization

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.074084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.657146Z digest=sha256:71cf4a50067837496d0a527375d63d956f6f50126875bc6b402880e057e90b0c

Observation fe50d3e3-97e6-4ba1-989f-c0f4b8c4e66e · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic rep- resentations for vision-and-language tasks.Advances in Neural Information Processing Systems, 32, 2019.

GeoMM: On Geodesic Perspective for Multi-modal Learning Vilbert: Pretraining task-agnostic visiolinguistic rep- resentations for vision-and-language tasks.Advances in Neural Information Processing Systems, 32, 2019

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.059524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.661372Z digest=sha256:ddf811d61a9ec9a9cef7c405a03709db8d82f6633ed11af879290c6f67948dd0

Observation 5246da00-04c3-45f3-8e9a-2f135549850b · outbound

This paper cites Computing geodesics on triangular meshes.Comput- ers & Graphics, 29(5):667–675, 2005.

GeoMM: On Geodesic Perspective for Multi-modal Learning Computing geodesics on triangular meshes.Comput- ers & Graphics, 29(5):667–675, 2005

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.042868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.665984Z digest=sha256:b0b2169bc3a80cda154fe9462d1f56a9e9f975d331fc586436f66f4b194d7cbc

Observation b5cd8895-692e-4312-9a36-eb02c979757c · outbound

This paper cites Bron- stein, and Pierre Vandergheynst.

GeoMM: On Geodesic Perspective for Multi-modal Learning Bron- stein, and Pierre Vandergheynst

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.027471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.670412Z digest=sha256:72b4c95f4907f2caf930b4203005bef97e1c55b19e3dc8f9066eb93b5a678b49

Observation db0dae6f-1b92-491a-9eec-622b4be87742 · outbound

This paper cites Jensen’s inequality.

GeoMM: On Geodesic Perspective for Multi-modal Learning Jensen’s inequality

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.011409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.674622Z digest=sha256:4a4f3e82bee34945fec78f50c9075a54e39e928bdc823ea90b43a466df53e488

Observation 53b6afc6-3b03-4001-99e5-aa01f5ddb3f2 · outbound

This paper cites Towards bridging sample complexity and model capacity.

GeoMM: On Geodesic Perspective for Multi-modal Learning Towards bridging sample complexity and model capacity

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.995610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.679359Z digest=sha256:b48b023a702b5108182a314b5f7d126f5aabc35d6f74921ddd41fe254fb5f93a

Observation 037bcfe0-7300-451f-8c0e-90acd50e9d26 · outbound

This paper cites Towards interpreting and utiliz- ing symmetry property in adversarial examples.

GeoMM: On Geodesic Perspective for Multi-modal Learning Towards interpreting and utiliz- ing symmetry property in adversarial examples

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.980284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.684246Z digest=sha256:a05101f7e3ff72eddfb05c737a66ce1d513e39fc3bfb03b8ef3c0238ffcc7a0f

Observation eb72ad7e-e4f0-4cb7-ac55-d7d92f501aee · outbound

This paper cites Exploring and utilizing pattern imbal- ance.

GeoMM: On Geodesic Perspective for Multi-modal Learning Exploring and utilizing pattern imbal- ance

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.965557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.688557Z digest=sha256:931be26e96e503ca39a0300f8e48f83f24e08690097ecef5d397579aa12c5b7d

Observation fc249f86-bb9e-4169-b426-df1f88ff3b4b · outbound

This paper cites MSSIDD: A Benchmark for Multi-Sensor Denoising.

GeoMM: On Geodesic Perspective for Multi-modal Learning MSSIDD: A Benchmark for Multi-Sensor Denoising

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.693131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.693131Z digest=sha256:e0f4e4d7d32c27657a0b49b9fc2a643eabeb5ecd7dceb7ef29aa2cf7c518907f

Observation 8667fbc0-cf70-4172-a709-7e11fe756667 · outbound

This paper cites Object- oriented anchoring and modal alignment in multi- modal learning.

GeoMM: On Geodesic Perspective for Multi-modal Learning Object- oriented anchoring and modal alignment in multi- modal learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.950715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.697887Z digest=sha256:a9b1528dc2310d46e1a092f99e2f794f3623f8a0902cf9a43791fff25e9eb922

Observation ab70750a-7f52-41d8-b35d-9c13d9cf1988 · outbound

This paper cites an unresolved cited work.

GeoMM: On Geodesic Perspective for Multi-modal Learning Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:01:39.936600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.702453Z digest=sha256:c74347aa4d0baf8f7aebd5f84b983d6eab1a8fc42ed6c50d919d78b57c02bb84

Observation dfa2bc74-bcff-4279-9090-48d22c9a014a · outbound

This paper cites Analytic inequalities.

GeoMM: On Geodesic Perspective for Multi-modal Learning Analytic inequalities

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.921950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.706836Z digest=sha256:1b87050590ae03194542862b80e36d827a648e9a676590d77488b4b076487071

Observation 779b1cb0-4418-4c9a-b5ee-c5bbc4d44e24 · outbound

This paper cites Slip: Self-supervision meets language- image pre-training.

GeoMM: On Geodesic Perspective for Multi-modal Learning Slip: Self-supervision meets language- image pre-training

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.907114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.711315Z digest=sha256:2d96cf816c7e6d20360946f65cd4c63cb6a215f4a3b8cd76212d3a808f2c0019

Observation fa827551-edcc-4c9e-9912-3714a5411439 · outbound

This paper cites Geodesic-former: A geodesic-guided few-shot 3d point cloud instance segmenter.

GeoMM: On Geodesic Perspective for Multi-modal Learning Geodesic-former: A geodesic-guided few-shot 3d point cloud instance segmenter

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.891302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.715810Z digest=sha256:5e03ea8d4c2c1c761d55d876844f0deff55c72f9d7b34fac54928c06fb4e3f50

Observation 365b3a40-91df-420d-8e36-a9d33cf78743 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

GeoMM: On Geodesic Perspective for Multi-modal Learning Representation Learning with Contrastive Predictive Coding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.720127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.720127Z digest=sha256:69b7755bbb67e4b5fe4e9cd650a48e5a2a94b4e4015e19e1a31c4478237f03fb

Observation 02cfc387-40af-4776-aaca-1ce14fa83ce9 · outbound

This paper cites Im2text: Describing images using 1 million cap- tioned photographs.

GeoMM: On Geodesic Perspective for Multi-modal Learning Im2text: Describing images using 1 million cap- tioned photographs

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.874796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.724243Z digest=sha256:d58748449454694cb1b8da50454c463cfd54680a6a89099c0c45ff80faef7d66

Observation c050648a-b3e4-47df-b98c-50d09c10e2f9 · outbound

This paper cites Automatic differentiation in pytorch.

GeoMM: On Geodesic Perspective for Multi-modal Learning Automatic differentiation in pytorch

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.860351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.728504Z digest=sha256:49b14370cadf0cbfc2bc4ee93c9468a64f1a2ab6d4a4566ab5ce29d7e05b8fa6

Observation cdca1922-b4c5-44fe-b249-9b550c72b70f · outbound

This paper cites BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers.

GeoMM: On Geodesic Perspective for Multi-modal Learning BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.732801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.732801Z digest=sha256:7d73ad4fc971a1c9ba04a6739fd91656c03b6ee90a1ce5811e0628f9ba695478

Observation 963c7e38-774e-441b-a3ab-26fa8a413f82 · outbound

This paper cites Computational optimal transport: With applications to data science.

GeoMM: On Geodesic Perspective for Multi-modal Learning Computational optimal transport: With applications to data science

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.845941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.737328Z digest=sha256:f2c719117fe59ed64d415f96c3f4b827a128bb5542a715bdffe03293bfea5f71

Observation b1c8fd83-b920-4b38-bd53-203dd2f7e424 · outbound

This paper cites Combined scaling for zero-shot transfer learn- ing.

GeoMM: On Geodesic Perspective for Multi-modal Learning Combined scaling for zero-shot transfer learn- ing

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.831296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.741852Z digest=sha256:69061560ede88383ffdd140a0ce3b1dfb5181897f210235ec0ec0e9dbd006c8b

Observation a45f9e51-ef3b-4c23-9477-b6dc223754f4 · outbound

This paper cites Flickr30k entities: Collecting region-to- phrase correspondences for richer image-to-sentence models.

GeoMM: On Geodesic Perspective for Multi-modal Learning Flickr30k entities: Collecting region-to- phrase correspondences for richer image-to-sentence models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.816611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.746508Z digest=sha256:a136ff74072e309b718269bf862863ea7e149c00b31642ae7d43f0e49fa08e8e

Observation 3aa8b2b5-f6c8-432a-bfad-9389940629e9 · outbound

This paper cites Straightest geodesics on polyhedral surfaces.

GeoMM: On Geodesic Perspective for Multi-modal Learning Straightest geodesics on polyhedral surfaces

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.800960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.751231Z digest=sha256:db457667d6fc71bba73d2e1fb56b75e5d7c783cea0bb5e6690243913b9e80ec3

Observation 69d4cb8f-1a1c-4d7f-beaf-ea39ec9da25a · outbound

This paper cites Graphwalks: Efficient shape agnostic geodesic shortest path estimation.

GeoMM: On Geodesic Perspective for Multi-modal Learning Graphwalks: Efficient shape agnostic geodesic shortest path estimation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.786494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.755787Z digest=sha256:5f74533f518ceeab2166c4aee5eae9b04aea431bd0e7242edd6f472e1864a945

Observation 83ece59c-1713-4f08-9db8-d05488a96e43 · outbound

This paper cites ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data.

GeoMM: On Geodesic Perspective for Multi-modal Learning ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.760337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.760337Z digest=sha256:830dabaad734a56b518e33427288245a8335fecdd31fc3b3c31659ee337ed101

Observation b91a6b2b-8976-4bc4-9df4-028ed954f459 · outbound

This paper cites Improving language under- standing by generative pre-training.

GeoMM: On Geodesic Perspective for Multi-modal Learning Improving language under- standing by generative pre-training

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.771893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.764854Z digest=sha256:34cd955554ecff2df861de1c1cb54a7335442f675ef1004c2fc2174e64b586b6

Observation cf238571-3c66-4790-a28d-b0db7bb9eb3c · outbound

This paper cites Learning transferable visual models from natural language supervision.

GeoMM: On Geodesic Perspective for Multi-modal Learning Learning transferable visual models from natural language supervision

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.756264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.770102Z digest=sha256:9f80c0f69d795b1103b0827ee8fd550be575de15fb03a9ed0959aeb70d624ec5

Observation 34b842d2-b60f-445a-a601-22e99fb888e9 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

GeoMM: On Geodesic Perspective for Multi-modal Learning Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.739955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.774550Z digest=sha256:b7233923db89c6e8d91f09f5e2c3c4a6d747d575c021a123216303bd22ad9c27

Observation 8da1e150-ae73-490d-a9c7-3157e38dbdf1 · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

GeoMM: On Geodesic Perspective for Multi-modal Learning Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.724753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.779246Z digest=sha256:e38332d9ceec31bfe76731bb2270214b004d5cf7411294698a25d2194e4d47f3

Observation 71a53bd5-3456-4d4e-aa00-1ceb46daee9c · outbound

This paper cites Conceptual captions: A cleaned, hy- pernymed, image alt-text dataset for automatic image captioning.

GeoMM: On Geodesic Perspective for Multi-modal Learning Conceptual captions: A cleaned, hy- pernymed, image alt-text dataset for automatic image captioning

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.709325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.783626Z digest=sha256:61d2587a833830b76f3fcc23d8875887b61230b82d8f3fc84202ca9cd514a1b6

Observation ef6e53a7-c2d1-4362-8650-182f8a060ede · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

GeoMM: On Geodesic Perspective for Multi-modal Learning VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.788840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.788840Z digest=sha256:5aca511454fe398558a5a98ab98fb70e64a1fc5348f633e3356ae4ee2d09abcf

Observation 37bb562b-fa2e-443f-b0de-ce00af6ef2dd · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

GeoMM: On Geodesic Perspective for Multi-modal Learning PandaGPT: One Model To Instruction-Follow Them All

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.793508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.793508Z digest=sha256:b71cbc200cbdbdabfed063306a7df851daf98ec31e309b0a123703a9d26e04f1

Observation 2d304298-9594-4627-bfbf-0b7daf9418fe · outbound

This paper cites A Corpus for Reasoning About Natural Language Grounded in Photographs.

GeoMM: On Geodesic Perspective for Multi-modal Learning A Corpus for Reasoning About Natural Language Grounded in Photographs

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.798297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.798297Z digest=sha256:69dfbb3e88294edfce14f29b41f41a8ddf58e1789575be9f32acdda3ad52d6bc

Observation 140a7beb-1be5-4a53-ad80-278df7e9ca50 · outbound

This paper cites Revisiting unreasonable effective- ness of data in deep learning era.

GeoMM: On Geodesic Perspective for Multi-modal Learning Revisiting unreasonable effective- ness of data in deep learning era

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.694704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.802864Z digest=sha256:3fa6af007ccfb3e10dbe4827a171d6504718f16dffae04387df2a8d314294caf

Observation 80e6a150-979d-4e34-96b5-8e289fb4060b · outbound

This paper cites Gortler, and Hugues Hoppe.

GeoMM: On Geodesic Perspective for Multi-modal Learning Gortler, and Hugues Hoppe

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.680072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.807626Z digest=sha256:91fd64a8dad9323878fe58e0f74d518ffc6c609b45a8c535965d427d7219c4e4

Observation 71871314-e8da-4eb1-821a-1a4f8785f83e · outbound

This paper cites LXMERT: learning cross-modality encoder representations from trans- formers.

GeoMM: On Geodesic Perspective for Multi-modal Learning LXMERT: learning cross-modality encoder representations from trans- formers

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.664840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.812216Z digest=sha256:1cad4dd16e2d510d53b40303da197a9e1836ab468ff00ff4f99da3575cbaaa3c

Observation 06bf8c74-6ea4-4cce-afc8-ea3c2782f5ee · outbound

This paper cites Tenenbaum, Vin de Silva, and John C.

GeoMM: On Geodesic Perspective for Multi-modal Learning Tenenbaum, Vin de Silva, and John C

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.650272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.816540Z digest=sha256:c4c1957ee63cfb4a2f30355484ba588a3de7387dc6343f9f55e97554a33a313b

Observation 65913de9-b905-4a15-910a-51ccf44688da · outbound

This paper cites Pigeon hole principle.

GeoMM: On Geodesic Perspective for Multi-modal Learning Pigeon hole principle

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.635830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.821250Z digest=sha256:bd6130f6013a3ce5239beb174f9b3c8b941e0466ebc242259e53a9c0ec5acdc2

Observation 403d64ff-a537-4552-8f49-4041e217dbac · outbound

This paper cites Attention is all you need.

GeoMM: On Geodesic Perspective for Multi-modal Learning Attention is all you need

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.621277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.826010Z digest=sha256:dcad6d218b84ad0b755cc529bf0eaf58dcc677e7cd39bf5e98f40159af3bc9ee

Observation db607275-2142-4af8-8dd8-14cb766e6385 · outbound

This paper cites Optimal transport: old and new.

GeoMM: On Geodesic Perspective for Multi-modal Learning Optimal transport: old and new

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.606340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.831030Z digest=sha256:58bd7a49940b772927164aba569d4d5d2c6ecb7e353550f3fbce7c0351ea1269

Observation 2dcf3c28-e99f-4ec3-ae7c-e8c33c20b304 · outbound

This paper cites Learning to combine: Knowledge aggrega- tion for multi-source domain adaptation.

GeoMM: On Geodesic Perspective for Multi-modal Learning Learning to combine: Knowledge aggrega- tion for multi-source domain adaptation

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.591726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.835846Z digest=sha256:7171c192a6fe7bbffed7dd89675db0ec58585aec25b7c63b355f2c8e09a2bd4e

Observation 7c0b0f61-e2b7-46d9-bb19-459480b1a9e1 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

GeoMM: On Geodesic Perspective for Multi-modal Learning Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.840420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.840420Z digest=sha256:27b8efb579cb7e1849ad46d33ccf68607e2ec80e183c7661b6fe126bfb93a278

Observation 55b012a0-b38b-49f9-b321-51c4f2c794b8 · outbound

This paper cites Mvp: Multimodality-guided visual pre-training.

GeoMM: On Geodesic Perspective for Multi-modal Learning Mvp: Multimodality-guided visual pre-training

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.577128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.845368Z digest=sha256:3ca315944562522d96f1e50a658e253be4244e56d9c2db6e12dab81ad218c23d

Observation deaa2c2d-f7ac-4aec-a500-d1e3e843a593 · outbound

This paper cites Visual Entailment: A Novel Task for Fine-Grained Image Understanding.

GeoMM: On Geodesic Perspective for Multi-modal Learning Visual Entailment: A Novel Task for Fine-Grained Image Understanding

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.849934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.849934Z digest=sha256:30c995c208b4da3fd6d15164c0e88fb77f4c790e8a075e3dab928f9adff31f90

Observation 35dd7e54-89fe-4315-8fb4-62c1fc13119e · outbound

This paper cites A fast proximal point method for computing exact wasserstein distance.

GeoMM: On Geodesic Perspective for Multi-modal Learning A fast proximal point method for computing exact wasserstein distance

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.562192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.854916Z digest=sha256:bf12459e66106197f1cb5974216d73840ec0deca6ec9294b7d4608902b34d5a8

Observation 9f1ffb73-0c4f-45eb-8907-460c0cd33c99 · outbound

This paper cites Vision-language pre- training with triple contrastive learning.

GeoMM: On Geodesic Perspective for Multi-modal Learning Vision-language pre- training with triple contrastive learning

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.547744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.859498Z digest=sha256:37dcc6b79ac6a5dee674aa27c67cee1e5adaf004ffbd3a3d8a317844c5d21be5

Observation 89478e07-0d26-4164-881d-72924a471f13 · outbound

This paper cites Unified contrastive learning in image-text-label space.

GeoMM: On Geodesic Perspective for Multi-modal Learning Unified contrastive learning in image-text-label space

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.533274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.864421Z digest=sha256:342eade588f7235c22f7adc3b4357a0b1db750a2ee6beb76f3762b1b5a2ad0f3

Observation 957b594d-6f13-47fe-96d9-93b44a7a7eb9 · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

GeoMM: On Geodesic Perspective for Multi-modal Learning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.868708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.868708Z digest=sha256:e3b5267711c6c9c5ef4b0d344fa28cbc7b3c3becff11cdce6cd8ffadceb82d22

Observation 5c558f01-be97-4bf8-aa03-68cd271fc20d · outbound

This paper cites FILIP: fine-grained interactive language-image pre-training.

GeoMM: On Geodesic Perspective for Multi-modal Learning FILIP: fine-grained interactive language-image pre-training

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.518572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:01:38.873663Z digest=sha256:e972a1768fc7c113f95c2e4258e9a30cc5f639b4593d16bd36cc9b67e729ea13

Observation 132e1aec-03d9-42ba-a78b-d52284e19a45 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

GeoMM: On Geodesic Perspective for Multi-modal Learning CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.878311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.878311Z digest=sha256:89eb189b85b40ee5bfb28bfe244649edae0b5a5ef008aa770de79bf2ceeeee7c

Pith citing papers

No inbound Pith citation observations are available.