Pith. sign in

Paper Citation Record · LEDGER

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks

As of 8 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2506.03174.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03174 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:00:57.269520Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T18:48:40.813486Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T18:53:08.558584Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 16a36a81-077d-46e7-a36f-4083957d07ed · outbound

This paper cites Imu2doppler: Cross-modal domain adaptation for doppler-based activity recogni- tion using imu data.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Imu2doppler: Cross-modal domain adaptation for doppler-based activity recogni- tion using imu data

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:00.850539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:54.638915Z digest=sha256:915a55dd1affd01c05c3009a8b22c2ffa6c435fae70fad2e47d3c038efa4f82e

Observation 8bff663f-f3eb-43fa-b3b0-4ef0b53668d9 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks On the Opportunities and Risks of Foundation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:54.721113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:54.721113Z digest=sha256:aa76c70111adc1566aabeaf259c311de7cb145afbf4030fb4c5aab7e6ad734a1

Observation 1b25788c-f478-4172-92fd-1146224a4682 · outbound

This paper cites Language models are few-shot learners.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Language models are few-shot learners

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:00.755749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:54.905897Z digest=sha256:0734c28ac284762ecdb1f0f18198291e24a6b5d9b400a785ad5ccedebd2b606e

Observation c3f689f4-cc84-4fd9-be40-1aa9686ba0cc · outbound

This paper cites A simple framework for contrastive learning of visual rep- resentations.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks A simple framework for contrastive learning of visual rep- resentations

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:00.630434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:54.994473Z digest=sha256:26e4f4063b07429cdb20b1ddfed90ae89a5f40cb1e7f9e1d947c12108c538476

Observation 68343256-f6bf-4e42-84e0-2d6f2257020b · outbound

This paper cites Cocoa: Cross modality contrastive learning for sensor data.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Cocoa: Cross modality contrastive learning for sensor data

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:00.504764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:55.164807Z digest=sha256:57bd79290102e1cc4a3766200ad8faca5abf57deeac2dd8972cc59dfdc4b6ae2

Observation 8b01799c-bc06-48b2-bc2a-217120cae277 · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks BERT: Pre-training of deep bidirectional transformers for language understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:00.389718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:55.333310Z digest=sha256:8da8ed015be2d1e7867d7136b4aa9c665fd9d1d458ccb2a130193ce0cd172712

Observation 891f8e61-de52-4458-bb5d-6f8b6a13cf41 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks An image is worth 16x16 words: Transformers for image recognition at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:00.283744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:55.467050Z digest=sha256:e36e48e491ff1eba73e97dcdbb2f563d672a5a64a88321749da486b360d16229

Observation 41046ee5-395c-46a1-8b43-a8a24aefe8de · outbound

This paper cites Unsupervised representation learning by predicting image rotations.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Unsupervised representation learning by predicting image rotations

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:00.153940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:55.592184Z digest=sha256:66d58e219f2c052751bff0230bfe75ba4a1797a1f2e3f126297bd401b842fd83

Observation dcd26bd9-5bc2-4823-a973-f76ed14f6b48 · outbound

This paper cites Image- bind: One embedding space to bind them all.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Image- bind: One embedding space to bind them all

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:00.054850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:55.771780Z digest=sha256:7685add5683224d07765199d02bceecbe48b5581e3dfa5da8905ab24636c90a5

Observation ac55b056-73d0-4946-a78a-6673ca449ed9 · outbound

This paper cites Ego4d: Around the World in 3,000 Hours of Egocentric Video.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Ego4d: Around the World in 3,000 Hours of Egocentric Video

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:59.932890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:55.908690Z digest=sha256:c215773f46356cf979aa014b7a1253b7e1566b8262c3a453a481ee486c426804

Observation 4d119b10-d556-433c-9b7a-9e3298c30105 · outbound

This paper cites Bikram Boote.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Bikram Boote

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:59.802040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:56.014994Z digest=sha256:8099af5b20760dffef0c9c47bc0ba83d5e95a58e4d8e6cd6c4a119770dbf4d04

Observation 32cc8468-8017-4f3d-a36d-83beeb0ac547 · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Momentum contrast for unsupervised visual representation learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:59.678019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:56.134192Z digest=sha256:a67606fa90a19013eae53bebf1245871d3ff222ba85ee97cea97ab0bc1cc21ab

Observation f74cddfb-b989-4817-b9a7-ad8596a693a9 · outbound

This paper cites Abowd, Nicholas D.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Abowd, Nicholas D

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:59.480671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:56.257477Z digest=sha256:c5750aff31489b968b1840cca1d3fc1cd72463fc1d5b5d2647878cec7933de38

Observation 18cf59d9-431f-46fe-9bab-935f3ec5e0f9 · outbound

This paper cites Imu2clip: Language- grounded motion sensor translation with multimodal contrastive learning.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Imu2clip: Language- grounded motion sensor translation with multimodal contrastive learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:59.268626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:56.403102Z digest=sha256:ac77c5962a7fdfe4bb451ab74740ed5e4134165a603882338b2ee12345194d25

Observation b69b3137-9803-4b1e-be10-d40b09221670 · outbound

This paper cites Unsupervised learning of visual representations by solving jigsaw puzzles.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Unsupervised learning of visual representations by solving jigsaw puzzles

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:59.092399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:56.521970Z digest=sha256:8a5b85548725f8eee1a8de00f72a7bfb2417f1aeceb3645b5eec50b7dc0ec374

Observation a2d4a8ff-dde6-4724-abf4-ec026832e11e · outbound

This paper cites Berkeley mhad: A comprehensive multimodal hu- man action database.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Berkeley mhad: A comprehensive multimodal hu- man action database

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:58.953308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:56.611697Z digest=sha256:8c6d68e3ad6542d71b56609ba5bc7301ff3d4918362daca89fe1287ef598a2f2

Observation 49b5eea2-86a3-462f-9330-259b650d4555 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Representation Learning with Contrastive Predictive Coding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:56.725182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:56.725182Z digest=sha256:d39a0b8cab7df39b715df924a5789f7e380dc4f1ac7083a105472f4046650e53

Observation 83348343-c0bb-4265-9d8a-5b56c66c07a1 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Learning transferable visual models from natural language supervi- sion

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:58.748914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:56.873258Z digest=sha256:5245b72fc596b3b2b92bac6f6fb529672a6fd10b643910faa65ce44c1f8dca82

Observation 39090540-1c75-47bf-b82a-a74a8a7cd8d0 · outbound

This paper cites Introducing a new benchmarked dataset for activity monitoring.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Introducing a new benchmarked dataset for activity monitoring

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:58.322693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:57.039481Z digest=sha256:30b39a3efe335580457ef549ea3696f2e2b046030558d005427d55b601dd6341

Observation e37e35de-baae-4f25-b6c2-d3de0b12a788 · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Facenet: A unified embedding for face recognition and clustering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:58.093811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:57.144155Z digest=sha256:c207f1e309c6369876910fcb23b026545857a32256304dabf2259827f993ac0f

Observation 3665b6c9-fc31-4d2f-b6b7-d7faa8705c47 · outbound

This paper cites an unresolved cited work.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:00:57.842932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:57.189278Z digest=sha256:5754146800f0dcd27b42022941e6b7c2ba3d241aa330aada9cc172dd1f6dce60

Observation 89a62ae4-eec6-457f-8d73-b4eef2a6e5cd · outbound

This paper cites Atten- tion is all you need.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Atten- tion is all you need

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:57.583075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:57.269520Z digest=sha256:a067c1ec0baeb589e8248b0f89b419ace7f4aea80c46a2e793fa2126a9aff152

Observation 402185f3-1984-4164-915a-eb865fc894f0 · outbound

This paper cites Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks IJABC: International Journal of Activity and Behavior Computing 25.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks IJABC: International Journal of Activity and Behavior Computing 25

Reference 8763

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:58.539978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:56.971111Z digest=sha256:601765e4597c35d89d42757f9dfa158e42791bacc0ac6cf6aff4f7d3ecc0eb07

Pith citing papers

Observation b21a931f-3449-4056-aa31-7922c9c1bef5 · inbound

Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And Outlook cites this paper.

Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And Outlook Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:53:08.561045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T18:48:40.813486Z digest=sha256:9e671bab1dee8b06eec2e32ae6d7dfb717fbd59d9685eff446cd76c581774596