Pith. sign in

Paper Citation Record · LEDGER

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks

As of 18 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2506.03174.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03174 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:00:57.269520Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T18:48:40.813486Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T18:53:08.558584Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 16a36a81-077d-46e7-a36f-4083957d07ed · outbound

This paper cites Imu2doppler: Cross-modal domain adaptation for doppler-based activity recogni- tion using imu data.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Imu2doppler: Cross-modal domain adaptation for doppler-based activity recogni- tion using imu data

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:00.850539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:54.638915Z digest=sha256:4aff1a911533b9a7579e66280b3f19b253cf4c939c98d61961939908610bb4c5

Observation 8bff663f-f3eb-43fa-b3b0-4ef0b53668d9 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks On the Opportunities and Risks of Foundation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:54.721113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:54.721113Z digest=sha256:388dc001dba4b1292c651ce5f1e33d87e2992701fa00f1d1a96b5c4852d301a1

Observation 1b25788c-f478-4172-92fd-1146224a4682 · outbound

This paper cites Language models are few-shot learners.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Language models are few-shot learners

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:00.755749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:54.905897Z digest=sha256:60a4dd221d1c74ec1983869fa7ef4b7310a1841de125cd9082488c121528a7d0

Observation c3f689f4-cc84-4fd9-be40-1aa9686ba0cc · outbound

This paper cites A simple framework for contrastive learning of visual rep- resentations.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks A simple framework for contrastive learning of visual rep- resentations

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:00.630434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:54.994473Z digest=sha256:7bc78fe463d260141525c4e5cb4daa8277fbd23fa778da4214d5e0846216e86b

Observation 68343256-f6bf-4e42-84e0-2d6f2257020b · outbound

This paper cites Cocoa: Cross modality contrastive learning for sensor data.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Cocoa: Cross modality contrastive learning for sensor data

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:00.504764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:55.164807Z digest=sha256:8d01047c74700a2100bc2bbb2e0aaaea4745868504f46dac196b9c6279749972

Observation 8b01799c-bc06-48b2-bc2a-217120cae277 · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks BERT: Pre-training of deep bidirectional transformers for language understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:00.389718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:55.333310Z digest=sha256:43e3742139dd49a16beb0deaeca6fd211e544ca1495e665af69575cf5409769f

Observation 891f8e61-de52-4458-bb5d-6f8b6a13cf41 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks An image is worth 16x16 words: Transformers for image recognition at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:00.283744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:55.467050Z digest=sha256:54e33497a51653c68b0a0ca32ffbe10849b4cf406f79990d40c4351343e7a24d

Observation 41046ee5-395c-46a1-8b43-a8a24aefe8de · outbound

This paper cites Unsupervised representation learning by predicting image rotations.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Unsupervised representation learning by predicting image rotations

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:00.153940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:55.592184Z digest=sha256:73bfdea22cb1a7a9e214396009362b99d2ffe5ac757f30fb3484e12d7eb299b5

Observation dcd26bd9-5bc2-4823-a973-f76ed14f6b48 · outbound

This paper cites Image- bind: One embedding space to bind them all.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Image- bind: One embedding space to bind them all

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:00.054850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:55.771780Z digest=sha256:048913adae1a2afbeebecaef527644110cdd73b0520b9b4235a597eb2fd71e4f

Observation ac55b056-73d0-4946-a78a-6673ca449ed9 · outbound

This paper cites Ego4d: Around the World in 3,000 Hours of Egocentric Video.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Ego4d: Around the World in 3,000 Hours of Egocentric Video

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:59.932890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:55.908690Z digest=sha256:9a2a0f701397121ae53d6fbefb230e888516cd5e31cadf15bf3c6005d6c42ca9

Observation 4d119b10-d556-433c-9b7a-9e3298c30105 · outbound

This paper cites Bikram Boote.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Bikram Boote

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:59.802040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:56.014994Z digest=sha256:e7a83362eb8b7b39c3363345c7292381d0453c96211efe610dcce3884e750e19

Observation 32cc8468-8017-4f3d-a36d-83beeb0ac547 · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Momentum contrast for unsupervised visual representation learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:59.678019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:56.134192Z digest=sha256:50d7370b5c9fc70e953d86a6d4bc61c7b12e5701cfa7716425ae3b09a5ede647

Observation f74cddfb-b989-4817-b9a7-ad8596a693a9 · outbound

This paper cites Abowd, Nicholas D.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Abowd, Nicholas D

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:59.480671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:56.257477Z digest=sha256:1367cdc8e72ca92620a019ab465882c5d445958d352000292ae4058e4e5a3f7d

Observation 18cf59d9-431f-46fe-9bab-935f3ec5e0f9 · outbound

This paper cites Imu2clip: Language- grounded motion sensor translation with multimodal contrastive learning.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Imu2clip: Language- grounded motion sensor translation with multimodal contrastive learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:59.268626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:56.403102Z digest=sha256:7efc948fd5c3ec15095544557d65abd086a95ef344af18d80a94f03b5c5a106c

Observation b69b3137-9803-4b1e-be10-d40b09221670 · outbound

This paper cites Unsupervised learning of visual representations by solving jigsaw puzzles.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Unsupervised learning of visual representations by solving jigsaw puzzles

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:59.092399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:56.521970Z digest=sha256:ef5fc42d228875d2246469f060fc7d696842cee353285014ae4998e13612a791

Observation a2d4a8ff-dde6-4724-abf4-ec026832e11e · outbound

This paper cites Berkeley mhad: A comprehensive multimodal hu- man action database.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Berkeley mhad: A comprehensive multimodal hu- man action database

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:58.953308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:56.611697Z digest=sha256:4a956577feb6d45474dd554ce40c4f55669f8d3b55024e560423d3ca8c4becf3

Observation 49b5eea2-86a3-462f-9330-259b650d4555 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Representation Learning with Contrastive Predictive Coding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:56.725182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:56.725182Z digest=sha256:b3338392ccabcee31a785fb49e355849d068fc920d71a0ffe17bb26595c9ef75

Observation 83348343-c0bb-4265-9d8a-5b56c66c07a1 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Learning transferable visual models from natural language supervi- sion

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:58.748914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:56.873258Z digest=sha256:3fe720bafe3a4e9fa35c879ca4dd6e03685cb572c621875cfe5a560d2d09e646

Observation 39090540-1c75-47bf-b82a-a74a8a7cd8d0 · outbound

This paper cites Introducing a new benchmarked dataset for activity monitoring.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Introducing a new benchmarked dataset for activity monitoring

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:58.322693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:57.039481Z digest=sha256:ff6f1abb79416fa0daa7e265002a6c9422271e30599cd45ad7e6b1f0e298a591

Observation e37e35de-baae-4f25-b6c2-d3de0b12a788 · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Facenet: A unified embedding for face recognition and clustering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:58.093811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:57.144155Z digest=sha256:edfde717982f9adf19bbffbffb85cc6023482188a19ef9957f8527e32dae1064

Observation 3665b6c9-fc31-4d2f-b6b7-d7faa8705c47 · outbound

This paper cites an unresolved cited work.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:00:57.842932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:57.189278Z digest=sha256:f44c481382e532146296fa847bb905f9e91b95bba6ae3356151694d7cb455c0f

Observation 89a62ae4-eec6-457f-8d73-b4eef2a6e5cd · outbound

This paper cites Atten- tion is all you need.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Atten- tion is all you need

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:57.583075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:57.269520Z digest=sha256:c5e2843a9d76a30f4dee4635d86579c4dd40c00c64b7e91e86248c16a7506f91

Observation 402185f3-1984-4164-915a-eb865fc894f0 · outbound

This paper cites Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks IJABC: International Journal of Activity and Behavior Computing 25.

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks IJABC: International Journal of Activity and Behavior Computing 25

Reference 8763

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:58.539978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:00:56.971111Z digest=sha256:7fb2f79cec26f2bbeafbd8433bfa2d65bb31c81b9a2c099a398072cf94be357e

Pith citing papers

Observation b21a931f-3449-4056-aa31-7922c9c1bef5 · inbound

Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And Outlook cites this paper.

Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And Outlook Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:53:08.561045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T18:48:40.813486Z digest=sha256:a80fde4e399867e3990f18750e2749c22a6f4ae2d6ca48329b0ec11bdbd65820