Pith. sign in

Paper Citation Record · LEDGER

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text

As of 7 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2508.00447.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.00447 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:10:57.654052Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 149c51b0-aecb-49d5-a262-ce0a338be6f0 · outbound

This paper cites Large Language Models: A Survey.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Large Language Models: A Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.566045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.566045Z digest=sha256:587860df7c65cf6964cce3b7cbe4fcb95a337255943847fcb0d246e9812db46e

Observation dd2710f2-4dd6-469b-9396-5ac2e867eca4 · outbound

This paper cites Learning transferable visual models from natural language supervision.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Learning transferable visual models from natural language supervision

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.570247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.570247Z digest=sha256:821711cf093306cb617a56cf27043d7e508cf8c7c604258c84102736e88d00ed

Observation ab1a973f-5791-4f81-af60-63ef50eb0d9b · outbound

This paper cites Vlm agents generate their own memories: Distilling experience into embodied programs of thought.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Vlm agents generate their own memories: Distilling experience into embodied programs of thought

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.912864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.574095Z digest=sha256:65bc2836fb5bc86b6a9043899320b983a54555789b547387c2550df7e17ff7c1

Observation ba806b59-659a-4982-92a8-73a7bd038118 · outbound

This paper cites Rethinking vlms and llms for image classification.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Rethinking vlms and llms for image classification

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.904957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.577567Z digest=sha256:75e5b6c26d461086f0a7d01a2a97b038ff5a2648cd7cd046827abd6003c74725

Observation 84c0b498-b7a4-40ae-893a-5076fbf49f16 · outbound

This paper cites Regression in EO: Are VLMs Up to the Challenge?.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Regression in EO: Are VLMs Up to the Challenge?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.580373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.580373Z digest=sha256:349d17e2a8b3f102e8e896dbba08c9370f0ef12b762708634d00acb2a40e61b8

Observation 37c8c514-ea71-4243-888f-106e2da05f68 · outbound

This paper cites From Images to Signals: Are Large Vision Models Useful for Time Series Analysis?.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text From Images to Signals: Are Large Vision Models Useful for Time Series Analysis?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.584389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.584389Z digest=sha256:236db0035b75a2ffd9ea9d1edfa22932eb05967b45c5bc864889b9ba62de6123

Observation 06d4ebcc-338b-4ed5-8fcd-de514be79caa · outbound

This paper cites Multi-step time series forecasting with an ensemble of varied length mixture models.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Multi-step time series forecasting with an ensemble of varied length mixture models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.896346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.588226Z digest=sha256:8095a51f7570fbfa3e7b304e0cd6b2aad39cfa2f922575265ea9f0e99261bd25

Observation 88a1cdb1-3f66-4c60-9aaf-442fd10cca4d · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text CogVLM2: Visual Language Models for Image and Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.591421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.591421Z digest=sha256:834a91759b1f8f8c1fef90ea018c6c9a1c600dff0294d91703f3541eae8ef312

Observation fefd3627-07d8-4bf1-b645-0d9866c49b91 · outbound

This paper cites Integrating llms with its: Recent advances, potentials, challenges, and future directions.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Integrating llms with its: Recent advances, potentials, challenges, and future directions

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.889206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.594638Z digest=sha256:04e8ac8682e6f6d3a168ae7a141e07862dcd5518b0fdd1bb6a5dcd3536ef9325

Observation ad8958b7-ae7b-4f20-9d7c-4532d0a30b87 · outbound

This paper cites iTransformer: Inverted Transformers Are Effective for Time Series Forecasting.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text iTransformer: Inverted Transformers Are Effective for Time Series Forecasting

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.597170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.597170Z digest=sha256:3bd8953063a59428f3d297a17d789af0e19cbc564013223eeb86b14311d18da2

Observation 85e09ceb-a398-446b-bfe7-f9155f22000c · outbound

This paper cites PV-VLM: A Multimodal Vision-Language Approach Incorporating Sky Images for Intra-Hour Photovoltaic Power Forecasting.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text PV-VLM: A Multimodal Vision-Language Approach Incorporating Sky Images for Intra-Hour Photovoltaic Power Forecasting

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.600421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.600421Z digest=sha256:211bd562d7ef5ec2008ce10e6a54c5940820be88dca77138069970c61b76a6b9

Observation 708f493d-fa75-4949-879b-c5ed5b98eb1a · outbound

This paper cites Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.603925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.603925Z digest=sha256:0c1a6b656411ef15b0c96db1a92aed2b3bb7e2f653d294e469de7fc854335b23

Observation 9a3baa57-ccf5-416c-b4aa-8ffdbb9b11d3 · outbound

This paper cites Leveraging temporal contextualization for video action recognition.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Leveraging temporal contextualization for video action recognition

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.880538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.607461Z digest=sha256:904ff9520252207c29af6cf6df9b89096b8e16ac4e54ae42f537983e3d6c0829

Observation 200ee217-6049-461d-97c3-20ac5323b43e · outbound

This paper cites C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.611080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.611080Z digest=sha256:2ed9a226276a6e9cea5d83d85caf4321a0332854021cc67c7502991e971f921a

Observation 3af5bf5a-1e2a-4926-ac49-757828e85738 · outbound

This paper cites Pattern recognition and prediction in time series data through retrieval- augmented techniques.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Pattern recognition and prediction in time series data through retrieval- augmented techniques

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.871564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.614083Z digest=sha256:763419460253df27e57dfa342d0f1e99e20db411c605708e0e17e71be2185879

Observation 30b41a4b-27b0-49a2-8b9e-55a1c9f526df · outbound

This paper cites Driver intention prediction using text prompts with in-cabin and out-cabin cameras.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Driver intention prediction using text prompts with in-cabin and out-cabin cameras

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.862058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.617359Z digest=sha256:b85abba20135fd841c443a40d2683a07c2ec7a132e02a06e93eac4c9a40894f9

Observation 02535a6d-e486-452a-b1d2-2551c8154a34 · outbound

This paper cites Diverse data augmentation with dif- fusions for effective test-time prompt tuning.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Diverse data augmentation with dif- fusions for effective test-time prompt tuning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.851547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.620833Z digest=sha256:d1ae8108435477a771411cbb40b286781e7baa7cbd23485ff55546d1e16214ba

Observation 2a3e5744-844b-4eea-a6ae-4b2e99768ebb · outbound

This paper cites FungalZSL: Zero-Shot Fungal Classification with Image Captioning Using a Synthetic Data Approach.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text FungalZSL: Zero-Shot Fungal Classification with Image Captioning Using a Synthetic Data Approach

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:10:57.692926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.623645Z digest=sha256:ee1e499f2ffc1162d3e6ee65494a651db82bf16d5000e96da580218b55a6163c

Observation e39d859c-8326-4a61-ad81-65683ad4de9a · outbound

This paper cites Fungal biology.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Fungal biology

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.842797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.626549Z digest=sha256:58df4b36d15ab73fb855befbe7801e01601a5d2088019dd179f8135eaa647480

Observation 76d98e21-6e5f-4a84-8f23-c50cecc9dc07 · outbound

This paper cites Assembling the fungal tree of life: progress, classification, and evolution of subcellular traits.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Assembling the fungal tree of life: progress, classification, and evolution of subcellular traits

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.834678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.629001Z digest=sha256:83dcab45f28237e5133103aae263db72c2d6d4c071347ed39180be79901de3fd

Observation 31ccd902-c5bc-4f0a-8d96-7db992310a9c · outbound

This paper cites Developments in fungal taxonomy.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Developments in fungal taxonomy

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.825328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.631540Z digest=sha256:b61884f50d7da43f473f8207b1adb4aa2e9fc329ab2463847a538f9bc8991670

Observation 26edf3e3-5f98-40be-b0df-33f4c42e9440 · outbound

This paper cites Assessment of fungal spores and spore-like diversity in environmental samples by targeted lysis.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Assessment of fungal spores and spore-like diversity in environmental samples by targeted lysis

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.815366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.634084Z digest=sha256:6b77fb0c93bd9d15642ecbcc03196c9dc491d157b974ac360baf6ccdd7f65ab7

Observation 5e4e9dc0-eb84-4df6-b842-fdc7a8f03d09 · outbound

This paper cites Architecture and developmental dynamics of the external mycelium of the arbuscular mycorrhizal fungus glomus intraradices grown under monoxenic conditions.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Architecture and developmental dynamics of the external mycelium of the arbuscular mycorrhizal fungus glomus intraradices grown under monoxenic conditions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.806844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.637069Z digest=sha256:f3222457fcbb9695fc165a0b6ac89d78569d7dc2613a99e2dc07ba670d061d1a

Observation 85fc7581-3e42-4ef4-851f-6102411aadc6 · outbound

This paper cites Study of kinetic model for fungal spore germination under dynamic conditions: Case study on germination of penicillium expansum spores.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Study of kinetic model for fungal spore germination under dynamic conditions: Case study on germination of penicillium expansum spores

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.796630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.640618Z digest=sha256:2fde1be09e1d8c2dd00744073ae4ef2dfaca673ad98f7fde3cb6bf10ade50682

Observation 662c32db-1362-43e2-916a-dff97d1ddc09 · outbound

This paper cites Pathways of pathogenicity: transcriptional stages of germination in the fatal fungal pathogen rhizopus delemar.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Pathways of pathogenicity: transcriptional stages of germination in the fatal fungal pathogen rhizopus delemar

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.787085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.643506Z digest=sha256:a2d5647511ccf5f0aadb4889834df0ffc241ab55eadf443a333bdfb84869eaa4

Observation 1f2ddbcc-5c21-465a-a143-ca725ba058e7 · outbound

This paper cites A model for growth of a single fungal hypha based on well-mixed tanks in series: simulation of nutrient and vesicle transport in aerial reproductive hyphae.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text A model for growth of a single fungal hypha based on well-mixed tanks in series: simulation of nutrient and vesicle transport in aerial reproductive hyphae

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.777381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.645926Z digest=sha256:f27f6fae8a1e24ba1fd036686abca9727bec12441f2165c25ef6b1e1f0c1de57

Observation c75dbd80-d519-4699-9dc0-8a43782189e2 · outbound

This paper cites Neurospora crassa nadph oxidase nox-1 is localized in the vacuolar system and the plasma membrane.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Neurospora crassa nadph oxidase nox-1 is localized in the vacuolar system and the plasma membrane

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.766849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.648290Z digest=sha256:e82f0a2f0fe11753192325a8ef48e9ca99ed15b83d8f46c92604c6222e2eb3c5

Observation 1f1b6b55-5eea-4796-bb2b-4c01b29d5ef3 · outbound

This paper cites Synthetic Fungi Datasets: A Time-Aligned Approach.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Synthetic Fungi Datasets: A Time-Aligned Approach

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:10:57.680562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.650687Z digest=sha256:82a95997822d4e734a3c8affabfaccdbf7d56816e21a9de60b3b19e132ee0b98

Observation 204e3eb2-1690-49f5-bbd4-d825d233d203 · outbound

This paper cites Fungalzsl: A fine-grained synthetic dataset for zero-shot fungi classification.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Fungalzsl: A fine-grained synthetic dataset for zero-shot fungi classification

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.756848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:10:57.654052Z digest=sha256:c5c798a15cea9fcb2b17fb903fe083a32b4e4aed70fbaf631141c224679165a8

Pith citing papers

No inbound Pith citation observations are available.