Pith. sign in

Paper Citation Record · LEDGER

UNIV: Unified Foundation Model for Infrared and Visible Modalities

As of 13 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2509.15642.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.15642 v3

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T16:28:48.494939Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T10:27:20.578105Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact9
  • verified fuzzy35
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7b883189-b104-42b9-a5d1-35a47287ce9f · outbound

This paper cites https://github.com/ultralytics/ultralytics.

UNIV: Unified Foundation Model for Infrared and Visible Modalities https://github.com/ultralytics/ultralytics

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:38.564399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:2b137c1fca2a3818348e446771fd8e50eb213bbb15850753e5838b65b9ba2f96

Observation 80dd73f5-dbb2-454c-a8e1-92152eaf5756 · outbound

This paper cites Self-supervised learning from images with a joint-embedding predictive architecture.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Self-supervised learning from images with a joint-embedding predictive architecture

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:38.576943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:29fbc7c182b4fb048cecfc10e11534c22f1e1c0327939c176216e1e42e975174

Observation fc574d61-9aea-4975-b4f1-62b4c9830e81 · outbound

This paper cites End- to-end object detection with transformers.

UNIV: Unified Foundation Model for Infrared and Visible Modalities End- to-end object detection with transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:38.579698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:6340f6e45d1a826449c2957961f9595f12d7122b4bf2398cde8ca5e2591fdd5c

Observation a526a00f-a6b0-46a8-ac96-10b9b0767899 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Emerg- ing properties in self-supervised vision transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:38.582665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:48c06dc7f510a8e757a38f7b85c8d70b7e22eab4c59d35f3d11279812718b769

Observation ff014e28-f218-460a-a748-ec74bab49fcc · outbound

This paper cites Encoder-decoder with atrous separable convolution for semantic image segmentation.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Encoder-decoder with atrous separable convolution for semantic image segmentation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:38.573402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:f258faccccb8e1c1eb424331c2d2e434ccef09f4862c30a62fbe9a94626ca3f7

Observation 7302c942-5950-499b-a37c-041edcc9d880 · outbound

This paper cites An empiri- cal study of training self-supervised vision transformers.

UNIV: Unified Foundation Model for Infrared and Visible Modalities An empiri- cal study of training self-supervised vision transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:38.561437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:0f6366a9208596c1a859efe8dd0b6bf6cdab18f8e9ef03993200f29759603e3e

Observation 5ba99090-640a-4adc-9bab-5f7657fcc4db · outbound

This paper cites Tagging before alignment: Integrating multi- modal tags for video-text retrieval.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Tagging before alignment: Integrating multi- modal tags for video-text retrieval

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:38.570775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:67129610bd55a23b0090e3a0c35a467f981700b4322657757dce5193aa38d6b2

Observation c33a4ad2-9401-4d77-935b-1f9285ffaea0 · outbound

This paper cites Learning a similarity metric discriminatively, with application to face verification.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Learning a similarity metric discriminatively, with application to face verification

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:38.567439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:42b364a653b885e89af35768b84469d27882e248296c26a8ebcff8f29ddacfef

Observation e6dedf02-82dc-41ff-8e79-68d447217abd · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Imagenet: A large-scale hierarchical image database

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:38.558694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:1be18937242e26c73f75724eeac61c7505b532348bc8568caeb794f97601dbd9

Observation 4ff63ccc-9812-40a3-bf1f-492d26fa70d9 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

UNIV: Unified Foundation Model for Infrared and Visible Modalities BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:31:37.143616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:72885e957245c1f7e108345608819dc692a24b56f975b29df815ca75fad29463

Observation 4b2b1fdb-cb07-44e4-ae85-d5fc04451865 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

UNIV: Unified Foundation Model for Infrared and Visible Modalities An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:31:37.180208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:4d8495e3c2699cd5cf77d7aed1b537d6e30ed4fffaf35f62ba2b0ce4485769a4

Observation 95a9f9d2-43b0-43f6-8dd1-3f4dac6969a9 · outbound

This paper cites Flir thermal dataset version 1.3 [dataset].

UNIV: Unified Foundation Model for Infrared and Visible Modalities Flir thermal dataset version 1.3 [dataset]

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.373282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:bfe25f9b86dcbe126d2bf1943f274edc00b9787a67c6d3889b2091961ad0c412

Observation fd48258d-c686-4772-aeea-480187373199 · outbound

This paper cites an unresolved cited work.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-18T16:32:45.391805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:c633d13418c6ac8f28c7e4581b7a7f754f7460d3bb1b31eb07c44217418aeb5f

Observation 4fb88795-19fd-455e-a07d-0bbefa7858d2 · outbound

This paper cites Catastrophic forgetting in connectionist networks.Trends in cognitive sciences, 3(4):128–135.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Catastrophic forgetting in connectionist networks.Trends in cognitive sciences, 3(4):128–135

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.384910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:25a22fb088b54cf468608095d9d0191023ba528db697f96c85fc71eb581e3ee1

Observation 53301122-e33a-4ca9-8e36-8d2968edc846 · outbound

This paper cites ConvMAE: Masked Convolution Meets Masked Autoencoders.

UNIV: Unified Foundation Model for Infrared and Visible Modalities ConvMAE: Masked Convolution Meets Masked Autoencoders

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:31:37.170898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:6e01144e79a287a312bf83129d5c15e2fceaa070f7c925aabfe11ca5046102b1

Observation 3b11516e-3565-4532-81b2-48b21306eb97 · outbound

This paper cites Mfnet: Towards real-time se- mantic segmentation for autonomous vehicles with multi- spectral scenes.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Mfnet: Towards real-time se- mantic segmentation for autonomous vehicles with multi- spectral scenes

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.401635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:38023947d016a2f660308fbd0eee116cc0c0a3a5ae52b0998e01c5ae54ee310c

Observation 129afbed-c491-4414-b7ad-14c51a39ea75 · outbound

This paper cites Mask r-cnn.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Mask r-cnn

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.398520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:69303639322b8faed55983537fbb97ead158fb6b00791fcc2049403f23a42286

Observation 499f363c-40fe-489a-a385-fa2c98a46f06 · outbound

This paper cites Masked autoencoders are scalable vision learners.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Masked autoencoders are scalable vision learners

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.368227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:939b6954d86cef0e5b5f9efcf2aa116d9021875e54b748088dee8099e3876325

Observation d694142c-0bbd-44cf-974d-9455684d00d3 · outbound

This paper cites Vic-mae: Self-supervised representation learning from im- ages and video with contrastive masked autoencoders.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Vic-mae: Self-supervised representation learning from im- ages and video with contrastive masked autoencoders

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.351203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:43835733234193128fdf1b2c702c45132e5d2cb160a002c0357b226ce0f480e0

Observation 539bfa7c-0004-4261-a864-e53e8b78d681 · outbound

This paper cites MILAN: Masked Image Pretraining on Language Assisted Representation.

UNIV: Unified Foundation Model for Infrared and Visible Modalities MILAN: Masked Image Pretraining on Language Assisted Representation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:31:37.148483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:0cdeb052ac29902d008520bc1e5ae2b538fb49b4a69baa8e349d05e10079d91b

Observation 2fc8c05f-8cba-498c-93cf-0c9663f1fd1d · outbound

This paper cites Interlaced Sparse Self-Attention for Semantic Segmentation.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Interlaced Sparse Self-Attention for Semantic Segmentation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:31:37.161675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:8c1fe867831e943962f109303dae3d89b1dc7b8396212a5597d25f7adb1f54db

Observation 670986c7-6e40-4952-979e-7b0a041f4b73 · outbound

This paper cites Multispectral pedestrian detection: Benchmark dataset and baseline.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Multispectral pedestrian detection: Benchmark dataset and baseline

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.354290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:da2292ddc2c68c02c61f9c73cdb2cf7c66140bd94040e55c3cd8a7f618a05005

Observation 0ae2a3c4-5690-49d9-b3b6-153260e0064f · outbound

This paper cites Llvip: A visible-infrared paired dataset for low-light vision.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Llvip: A visible-infrared paired dataset for low-light vision

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.317229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:3fadb85ea28cb65d8bea5d66545ac5beb857901fb30fe42a44050526efdb69dd

Observation a47151a0-7d07-4b6b-b746-a5aa31a36cd9 · outbound

This paper cites Tencent text-video retrieval: hierarchi- cal cross-modal interactions with multi-level representations.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Tencent text-video retrieval: hierarchi- cal cross-modal interactions with multi-level representations

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.357759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:780da03fcee9f9db9035d1a1e1b11b2275fbefb767e90d799ac663bcdeceaba3

Observation 9849a1d1-0c09-485e-b2e2-114b3c0f1c3a · outbound

This paper cites Segmenting objects in day and night: Edge-conditioned cnn for thermal image semantic segmentation.TNNLS, 32(7): 3069–3082.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Segmenting objects in day and night: Edge-conditioned cnn for thermal image semantic segmentation.TNNLS, 32(7): 3069–3082

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.388556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:98b0700c68091473a60a9ab43034ed0664f89dbbe92476575a7ba6e2ff97cb2e

Observation bba935d6-8cb0-4e9c-9254-918d57dc825e · outbound

This paper cites Lasher: A large-scale high- diversity benchmark for rgbt tracking.TIP, 31:392–404.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Lasher: A large-scale high- diversity benchmark for rgbt tracking.TIP, 31:392–404

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.321893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:50de0fe417f378d8a4276d96997f8cb445c4d6e091f8c54054788d336398d355

Observation 977d941b-9127-4cca-9e4b-56d854c83d9d · outbound

This paper cites Feature pyramid networks for object detection.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Feature pyramid networks for object detection

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.377500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:d0700a759e66ea8cd33ce56f11e969d44943b67c9c58b376c67498b0328df53c

Observation ecb22407-e09d-44e5-8a50-0ef0665779ea · outbound

This paper cites Infmae: A foundation model in the infrared modality.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Infmae: A foundation model in the infrared modality

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.337725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:154fcaf5d74bcf469b174de27377db3307957d5c54d4bf76c436883404e0479a

Observation 9c8f12f3-e134-4288-9aab-f4e382cf4c66 · outbound

This paper cites Target-aware dual adversarial learning and a multi-scenario multi-modality benchmark to fuse infrared and visible for object detection.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Target-aware dual adversarial learning and a multi-scenario multi-modality benchmark to fuse infrared and visible for object detection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.345230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:7a4f3131e521f99af9af8c91f5904b37e30fc0c84f38e3e63716de7584d3bc35

Observation 40059d25-3d69-4b6c-8f97-7356b393403e · outbound

This paper cites X-clip: End-to-end multi-grained con- trastive learning for video-text retrieval.

UNIV: Unified Foundation Model for Infrared and Visible Modalities X-clip: End-to-end multi-grained con- trastive learning for video-text retrieval

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.330038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:a3a15fb1c46392dcc99488340111535fc7220a487611494ada627c3830cf3816

Observation c14689af-6625-4c21-9dd5-2a781b3a2e02 · outbound

This paper cites Multi- level cross-modal semantic alignment network for video–text retrieval.Mathematics, 10(18):3346.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Multi- level cross-modal semantic alignment network for video–text retrieval.Mathematics, 10(18):3346

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.333648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:e1587bf62fb3259200a0f3456ea88bd3bdfb4f3f9542d20abb15479218b878e9

Observation 9bd0514a-bca2-42ce-b79d-67fde6eebb96 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

UNIV: Unified Foundation Model for Infrared and Visible Modalities DINOv2: Learning Robust Visual Features without Supervision

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:31:37.152875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:b5be7620b7d8c9d354c57a93cd5a9c3bd7caa67e7be4a6ebc44baa5f6b7f2656

Observation 82c40e70-4f60-47a8-a1df-8dbb573fb9df · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Learn- ing transferable visual models from natural language super- vision

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:38.588780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:3a5e7396c9bf51dac1b3568d14cf76c5ee799dcf282382e2ee6017e1a10af24c

Observation c225f6c9-7854-4ae2-9c34-395fdcff2f90 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.PAMI, 39(6):1137–1149.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Faster r-cnn: Towards real-time object detection with region proposal networks.PAMI, 39(6):1137–1149

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.341435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:15ad266396c534439832b1376cb26c7debe7426b66312ca21b1c63345dde9b85

Observation 212aae39-2afd-4154-81aa-c65fd10cd165 · outbound

This paper cites The earth mover’s distance as a metric for image retrieval.In- ternational journal of computer vision, 40(2):99–121.

UNIV: Unified Foundation Model for Infrared and Visible Modalities The earth mover’s distance as a metric for image retrieval.In- ternational journal of computer vision, 40(2):99–121

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:38.592078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:58fc9f91926556e900f3b4424f797d50b1b806869b1c53544f3deb25a97fedf5

Observation 2bcf8726-9a65-4f41-b42a-7d74d21c2c20 · outbound

This paper cites Drone-based rgb-infrared cross-modality vehicle detection via uncertainty-aware learning.TCSVT, 32(10):6700–6713.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Drone-based rgb-infrared cross-modality vehicle detection via uncertainty-aware learning.TCSVT, 32(10):6700–6713

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:38.585623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:20fdcdd6379390a5dfdad6e99b50fdf417d13bd3eb1f24a2674f1836803b0ab5

Observation 8f350059-768c-4a97-beb8-26d99a61f0e0 · outbound

This paper cites Piafusion: A progressive infrared and visible im- age fusion network based on illumination aware.Information Fusion.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Piafusion: A progressive infrared and visible im- age fusion network based on illumination aware.Information Fusion

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:38.595180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:43a4238d1649d9cd088a79813693d53965370c8f4b8287725501166b46edb0ff

Observation 60b6caac-a745-4695-b6ea-4b0dbbd09b16 · outbound

This paper cites Fine-grained action retrieval through multiple parts- of-speech embeddings.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Fine-grained action retrieval through multiple parts- of-speech embeddings

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.326472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:4f7cde1d66eebe0fa2954d0c21cd3d2d30568e950c4773d203fb2f12b4645833

Observation 40e74b97-18ab-4f2e-b782-b815d7afca10 · outbound

This paper cites Unified perceptual parsing for scene understand- ing.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Unified perceptual parsing for scene understand- ing

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.395213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:5d66de48bd6dd03df29c130f164fb042b2d4ca2002fb42061a68d9872f238742

Observation 0a8b5769-fd20-45b2-a733-9da959ddd5e7 · outbound

This paper cites Hitea: Hierarchical temporal- aware video-language pre-training.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Hitea: Hierarchical temporal- aware video-language pre-training

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.360978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:aa3fd8d9a40da482a94964d8c5dad40a02002be3c63b46fca4d413e04eca332a

Observation 36226bf4-198f-49f8-b791-891a9717c59d · outbound

This paper cites Sigmoid loss for language image pre-training.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Sigmoid loss for language image pre-training

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.348380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:b1509741c513ea1ec1f009bc812f49639785f921a3a974e4045c3768e04bbc55

Observation 2f978e6a-9ae7-466f-9d25-219877188654 · outbound

This paper cites PAD: Self-Supervised Pre-Training with Patchwise-Scale Adapter for Infrared Images.

UNIV: Unified Foundation Model for Infrared and Visible Modalities PAD: Self-Supervised Pre-Training with Patchwise-Scale Adapter for Infrared Images

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:31:37.157492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:80bfdc3f4ed68a451cc0ad6d87f2469d987d34f688cc4c6bc5f57775e8d06123

Observation deedfaa1-d2b8-4a73-a7ab-be4fab009306 · outbound

This paper cites UNIP: Rethinking Pre-trained Attention Patterns for Infrared Semantic Segmentation.

UNIV: Unified Foundation Model for Infrared and Visible Modalities UNIP: Rethinking Pre-trained Attention Patterns for Infrared Semantic Segmentation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:31:37.166147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:7faf9eb20722065c65581a3ee1943df6fd6abf024c5a909642994db7f3fb69c2

Observation 5c14ab0c-542b-4b1d-af81-d877ec393311 · outbound

This paper cites Scene parsing through ade20k dataset.

UNIV: Unified Foundation Model for Infrared and Visible Modalities Scene parsing through ade20k dataset

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:32:45.381215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:952ae8a6c943b2e855c63b7cffe829ac452397f7107913dda83b40e8c494aef2

Observation fb6eea95-5fe6-4981-b522-70b61f53c47b · outbound

This paper cites iBOT: Image BERT Pre-Training with Online Tokenizer.

UNIV: Unified Foundation Model for Infrared and Visible Modalities iBOT: Image BERT Pre-Training with Online Tokenizer

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:31:37.175410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:28:48.494939Z digest=sha256:20713a745ccceb98401902c9ac75df610a9ded89ea590ae059e143be28686a82

Pith citing papers

Observation 5d2a49d0-3772-4110-b706-4cb64ef106b9 · inbound

Not All Patches are Equal: Sampling Matters for Visible-Infrared Pre-Training cites this paper.

Not All Patches are Equal: Sampling Matters for Visible-Infrared Pre-Training UNIV: Unified Foundation Model for Infrared and Visible Modalities

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T10:27:20.578105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:27:20.578105Z digest=sha256:018b600153eaac73ac91734aab5218a8dff5710e273dc9ee1eba0f0198aacaed