Pith. sign in

Paper Citation Record · LEDGER

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation

As of 13 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2412.02262.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02262 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:41:41.515573Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:08:58.050673Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T21:08:58.212497Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact8
  • verified fuzzy9
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29b320c0-d0ae-48dd-b7d2-af8f30e2480c · outbound

This paper cites Wild salmon enumeration and monitoring using deep learning empow- ered detection and tracking.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Wild salmon enumeration and monitoring using deep learning empow- ered detection and tracking

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.307801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.307801Z digest=sha256:c931081bfa4733f818d65a0e8b0ddd0462f872a1521793db4d8543444c426cfd

Observation 5e6bbc10-6c55-4375-a10d-d356830a6be0 · outbound

This paper cites Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.316085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.316085Z digest=sha256:10927776d3180c99a9e9ef5e2055b4b1f0d096ada3aaf815c4c085909616edea

Observation 8bb92a7c-af3d-405c-ad44-385ac67d098b · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recogni- tion at Scale.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation An Image is Worth 16x16 Words: Transformers for Image Recogni- tion at Scale

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.791378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.329852Z digest=sha256:518d3ba78137baef32a8a2c170eea86f5686e11448f6cd49fda331d4ce4f91d6

Observation d48afad5-34ad-4bda-99b6-4f6c30e29f1a · outbound

This paper cites Knowledge Augmented Instruction Tuning for Zero-shot Animal Species Recognition.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Knowledge Augmented Instruction Tuning for Zero-shot Animal Species Recognition

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.767809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.336052Z digest=sha256:9e631494cab598c89508c0ae4e982594867d989ca87a677421bd3c91a5c1ebe3

Observation 67d7ca62-2d10-4680-9ecc-073068561787 · outbound

This paper cites Joint SDG Fund | Goal 14: Life below water.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Joint SDG Fund | Goal 14: Life below water

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.747194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.341976Z digest=sha256:dec4056b88b249909d73c86586cab75acb39b7024f339e2acea5de16227439c2

Observation aeee8331-6626-4a1d-8d8f-8795f07f70cd · outbound

This paper cites REALM: Retrieval-Augmented Language Model Pre-Training.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation REALM: Retrieval-Augmented Language Model Pre-Training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.348366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.348366Z digest=sha256:b9f2336934ca60b3a2bc87da92f1622abfa99b8eaa053159369b17f3209e084b

Observation d30c9162-2cd7-4f31-922f-374e61a06702 · outbound

This paper cites Deep Residual Learning for Image Recognition.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Deep Residual Learning for Image Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.354741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.354741Z digest=sha256:3f48bc7087934000806f652939c57893e683d86b3a7b7ca334e1a4694131a289

Observation f0e327bb-4f50-41b1-966a-f9be952974be · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation LoRA: Low-Rank Adaptation of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.362035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.362035Z digest=sha256:ddabc21bbfe995c1f0470bc4aaa4097cea69463f8e60ef76e6640003141d2063

Observation f84ed92c-1b3a-4a48-bbb4-cb680dde8698 · outbound

This paper cites REVEAL: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge Memory.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation REVEAL: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge Memory

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.370174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.370174Z digest=sha256:d73783329a7a9aa269027d64f727b93f33a8920d1480f9fcf456f6110750947e

Observation ebef789a-774d-4c7b-9bdb-8ac6685034a8 · outbound

This paper cites Active Retrieval Augmented Generation.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Active Retrieval Augmented Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.377215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.377215Z digest=sha256:b0fa2f40fb6d9da7bb932e2295e6f0c3224285c72e6843addb19d805c08d3f5f

Observation 02567d19-32f8-421f-8c00-144407b67629 · outbound

This paper cites FathomNet: A global image database for enabling artificial intelligence in the ocean.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation FathomNet: A global image database for enabling artificial intelligence in the ocean

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.727632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.383031Z digest=sha256:4d4824bec828f300196a2369173deaacfd9b2055e6c0bdc1cd5e015a1a71a35c

Observation fc83bae2-e79c-479d-b7e2-898892a1590c · outbound

This paper cites The Fishnet Open Images Database: A Dataset for Fish Detection and Fine-Grained Categorization in Fisheries.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation The Fishnet Open Images Database: A Dataset for Fish Detection and Fine-Grained Categorization in Fisheries

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.396405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.396405Z digest=sha256:18fe4b4e1fe7dcf1761a91f5db7e88db9c7ece02ff5f75b58b8a58c815acdcd4

Observation 539593de-eb52-497f-b5b7-a67952e93b5b · outbound

This paper cites ImageNet Classification with Deep Convolutional Neural Networks.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation ImageNet Classification with Deep Convolutional Neural Networks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.708274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.404767Z digest=sha256:48c18ef530cd76bae91ebe96dcebb0da93ef3464e71104b0b36dd29461756e4b

Observation a72b9a4c-ce03-451d-b5d1-16ecc87e5e16 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.411781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.411781Z digest=sha256:74c2eb031e0da3c480060bfb0614677f63adf9ed61b950de29906ddf6b584669

Observation ac06c90b-220c-4922-bf3e-5aef7eb10c32 · outbound

This paper cites Grounded Language-Image Pre-training.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Grounded Language-Image Pre-training

Reference 15

Resolution
verified exact
raw_fallback, observed 2026-08-11T23:41:42.361860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.417891Z digest=sha256:2483bee6fc5398a7390639af275bc14868161e61b3b73b400a8b773494f63ac8

Observation 0076c009-5bf5-4e80-a9b4-351baf72c9a3 · outbound

This paper cites Visual Instruction Tuning.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Visual Instruction Tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.424916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.424916Z digest=sha256:f6d0f7560bdc6bbc4ce0b1017e73e2b9e1f3a3859f30cb65e7fcf3aec125a720

Observation af48f09c-8a81-4ec3-a30d-b8f6c070f987 · outbound

This paper cites KRISP: Integrating Implicit and Symbolic Knowledge for Open- Domain Knowledge-Based VQA.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation KRISP: Integrating Implicit and Symbolic Knowledge for Open- Domain Knowledge-Based VQA

Reference 17

Resolution
verified exact
raw_fallback, observed 2026-08-11T23:41:42.257825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.431235Z digest=sha256:dfaf185ecc2ae346b9cb5b96d4ba6757b2d638e5351cbc5d316771df514884fd

Observation b7cc8d75-a0ac-4c2b-a3a5-21f9578f77a9 · outbound

This paper cites OK-VQA: A Visual Question Answering Benchmark Requiring Exter- nal Knowledge.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation OK-VQA: A Visual Question Answering Benchmark Requiring Exter- nal Knowledge

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-11T23:41:42.171880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.437294Z digest=sha256:04a0462972e5d201121fe7713f298c52784ff95a2d321a2163e0284045226397

Observation 3c2bd38b-8bee-4502-87b0-2599141d6b49 · outbound

This paper cites New frontiers in AI for biodiversity research and conservation with multimodal language models.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation New frontiers in AI for biodiversity research and conservation with multimodal language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.689102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.442750Z digest=sha256:f766580916e9e3fa602a24251deba8e62e9ee87315749edb84a0cb775dad0416

Observation 5358996d-cb21-480c-a9a1-87b151644736 · outbound

This paper cites A deep active learning system for species identification and counting in camera trap images.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation A deep active learning system for species identification and counting in camera trap images

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:41:42.070193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.449340Z digest=sha256:c3fa824692b616eed23420779f5bc9110808aa1c51c662fb157de3e12256f9d0

Observation ffb93696-6913-4a47-8ed8-a1dae51f9936 · outbound

This paper cites Automatically identifying, counting, and describing wild animals in camera-trap images with deep learning.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Automatically identifying, counting, and describing wild animals in camera-trap images with deep learning

Reference 21

Resolution
malformed identifier
no resolver link, observed 2026-08-11T23:41:41.457182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.457182Z digest=sha256:537f8c7b20c4f233f34be73f553338dd2a2dca99e7ed4ec6f2cd324ee2b9c8ad

Observation 7a972156-5d2a-478c-9caf-1dd21b047f55 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervi- sion.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Learning Transferable Visual Models From Natural Language Supervi- sion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.668802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.462361Z digest=sha256:7a71ff4623c06697e5724ebca5bf2f5f0ae49f2ad857339bc280b4a0a11d8215

Observation 5ee30694-4b65-4bcf-b469-e2cc1dbe3124 · outbound

This paper cites SmallCap: Lightweight Image Captioning Prompted with Retrieval Augmentation.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation SmallCap: Lightweight Image Captioning Prompted with Retrieval Augmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.467754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.467754Z digest=sha256:71291f4c7051252dd791c4ba5e998c03b5f6faebb15d65740613f80be37d212c

Observation 4e4456c4-02a9-4c53-8876-89ce403ecf11 · outbound

This paper cites You Only Look Once: Unified, Real-Time Object Detection.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation You Only Look Once: Unified, Real-Time Object Detection

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.475227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.475227Z digest=sha256:1f1c7eabcf80264a8a14cea640f99cba7a97e60989ec57e578aac4bf7c49e149

Observation 444c47a0-07fa-433b-877c-f53d199ab57c · outbound

This paper cites Retrieval-Augmented Transformer for Image Captioning.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Retrieval-Augmented Transformer for Image Captioning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.481605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.481605Z digest=sha256:a1c354b364dcf6e25eb62c63e0daa0deb732866e09b8204104c4c5510951f1e0

Observation 67858487-0c8c-4611-af1d-7a4d38a671b3 · outbound

This paper cites K-LITE: Learning Transferable Visual Models with External Knowl- edge.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation K-LITE: Learning Transferable Visual Models with External Knowl- edge

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.647250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.487228Z digest=sha256:217e6dd25c80bc6820e6537ec5556c6904c2339f6d139c5a43100af96cab32ac

Observation f2b0088a-0b81-4756-829a-cacbfd8a1cb6 · outbound

This paper cites BioCLIP: A Vision Foundation Model for the Tree of Life.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation BioCLIP: A Vision Foundation Model for the Tree of Life

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.629866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.492660Z digest=sha256:1fc7ac5d4bd794cc55bddac8a910c55c5f6350a126ed1198eb96c2b4df5fa9f7

Observation b819b07f-d332-4c90-8e8d-9dbbf9d32e74 · outbound

This paper cites The iNaturalist Species Classification and Detection Dataset.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation The iNaturalist Species Classification and Detection Dataset

Reference 28

Resolution
verified exact
raw_fallback, observed 2026-08-11T23:41:41.895892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.497528Z digest=sha256:d8314200fbaaae53bffeb1d1dc6e814efee6280963882c007c6bbfbc09b3d0ee

Observation 9185048f-4e26-452c-b128-e2fcc4df740f · outbound

This paper cites Advancing artificial intelligence in fisheries requires novel cross-sector collaborations.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Advancing artificial intelligence in fisheries requires novel cross-sector collaborations

Reference 29

Resolution
verified exact
doi, observed 2026-08-11T23:41:41.575579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.505489Z digest=sha256:4a06d55e057259344c83c7acdec2cf59b6a8c372cc9b4662d310a630f467056b

Observation 430791f0-3760-4813-a2f7-4a52502f63b4 · outbound

This paper cites Multi-Modal Answer Validation for Knowledge-Based VQA.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Multi-Modal Answer Validation for Knowledge-Based VQA

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:41:41.800943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.510815Z digest=sha256:89d721f143aa350c3a46103dba396e3fc51bb4e6ed387fac606abdd0139f2eda

Observation a537c598-4f1f-48b1-9fd1-d6447bc5d05f · outbound

This paper cites MSR-VTT: A Large Video Description Dataset for Bridging Video and Lan- guage.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation MSR-VTT: A Large Video Description Dataset for Bridging Video and Lan- guage

Reference 31

Resolution
verified exact
raw_fallback, observed 2026-08-11T23:41:41.774782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:41:41.515573Z digest=sha256:35177b487b590d06d2cb1c981f52137c1fe84ddf64de41c2cfd284437755ed50

Pith citing papers

Observation c067fa4c-c462-4a07-b63b-423631a37309 · inbound

Environmental large language model Evaluation (ELLE) dataset: A Benchmark for Evaluating Generative AI applications in Eco-environment Domain cites this paper.

Environmental large language model Evaluation (ELLE) dataset: A Benchmark for Evaluating Generative AI applications in Eco-environment Domain Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T21:08:58.219723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:08:58.050673Z digest=sha256:3e1c901ad2ba7cc03452f7a0e11be21b6e5863e0a4224d4448a45573226a337f