Pith. sign in

Paper Citation Record · LEDGER

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation

As of 16 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2412.02262.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02262 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:41:41.515573Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:08:58.050673Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T21:08:58.212497Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact8
  • verified fuzzy9
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29b320c0-d0ae-48dd-b7d2-af8f30e2480c · outbound

This paper cites Wild salmon enumeration and monitoring using deep learning empow- ered detection and tracking.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Wild salmon enumeration and monitoring using deep learning empow- ered detection and tracking

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.307801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.307801Z digest=sha256:199f4a92460bc79769da44e3764014e03fe477df9c8684783557edb53cb9cae9

Observation 5e6bbc10-6c55-4375-a10d-d356830a6be0 · outbound

This paper cites Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.316085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.316085Z digest=sha256:a833074e0110567b94075dfad155243a49b66d0b9bb4c9f4958c28e4c81bfb1b

Observation 8bb92a7c-af3d-405c-ad44-385ac67d098b · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recogni- tion at Scale.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation An Image is Worth 16x16 Words: Transformers for Image Recogni- tion at Scale

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.791378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.329852Z digest=sha256:666021483886a1e259066563f07d09d5b8cf7a248386954f04683e32d2c4f234

Observation d48afad5-34ad-4bda-99b6-4f6c30e29f1a · outbound

This paper cites Knowledge Augmented Instruction Tuning for Zero-shot Animal Species Recognition.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Knowledge Augmented Instruction Tuning for Zero-shot Animal Species Recognition

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.767809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.336052Z digest=sha256:c91694897aea692de9940bb0af48b5ecd544dae7882a55a9786e69136a5a5b53

Observation 67d7ca62-2d10-4680-9ecc-073068561787 · outbound

This paper cites Joint SDG Fund | Goal 14: Life below water.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Joint SDG Fund | Goal 14: Life below water

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.747194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.341976Z digest=sha256:bdf7679e7c4fde7a178009cbfc88dc655c4b0807e077eb8fdcc5ca563d1ac46e

Observation aeee8331-6626-4a1d-8d8f-8795f07f70cd · outbound

This paper cites REALM: Retrieval-Augmented Language Model Pre-Training.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation REALM: Retrieval-Augmented Language Model Pre-Training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.348366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.348366Z digest=sha256:99abe23b5ef7f94a71961defb5d14543d76fbf65e6f70e3bbeb490b3709c15e5

Observation d30c9162-2cd7-4f31-922f-374e61a06702 · outbound

This paper cites Deep Residual Learning for Image Recognition.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Deep Residual Learning for Image Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.354741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.354741Z digest=sha256:db5363cef852578596cd0a12ef2ac85d20592f56ce00e0795cf27e2a62eb32f2

Observation f0e327bb-4f50-41b1-966a-f9be952974be · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation LoRA: Low-Rank Adaptation of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.362035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.362035Z digest=sha256:a388b0676b0aa13adb54329e7b546d9dfff03f4c7a0b3c92b63c8fa9e18ef901

Observation f84ed92c-1b3a-4a48-bbb4-cb680dde8698 · outbound

This paper cites REVEAL: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge Memory.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation REVEAL: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge Memory

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.370174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.370174Z digest=sha256:5dbd6b1a19c00930518b704bc087664ba6c35c50ca5e7790db2c7db8f6d74c0f

Observation ebef789a-774d-4c7b-9bdb-8ac6685034a8 · outbound

This paper cites Active Retrieval Augmented Generation.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Active Retrieval Augmented Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.377215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.377215Z digest=sha256:e6240cb3ba203949a5e1ce74d0f94c3a62379551641bcaab2ddaf478aafc0f4c

Observation 02567d19-32f8-421f-8c00-144407b67629 · outbound

This paper cites FathomNet: A global image database for enabling artificial intelligence in the ocean.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation FathomNet: A global image database for enabling artificial intelligence in the ocean

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.727632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.383031Z digest=sha256:984dcd54cb95b673a62654c1a41d1a0147c0e9907da0b3c4ab56fc916ebf7fba

Observation fc83bae2-e79c-479d-b7e2-898892a1590c · outbound

This paper cites The Fishnet Open Images Database: A Dataset for Fish Detection and Fine-Grained Categorization in Fisheries.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation The Fishnet Open Images Database: A Dataset for Fish Detection and Fine-Grained Categorization in Fisheries

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.396405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.396405Z digest=sha256:2d4d1634fa830046fbedcf87aeb146234e06d8cb71c9069f18fdf11c102a9c71

Observation 539593de-eb52-497f-b5b7-a67952e93b5b · outbound

This paper cites ImageNet Classification with Deep Convolutional Neural Networks.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation ImageNet Classification with Deep Convolutional Neural Networks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.708274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.404767Z digest=sha256:beeb6c8cf763e12978983a6b2cf369cb4df2732251b8e6e3564cfa5a686d5376

Observation a72b9a4c-ce03-451d-b5d1-16ecc87e5e16 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.411781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.411781Z digest=sha256:144c0a1b7078e9e85c7c19c8996420f7819d87e5b5cc006d253f5424ea3cdf18

Observation ac06c90b-220c-4922-bf3e-5aef7eb10c32 · outbound

This paper cites Grounded Language-Image Pre-training.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Grounded Language-Image Pre-training

Reference 15

Resolution
verified exact
raw_fallback, observed 2026-08-11T23:41:42.361860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.417891Z digest=sha256:0d4495460f6b062636cd29e894ac98b1c0816b46d1971b251550f74b0b00e39d

Observation 0076c009-5bf5-4e80-a9b4-351baf72c9a3 · outbound

This paper cites Visual Instruction Tuning.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Visual Instruction Tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.424916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.424916Z digest=sha256:37a179b70711d53eb23ef39cde65de417a10b5d14af00fc9c6e9a322f942fca4

Observation af48f09c-8a81-4ec3-a30d-b8f6c070f987 · outbound

This paper cites KRISP: Integrating Implicit and Symbolic Knowledge for Open- Domain Knowledge-Based VQA.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation KRISP: Integrating Implicit and Symbolic Knowledge for Open- Domain Knowledge-Based VQA

Reference 17

Resolution
verified exact
raw_fallback, observed 2026-08-11T23:41:42.257825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.431235Z digest=sha256:2770fbd9bb1fcdbcbdc656983794e4eb2e8a7f83d6e394448677f69191f2e975

Observation b7cc8d75-a0ac-4c2b-a3a5-21f9578f77a9 · outbound

This paper cites OK-VQA: A Visual Question Answering Benchmark Requiring Exter- nal Knowledge.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation OK-VQA: A Visual Question Answering Benchmark Requiring Exter- nal Knowledge

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-11T23:41:42.171880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.437294Z digest=sha256:ea6cb6d6954995f8a9b13149a3f2d141a242ad761a1a699cf0de1a1031653cc3

Observation 3c2bd38b-8bee-4502-87b0-2599141d6b49 · outbound

This paper cites New frontiers in AI for biodiversity research and conservation with multimodal language models.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation New frontiers in AI for biodiversity research and conservation with multimodal language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.689102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.442750Z digest=sha256:4f0d2cbce6a823cf6d56cb676e48640e43c16c99d128bc3daf257e15a4301676

Observation 5358996d-cb21-480c-a9a1-87b151644736 · outbound

This paper cites A deep active learning system for species identification and counting in camera trap images.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation A deep active learning system for species identification and counting in camera trap images

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:41:42.070193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.449340Z digest=sha256:d782606cd8c7d1df6e34860e1ec56e92b92c2226845e3c0e4285706c5fb10916

Observation ffb93696-6913-4a47-8ed8-a1dae51f9936 · outbound

This paper cites Automatically identifying, counting, and describing wild animals in camera-trap images with deep learning.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Automatically identifying, counting, and describing wild animals in camera-trap images with deep learning

Reference 21

Resolution
malformed identifier
no resolver link, observed 2026-08-11T23:41:41.457182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.457182Z digest=sha256:73315e2c553252449b1d927b50bb03fdeba3248ad0d694897ddf8b44dbdf0f38

Observation 7a972156-5d2a-478c-9caf-1dd21b047f55 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervi- sion.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Learning Transferable Visual Models From Natural Language Supervi- sion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.668802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.462361Z digest=sha256:577236b95447d2919c4c92162dfb522b55d9c23aa45294634679c176851d3ebc

Observation 5ee30694-4b65-4bcf-b469-e2cc1dbe3124 · outbound

This paper cites SmallCap: Lightweight Image Captioning Prompted with Retrieval Augmentation.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation SmallCap: Lightweight Image Captioning Prompted with Retrieval Augmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.467754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.467754Z digest=sha256:dad314c9da8ca6024c9b0b293adbc5167ccb19cb9e63fcc3e99c6494a4c6d1bb

Observation 4e4456c4-02a9-4c53-8876-89ce403ecf11 · outbound

This paper cites You Only Look Once: Unified, Real-Time Object Detection.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation You Only Look Once: Unified, Real-Time Object Detection

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.475227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.475227Z digest=sha256:0be9039aaaa6a34fc6a1dad4ae286b1f4944ad67936d3a609e9dc9077c66f3dc

Observation 444c47a0-07fa-433b-877c-f53d199ab57c · outbound

This paper cites Retrieval-Augmented Transformer for Image Captioning.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Retrieval-Augmented Transformer for Image Captioning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:41.481605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:41.481605Z digest=sha256:cfb72f5a0a0cdb1f935abe61c453f927036c29afa9ede43406b787327c275c33

Observation 67858487-0c8c-4611-af1d-7a4d38a671b3 · outbound

This paper cites K-LITE: Learning Transferable Visual Models with External Knowl- edge.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation K-LITE: Learning Transferable Visual Models with External Knowl- edge

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.647250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.487228Z digest=sha256:0af6b646cc984ed8d1ed3261f257fbfac563189283ab3e76615705205d0cf312

Observation f2b0088a-0b81-4756-829a-cacbfd8a1cb6 · outbound

This paper cites BioCLIP: A Vision Foundation Model for the Tree of Life.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation BioCLIP: A Vision Foundation Model for the Tree of Life

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:41:42.629866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.492660Z digest=sha256:fbdbee629c1b3c13a1c3268d092c0dc595f49e1faedbb37a6d7496712a0fd673

Observation b819b07f-d332-4c90-8e8d-9dbbf9d32e74 · outbound

This paper cites The iNaturalist Species Classification and Detection Dataset.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation The iNaturalist Species Classification and Detection Dataset

Reference 28

Resolution
verified exact
raw_fallback, observed 2026-08-11T23:41:41.895892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.497528Z digest=sha256:936bf0333132acd831ab1096fcb75e98f0597d6b2dbf684b5d1c65885f83dbf8

Observation 9185048f-4e26-452c-b128-e2fcc4df740f · outbound

This paper cites Advancing artificial intelligence in fisheries requires novel cross-sector collaborations.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Advancing artificial intelligence in fisheries requires novel cross-sector collaborations

Reference 29

Resolution
verified exact
doi, observed 2026-08-11T23:41:41.575579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.505489Z digest=sha256:58e34669881d12aa1682ced7c09d18071aff1f0e28ff21d79f3ba82ed9371051

Observation 430791f0-3760-4813-a2f7-4a52502f63b4 · outbound

This paper cites Multi-Modal Answer Validation for Knowledge-Based VQA.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation Multi-Modal Answer Validation for Knowledge-Based VQA

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:41:41.800943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.510815Z digest=sha256:8c29dde8883ba21b984f521b9bb4b64e86f83936c802a60609ab099cd358862e

Observation a537c598-4f1f-48b1-9fd1-d6447bc5d05f · outbound

This paper cites MSR-VTT: A Large Video Description Dataset for Bridging Video and Lan- guage.

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation MSR-VTT: A Large Video Description Dataset for Bridging Video and Lan- guage

Reference 31

Resolution
verified exact
raw_fallback, observed 2026-08-11T23:41:41.774782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:41:41.515573Z digest=sha256:343e05e1593c6c9e9642d623db3993fc41859d9859d2b726ef400a85196ee5f2

Pith citing papers

Observation c067fa4c-c462-4a07-b63b-423631a37309 · inbound

Environmental large language model Evaluation (ELLE) dataset: A Benchmark for Evaluating Generative AI applications in Eco-environment Domain cites this paper.

Environmental large language model Evaluation (ELLE) dataset: A Benchmark for Evaluating Generative AI applications in Eco-environment Domain Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T21:08:58.219723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T21:08:58.050673Z digest=sha256:348b98b853417976f3d7f0a978985c04b2e561fa33b771f3a42e1349a784a767