Pith. sign in

Paper Citation Record · LEDGER

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

As of 5 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 10 inbound Pith citation observations for arXiv:2604.24954.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.24954 v2

Coverage vector

measured 6 of 6 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T02:57:23.422932Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T12:51:52.021657Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

6 of 6 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 25c3deaa-4eef-483a-86de-7683bf222c92 · outbound

This paper cites Efficient video sampling: Pruning temporally redundant tokens for faster vlm inference.

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence Efficient video sampling: Pruning temporally redundant tokens for faster vlm inference

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.543471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:57:23.422932Z digest=sha256:5f706c1467d63a5d7d45a90900b8f6521abbabb5cf93be63de5eefb282d6a678

Observation 97ec84fa-a59a-4613-a0b4-8758b8d3bad6 · outbound

This paper cites OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning.

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-27T02:04:28.324233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:57:23.422932Z digest=sha256:792cc0af2a4672d5b5cafbb7cdbcb8e5fc3da8c8e0166f2da8ce402fbf887cbc

Observation 4f1600e1-8ece-46d0-9f47-a47c45428527 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 3

Resolution
metadata mismatch
doi, observed 2026-05-12T03:01:17.442185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:57:23.422932Z digest=sha256:2ce05dbe18c413bfad87759d8eeaccd1053ac4c6b942048b95e75747de1112a1

Observation 6461e1b3-d41a-4c09-a669-70d61a080421 · outbound

This paper cites Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning.

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:01:18.539530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:57:23.422932Z digest=sha256:80ed6ab94d96d08ac7df68d0aaf9289e3795db632e296c25e9664d75af5bbe21

Observation 66e69f25-48c2-415b-a054-bcd20e0da5e6 · outbound

This paper cites Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual Speech Recognition Evaluation.

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual Speech Recognition Evaluation

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.535846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:57:23.422932Z digest=sha256:a58b70e4130492589ae95849da50f66888590b37e4a8b94151e1f71a4e26e157

Observation 1ea3270b-503c-4878-affb-3f90dfc0ba9a · outbound

This paper cites Group Sequence Policy Optimization.

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence Group Sequence Policy Optimization

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T03:01:18.531781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:57:23.422932Z digest=sha256:4683b499b8f3ba0e94fd571e25b39bee1338421ef1a21fcdec5f6e6d134454fe

Pith citing papers

Observation 4bc8a74b-37ad-4fb0-b97d-7976ffe6066f · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-13T03:52:12.713065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T03:49:58.240883Z digest=sha256:856eef9b2b6441fb129732b35af248dd4513d242180c8bc7549bff4b763bc4b0

Observation 3a09842d-106c-43f7-a5de-9176324abb6f · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:09:50.349599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T06:06:20.658030Z digest=sha256:0fd236a75dcfea2c897836fbc46611ddc92fc7e4156e010993dd2158b9f85910

Observation 3d282a39-dedf-4f93-aee4-09a64aa113f2 · inbound

OmniDrop: Layer-wise Token Pruning for Omni-modal LLMs via Query-Guidance cites this paper.

OmniDrop: Layer-wise Token Pruning for Omni-modal LLMs via Query-Guidance Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:38:27.616614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:36:11.208299Z digest=sha256:1211161ebb1de10d68775cc4ea734f0d336c390294a81675eadf41a00e149a6a

Observation b4071e9e-7a53-412e-8b2d-d848d5921b11 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:01.949375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:1ee63a427a5caa9ed4bb1e8266b17f6a45f24af3c054a0366968007cfeed111e

Observation 870dfe79-622a-48e3-920c-530cf5a75ea9 · inbound

VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents cites this paper.

VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:33:15.302774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T08:28:50.818250Z digest=sha256:56d8db8115dc4eaa78f2dad1426bc2b740819de96b86ca71a437e66d2f10d250

Observation 50906f4e-c62f-4c97-8f1f-6835e06ec6e8 · inbound

VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents cites this paper.

VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T12:51:52.021657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:51:52.021657Z digest=sha256:8ee9c727a1b83c0a5233492cfa9df51aa168a8b1e67a1366eee784c3e4fbf837

Observation 87f5a9af-7634-4947-992d-b46a6c763b69 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

Reference 251

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.669220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:a19a68dac0ff22831b3a0f17b008b3ec9910892bbde07db978e5c63838f30ceb

Observation 6f7192cd-b669-4f83-a14c-a7cf9d12ba4f · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

Reference 251

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:1b7727ad23220896a63f78928084e078b4de157444aa045df82368f9d62ec882

Observation 18256e30-b327-48c8-bffc-5c7961ef86db · inbound

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages cites this paper.

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:15:30.183427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T19:08:06.748619Z digest=sha256:a7a9a6e02f4af659ced18bb8558c69542452a872e34041e530000bed95c14b87

Observation 69a30c48-55a7-4e38-b11e-07600fb82e9a · inbound

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos cites this paper.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:43.362060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:43.362060Z digest=sha256:d5639e3471f293f35dc5bb0a60015651482eb9055f7ecce0e22042b1b696e173