Pith. sign in

Paper Citation Record · LEDGER

EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2505.04623.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04623 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:13.645079Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T06:36:10.581673Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bb8cd184-7364-4401-a920-23f1239dae01 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.645079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.645079Z digest=sha256:1d0f336cc73ac1cdd0c44d72308154a123f6d5351df8df7d66e1c9f286430fbb

Observation 8809531c-3f7a-4032-8af3-4ab46ef679fb · inbound

FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design cites this paper.

FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:40.398754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:40.398754Z digest=sha256:e608d71767658eda9352173b3d77d76b7b043e21960779fa961cde38f362ba9d

Observation 2caf0f52-2c2f-4cb8-bcf5-c365bbbc7f9b · inbound

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context cites this paper.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:15.038514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:15.038514Z digest=sha256:6073a90b5afccd4a3da30d4f2d5e0a811b8df6dd1af2fa40c003f722e389ad3e

Observation 2f29198a-1efa-4caf-8739-a82b313520b7 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

Reference 264

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.404714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:f8992429b648b139992237d49af015362b5fc17a9f08e89140768ac41fcad6fe

Observation e0a7d057-b8c4-4e62-9dde-6caf0ffe500f · inbound

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models cites this paper.

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:45:56.120572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:45:07.700571Z digest=sha256:09f80c2bb17a5fc305bf03875e5ddcc5c638bf4065ac9f4c260d24711c82ff31

Observation 6cbf7b30-a75b-4436-97f2-cf6cb8b1392d · inbound

Development of a 3D-CNN-based Prediction Model for Migration Barriers in Plasma-Wall Interactions cites this paper.

Development of a 3D-CNN-based Prediction Model for Migration Barriers in Plasma-Wall Interactions EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T09:25:05.174833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:25:05.174833Z digest=sha256:83201c9c0ec74268d1de10e0d4dae576f5479e35458d4399a6bffdcc6736fa24

Observation 0a9dfaff-cf0c-4434-995f-49833a1bc469 · inbound

Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs cites this paper.

Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:25:49.792701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:10:52.629490Z digest=sha256:143ee8bb0f64dd0dbf5c5ece40e3acaadab276f3a5e548a5bf058c466bd250db

Observation cb2b7095-f2cf-4861-8605-52d8989c1bad · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:03.503815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:d8a93ee6848bf2a8b5d4e28230a8a6136b28cbd04ec4990482b9c2f6219fbdd0

Observation 87eebf59-fac5-45cc-b15c-4399c9236ab6 · inbound

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale cites this paper.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:03.279034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:8072c09f6fadfcec7860001680cc4e085cf733ee72f44a03255463e7b9f2b683

Observation 1afb3801-19f9-4cf1-a99e-336a9f694b0a · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.117814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:005c7b1a005dd1faa4d7337b74eb5215b00b8f782cb8021d3706a7b1a06c49e1

Observation 43e2786f-801b-4e1b-aab3-a201e66cfbbc · inbound

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers cites this paper.

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:18:32.200681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:02:53.574120Z digest=sha256:d2db20e662cc5a725492b7212db95a5074f0d6daade7d7afb2dc6a6206d84722

Observation dcd6a649-c1a0-4046-af9f-6f01963efd5f · inbound

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning cites this paper.

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:36:10.587746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T06:34:57.483234Z digest=sha256:90113e7e7907d7ee7955305555ab25adb14853ea920554ecf8d66d4646c8f49a