Pith. sign in

Paper Citation Record · LEDGER

Representations in vision and language converge in a shared, multidimensional space of perceived similarities

As of 8 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2507.21871.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21871 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:21:22.544303Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:32:28.422930Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T16:32:28.671394Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact2
  • verified fuzzy10
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9bdf7fb3-8a6f-4f18-8ba3-ce6f87e63edf · outbound

This paper cites Emerging evidence suggests that human brain representations in both vision and language are well predicted by semantic feature spaces obtained from large language models (LLMs).

Representations in vision and language converge in a shared, multidimensional space of perceived similarities Emerging evidence suggests that human brain representations in both vision and language are well predicted by semantic feature spaces obtained from large language models (LLMs)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.778065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:22.476705Z digest=sha256:64d48693ebd2d8c8ca88bad54c50a58b758bbf4a47daae8427549a3a60377eee

Observation 808fd061-f75c-4eae-8e7f-de22674eed3a · outbound

This paper cites (A) Cross-validated non-negative least squares regression was used to model the brain RDMs at every searchlight location using the behavioural RDMs derived from our MA tasks.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities (A) Cross-validated non-negative least squares regression was used to model the brain RDMs at every searchlight location using the behavioural RDMs derived from our MA tasks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.727176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:22.496414Z digest=sha256:2ade81a05440d3069abf0fe36a9ba6f7f2457f4dd45d7408c2d01c7e62756064

Observation 5bbd874c-07bd-406d-97ab-218d9a5d3fd2 · outbound

This paper cites an unresolved cited work.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:21:22.716796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:22.500243Z digest=sha256:3bf5a8411710f62d366826e00bd14f364ffb48e16ec247a76df03432b3b6e196

Observation 09891198-0646-410c-acc0-115c75458ddb · outbound

This paper cites linguistic modality.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities linguistic modality

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.748159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:22.488582Z digest=sha256:1eafd0f721d81dc36d4eba770d767e6448fc6c531bac322dc835e695b09f2f10

Observation 2a6fb438-672d-4b6a-b27c-01a0d165e43e · outbound

This paper cites (A) Participants completed the MA task either on 100 natural scene images (visual modality left) or 100 sentence captions 8 describing the images (linguistic modality right).

Representations in vision and language converge in a shared, multidimensional space of perceived similarities (A) Participants completed the MA task either on 100 natural scene images (visual modality left) or 100 sentence captions 8 describing the images (linguistic modality right)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.738082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:22.492359Z digest=sha256:7b44570e9e298b352be674458a13d82295a1f7e1fabd40c7f7f5f61d64b378e5

Observation 6cb16c37-3857-4953-adea-9511c086b9a4 · outbound

This paper cites Scaling Laws for Transfer.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities Scaling Laws for Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:22.532478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:22.532478Z digest=sha256:7c9639b848be43808cb43b274f25d5fc4757f3893f8f763eea105c8b157bf7e2

Observation 698df9c9-edc8-4cb8-b09b-0509435d986b · outbound

This paper cites VoLTA: Vision-Language Transformer with Weakly-Supervised Local-Feature Alignment.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities VoLTA: Vision-Language Transformer with Weakly-Supervised Local-Feature Alignment

Reference 128

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:21:22.581138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:22.544303Z digest=sha256:cf580803d825903912f65354b741f8a1e8b74bdb9f19dae2fca55e69887b77d2

Observation 678bc90a-5ea0-4743-acde-4b632dd97752 · outbound

This paper cites A., Schmitz, T.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities A., Schmitz, T

Reference 134

Resolution
verified exact
doi, observed 2026-08-06T12:21:22.636069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:22.525364Z digest=sha256:4a3a92aba53dda942b8d685a9fa075f464e6fa062850d794d11d2057dddf7427

Observation 24dcb9c7-e65c-45d3-be8e-6ace6e4c0c1b · outbound

This paper cites A., Kiani, R., Bodurka, J., Esteky, H., Tanaka, K., & Bandettini, P.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities A., Kiani, R., Bodurka, J., Esteky, H., Tanaka, K., & Bandettini, P

Reference 245

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.624190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:22.537199Z digest=sha256:388dccff8db6372793bb0713e5bd059ef6284825434099e9fb78bc71dcd2b23d

Observation 0f02e926-5b21-49b1-aaf4-c65258533fd1 · outbound

This paper cites visual modality.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities visual modality

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.758598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:22.485065Z digest=sha256:5959ffc6f1ec78643eb5af97b58ac821fa16a508613b92eb09e6952d23ada731

Observation c8bfdb37-929c-4e63-94b7-f280d91f9d21 · outbound

This paper cites We show that a similar relational structure emerges for both linguistic and visual inputs.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities We show that a similar relational structure emerges for both linguistic and visual inputs

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.706659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:22.503950Z digest=sha256:9762900a21da2be9e5324b3ca67b6eabff586677de9e427ac70d612539f6af14

Observation 5545cf45-c302-4c49-ac71-85dcceb97a61 · outbound

This paper cites The significance of correlations was tested using one-sided t-test across participants and corrected for multiple comparisons at FDR p < 0.05.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities The significance of correlations was tested using one-sided t-test across participants and corrected for multiple comparisons at FDR p < 0.05

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.659171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:22.518202Z digest=sha256:6cace7cafeb697232d137ca8376a988d733143fd9c4858bf42ada8a653e284d9

Observation b50d41b6-b50e-4c2d-bff9-78862206375c · outbound

This paper cites an unresolved cited work.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:21:22.682747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:22.510706Z digest=sha256:e8c7e19af7fe824fb0af0920c20d3260badbb1ed9faf115e4dd094f8cfba9cd4

Observation 044c86dc-3609-493a-a98d-061e1f87a0a4 · outbound

This paper cites It may be that the visual system translates sensory inputs into modality-agnostic representations that reflect stable, relational patterns observed in the real-world environment.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities It may be that the visual system translates sensory inputs into modality-agnostic representations that reflect stable, relational patterns observed in the real-world environment

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.695493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:22.507342Z digest=sha256:d1205ccd4ba49aeae788f305b4901e8662cfc5551790c472b1e3562a45a630d7

Observation c1b187da-df2f-4334-a4a4-b53ba3204fdc · outbound

This paper cites The sentence captions were collected from five human annotators as part of the Microsoft Common Objects in Context database (Lin et al., 2014).

Representations in vision and language converge in a shared, multidimensional space of perceived similarities The sentence captions were collected from five human annotators as part of the Microsoft Common Objects in Context database (Lin et al., 2014)

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.671282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:22.514391Z digest=sha256:7903e94dd575ae16591380f81a9ca2af33687f3dcda5c0a5aec16859a581d780

Observation 8746e166-da76-43e8-851d-94d81dc74e1e · outbound

This paper cites an unresolved cited work.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:21:22.767941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:22.481066Z digest=sha256:23ffa5530938a70d081d667e8e55b63bbef21758a0b88dfffaec333db5f12a8a

Observation be82ce1a-0acd-48f6-9705-f40062c5cb0f · outbound

This paper cites an unresolved cited work.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities Unresolved cited work

Reference 4081

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:21:22.647477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:22.522047Z digest=sha256:3fb2e13b80c5c9b3a2553a731f4af09fc48b74770ffade2757c335190175a9c4

Observation 233de5ba-3820-4e13-b61c-6c4c11479083 · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 6241

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:22.540405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:22.540405Z digest=sha256:aa1f5d0441faf6dd38cb41a4633b2ea55e565194c7094da35d40a2a07554fea4

Observation 2112a0ed-47c6-4cc6-acf6-2f903dea97d9 · outbound

This paper cites Visual representations in the human brain are aligned with large language models.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities Visual representations in the human brain are aligned with large language models

Reference 9383

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:22.528735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:22.528735Z digest=sha256:817fb847b192eecef073ac28f305cca8c69e66e5742ad1241980bcce0b8906dd

Pith citing papers

Observation e30cad62-9bbc-4c50-a7ce-b698b07ceb1e · inbound

Disentangling the Factors of Convergence between Brains and Computer Vision Models cites this paper.

Disentangling the Factors of Convergence between Brains and Computer Vision Models Representations in vision and language converge in a shared, multidimensional space of perceived similarities

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-05T16:32:28.674050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T16:32:28.422930Z digest=sha256:f9838761c36c23dfd3e49d8d5f4d85e6f6f575408ffec83a48625dbb97f7eb67