Pith. sign in

Paper Citation Record · LEDGER

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation

As of 14 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2605.02035.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.02035 v2

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T23:56:42.264337Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T05:09:49.727492Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T16:38:40.580304Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 40070c17-38a2-4bbd-9434-1f7bbc3b90b3 · outbound

This paper cites GPT-4o System Card.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation GPT-4o System Card

Reference 1

Resolution
malformed identifier
local_arxiv, observed 2026-07-01T00:55:12.245563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:06caa2593d1e732ffe959eb0297f19f9b27dee6fb5609aa67742531e5faab190

Observation 6bcec1fa-d70d-46f7-aec1-d6ac779e31c0 · outbound

This paper cites Describe how these elements con- nect to the text.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation Describe how these elements con- nect to the text

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:43:36.966597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:047d5804c72696aa42b036c584b69abc968fff409885d660b0cf48d6887cfb87

Observation ffa91c4a-1c91-4391-a7c7-77da7879cccf · outbound

This paper cites an unresolved cited work.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-07-07T09:43:36.972270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:6e8e3e145c1391c5ed16569b71f98f56a99b7f0a0034ca064372580e64235489

Observation baee153e-901a-425d-ad53-6c829253415e · outbound

This paper cites an unresolved cited work.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-07-07T09:43:36.979477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:a076eaeeef68f53a854c9707ce5be0e97e0fea4a780d30ce9915f1b1e53dc85e

Observation 29574f6d-5bd3-46c6-8a09-4f6879ed673f · outbound

This paper cites While visual grounding establishes a mapping between the image and the text, the initial translation can still leave some ambiguities un- resolved.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation While visual grounding establishes a mapping between the image and the text, the initial translation can still leave some ambiguities un- resolved

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:43:36.977469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:742039aaf39ff6da093130faab70c754d2dbc6dbdcd560ec9b9d05ff8cde42c9

Observation 875eb38f-105c-430d-a7f0-a4c8ea1d5f0c · outbound

This paper cites This constraint prevents unnecessary modifications to the sen- tence structure and helps maintain overall translation fluency.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation This constraint prevents unnecessary modifications to the sen- tence structure and helps maintain overall translation fluency

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:43:36.981451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:c9db26474ff071cc9c86d31347d39b1b612ca8b996cca20779c8416db9b43df6

Observation b48a8437-c4ee-4ce1-b567-5695632c63c7 · outbound

This paper cites object", which requires a concrete translation (.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation object", which requires a concrete translation (

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:43:36.974018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:1f33c3c099388f7926fed3bcd5557982a225a03cff9be7dd03c39bc2931a5419

Observation ad34308f-39b0-4f85-9cfb-345dc6f6df02 · outbound

This paper cites an unresolved cited work.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-07-07T09:43:36.968371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:e6f3abf6ba0473177c84d314ab7a273b4a1972c21ffb554a2e09ffe83852cd16

Observation 221f054e-3497-40db-8e25-77551046cd4e · outbound

This paper cites - The reason it is ambiguous (how multiple interpretations can arise).

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation - The reason it is ambiguous (how multiple interpretations can arise)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:43:36.960769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:f549a69d23f7f5a83a9f88584f1d55df061539447ece2bf02c8835bd9801b005

Observation 06bdd0a7-da72-4e84-b40a-cc6ebabc00e2 · outbound

This paper cites bank" = financial institution vs. river bank). - Syntactic: the sentence structure permits multiple interpretations (e.g.,.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation bank" = financial institution vs. river bank). - Syntactic: the sentence structure permits multiple interpretations (e.g.,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:43:36.957786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:0c05cdf567823a249d7e41c7f4da0c2c4d44a921992c2e44bc7b40919482e9ea

Observation d8ede4db-61a6-4fbd-a108-621e0677ad06 · outbound

This paper cites - If two ambiguities from different models share the same type and describe the same underlying issue, merge them.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation - If two ambiguities from different models share the same type and describe the same underlying issue, merge them

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:43:36.962785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:d6c62bc5500aa4ee9dd04643f044135e48c8c29e0599adcfec6ac11046e55ae5

Observation 74919a61-825b-479d-b46e-c9b955b201e0 · outbound

This paper cites - translations: union the translation candidates from both sources, removing exact duplicates.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation - translations: union the translation candidates from both sources, removing exact duplicates

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:43:36.964745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:b6bd48e89c0a2f1da0bc9c07fbd32346bcf9dc89c5cfa9b90e5033c880d7e127

Observation ab2db17e-0a7b-487a-8971-9a3cb49c586f · outbound

This paper cites en" field), extract the literal word(s) or phrase(s) that cause each ambiguity. - Save them into a new field.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation en" field), extract the literal word(s) or phrase(s) that cause each ambiguity. - Save them into a new field

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:43:36.970442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:3fe093134e1f39a475f36423168ccad6530000ff402171d1a352e5a99df9de7d

Observation 152725d9-c327-405e-9904-14e46a9718b6 · outbound

This paper cites an unresolved cited work.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-07-07T09:43:36.949963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:d3cbf8eea432d258e7a48dedd43999a110aba6fce9427d33ce995980e716c943

Observation 5a22858b-d367-42a3-86ce-ae2d2519272f · outbound

This paper cites an unresolved cited work.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-07-07T09:43:36.952488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:57d81dbf15b68625b9a629d9e767d5a28eeb254a922af3d4898f220522aadd92

Observation c2b33d34-f977-4857-8464-8ddb8647653a · outbound

This paper cites translation_zh.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation translation_zh

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:43:36.954434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:419dc53e4a22c53ec042447838e98daa4daf8c1031c5fa2e91a480261dda5f24

Observation 1dcab501-e346-41ce-b94f-22acb1b1f775 · outbound

This paper cites an unresolved cited work.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-07-07T09:43:36.956081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:820986acc4f61382039945e116dbeab1c42c316473eb4135da577e270013fd04

Observation 7b87a4d7-ee60-48cc-882b-c10fcc79cae5 · outbound

This paper cites an unresolved cited work.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-07-07T09:43:36.942773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:6c3f7968420898cd9f742f28797c9b6cd94d845c9165af5e78917dfc7aa042ed

Observation a777345e-1306-4bbc-9c06-b4fb0b6d47a2 · outbound

This paper cites an unresolved cited work.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-07-07T09:43:36.944875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:bb4b0511aee530b39e2b71bc7062c219910860762f23a3b06f6ae5c23b47ab81

Observation 4d764595-56c8-4f82-b276-7baa8454d721 · outbound

This paper cites an unresolved cited work.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-07-07T09:43:36.948271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:f62f1d875aec9a5219b4d0b1d3a8330dc3a1e62a8d93eca2171b38a47aa33277

Observation 8f997998-ade7-4145-95ec-d873526a864a · outbound

This paper cites Correct". - If the meaning is missing or distorted, return.

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation Correct". - If the meaning is missing or distorted, return

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:43:36.946680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:56:42.264337Z digest=sha256:bebb74dc1c302ca3fb1a58bf21e98e0bd398411990b69dbd16559c7b5082d25c

Pith citing papers

Observation 8cc521f9-df80-4b74-87f6-cc2b362792ab · inbound

IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products cites this paper.

IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T16:38:40.581750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T05:09:49.727492Z digest=sha256:88eb868a09e4fa194746be28733905965e05e1ae8f6fafc9dd08c308ae14f5f4