Pith. sign in

Paper Citation Record · LEDGER

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation

As of 22 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2505.13498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13498 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:06:59.670819Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e353f138-250f-4b03-9e29-db2994e5ab16 · outbound

This paper cites Evaluating large vision-and-language models on children’s mathematical olympiads,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Evaluating large vision-and-language models on children’s mathematical olympiads,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:00.181948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.538738Z digest=sha256:4bfffd924e9af6817c7da2ba7b022a1e82f8bc7201e6b25b3a9d4638c1b77f2c

Observation 3f767a34-2456-4a86-b61b-ddd47f7b6b36 · outbound

This paper cites Towards reasoning in large language models: A survey,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Towards reasoning in large language models: A survey,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:00.168449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.542424Z digest=sha256:e689c0219b8448d837cd2daed2e4ea470718c76d33135ee1f6ac8ec790c513e1

Observation 98326dee-95f0-4cfa-af71-1f5066f24ff4 · outbound

This paper cites The bitter lesson learned from 2,000+ multilingual benchmarks,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation The bitter lesson learned from 2,000+ multilingual benchmarks,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:00.156588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.546150Z digest=sha256:d59fa7dbe7caf9e84415c5ea141ee11a20d8f223f99d5a00b4b1411b719f7251

Observation b418d6ba-a19e-419f-8bd9-4012ac3c1c4e · outbound

This paper cites Challenging the boundaries of reasoning: An olympiad-level math benchmark for large language models,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Challenging the boundaries of reasoning: An olympiad-level math benchmark for large language models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:00.144957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.550321Z digest=sha256:f03e92bbc3498eb3bb23fcb0cf659d40ee2b7854277fc428ced922912d9f272a

Observation 9cc1d127-af62-41ed-bdd0-8bb2927c4619 · outbound

This paper cites Humanity’s last exam,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Humanity’s last exam,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:00.134305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.554166Z digest=sha256:f8b37a653ce1c0987728a1b02bd2fe9cc1f6e51f7f4a7655a12970f7939cd0da

Observation c8500227-a87f-4036-9a47-50806b63f63d · outbound

This paper cites The multilingual mind : A survey of multilingual reasoning in language models,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation The multilingual mind : A survey of multilingual reasoning in language models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:00.122556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.558077Z digest=sha256:645d53add3c8007625d61500cb25611cc01d172e3ed85bd9afafd4f810860962

Observation 9b1decdb-fb12-4d6c-bf1f-63298839868b · outbound

This paper cites Scaling test-time compute for low-resource languages: Multilingual reasoning in llms,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Scaling test-time compute for low-resource languages: Multilingual reasoning in llms,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:00.108586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.562226Z digest=sha256:d21b4aa9e4ce57378ff0800118de3d8f243bdc2e836e3181ee633803c0f295ab

Observation 7a05bebb-d546-4099-9fa7-09d67df29cd9 · outbound

This paper cites Towards measuring and modeling “culture.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Towards measuring and modeling “culture

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:00.094920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.565798Z digest=sha256:5be3462abd7c89fff1c053798fa654a12f6ebc7974bb64a4cc0a606994998ec7

Observation 65badd25-f43d-44ce-b1d1-1282e6b31169 · outbound

This paper cites Memory of Peoples Series, UNESCO, 3 ed., Feb.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Memory of Peoples Series, UNESCO, 3 ed., Feb

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:00.083275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.569442Z digest=sha256:c8e077bdfbb9eb5c80d748d822bc462db33b93f85312afb171d1126de6002bef

Observation 3e918326-2133-4940-9df9-6cd4de5ec00a · outbound

This paper cites SIB- 200: A simple, inclusive, and big evaluation dataset for topic classification in 200+ languages and dialects,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation SIB- 200: A simple, inclusive, and big evaluation dataset for topic classification in 200+ languages and dialects,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:00.071489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.573559Z digest=sha256:cde95d47f2a8341fef35c248b52cd856f5295ae1d61388bff9d106abf726518c

Observation ff32c060-5313-4954-af16-1751887b485f · outbound

This paper cites The belebele benchmark: a parallel reading comprehension dataset in 122 language variants,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation The belebele benchmark: a parallel reading comprehension dataset in 122 language variants,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:00.060652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.577410Z digest=sha256:de665beb0d910ad82c3856e2bad3da18711c7ee3f461b8e4c6d88dc9fd0c120c

Observation cb4f1d79-23a6-4346-84b6-706864247b03 · outbound

This paper cites Uccix: Irish-excellence large language model,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Uccix: Irish-excellence large language model,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:00.049558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.581577Z digest=sha256:2757efa905e742cfc68e08694297239d5cb615a83e6196636b969eb543024a94

Observation 94372199-e63e-4ac2-a3e2-1908b4e85d68 · outbound

This paper cites Survey of cultural awareness in language models: Text and beyond,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Survey of cultural awareness in language models: Text and beyond,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:00.037295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.584971Z digest=sha256:ce025a4e42503acd5f81705894127ed195ae1c77d0165b5a598ad68cf4709840

Observation 6b2daa91-c408-4fd6-9906-67f924b15966 · outbound

This paper cites Judging LLM-as-a-judge with MT-bench and chatbot arena,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Judging LLM-as-a-judge with MT-bench and chatbot arena,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:00.019823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.588177Z digest=sha256:d9e474c7126eeaa7034d950a5857e3c2e3ef228084477b4e5b13c1e195e542fb

Observation d562f72e-d14c-49d4-9065-e1613f418b99 · outbound

This paper cites Justice or prejudice? quantifying biases in LLM-as-a-judge,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Justice or prejudice? quantifying biases in LLM-as-a-judge,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:00.006224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.591265Z digest=sha256:15b66f89a21a7299205f15e443996ba79f3bc1d9d3e3759002d0f1dbf3780f5d

Observation 676b34f5-4ec0-4eb7-96d4-3a333addebfa · outbound

This paper cites Language models are multilingual chain-of-thought reasoners,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Language models are multilingual chain-of-thought reasoners,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.991668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.594626Z digest=sha256:842e9f98cca8915120daceb3c299c81335a4f04b0c34d5ad0d85124cc4bed53f

Observation cb9bcbe7-0ba7-4f67-9fdd-0ae6e6f326f0 · outbound

This paper cites Global mmlu: Understanding and addressing cultural and linguistic biases in multilingual evaluation,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Global mmlu: Understanding and addressing cultural and linguistic biases in multilingual evaluation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.980367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.597963Z digest=sha256:26e7ffabbe57f0287eabb6b6ee270084091859b8422475666d4f75f61e053262

Observation eca2b40d-4bdf-45c5-a489-f3d63625147f · outbound

This paper cites M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.969920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.601748Z digest=sha256:2b39f122c68b1da35c4b67493b074b2960591fb746c634bb93d7d80d2a9add90

Observation 9702f91d-7bcc-4094-bdfc-b89415656684 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Training Verifiers to Solve Math Word Problems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:06:59.604840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:06:59.604840Z digest=sha256:6ee174cad0553ac635d7ea862e9c3494c3c0b4376b234fa4cb7ba720fecdf12a

Observation b938eb7b-4fa9-49b8-85fc-9b9ac3ae1ed1 · outbound

This paper cites Measuring massive multitask language understanding,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Measuring massive multitask language understanding,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.956929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.608475Z digest=sha256:c438f24bed733e1d45e9f4e33e728caee0f65741f773ad549dc7d3db51ec9231

Observation 74846c90-94cf-4c22-a4ad-ccf8b9018d09 · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation CMMLU: Measuring massive multitask language understanding in Chinese,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.943675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.612067Z digest=sha256:80bd4a5666966783cc222d8241c27d80fcfab1340a770810da0621cb4c499861

Observation 2ed1a92a-772a-4ea1-9279-315c66d6fad8 · outbound

This paper cites KMMLU: Measuring massive multitask language understanding in Korean,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation KMMLU: Measuring massive multitask language understanding in Korean,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.931047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.615201Z digest=sha256:fc6ca2175ede8619b0e160336dc874a4eb4145fb86081efba082391b50682aa4

Observation e80bcadc-f996-4193-8989-e01a5157dcca · outbound

This paper cites ArabicMMLU: Assessing massive multitask language understanding in Arabic,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation ArabicMMLU: Assessing massive multitask language understanding in Arabic,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.917434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.618272Z digest=sha256:97fd1b2f4d72e35bca02773fff0ed9c1e33467090c7c8ee694f8ab247e591918

Observation c98c950c-70c1-479d-a39e-e9003b5ae395 · outbound

This paper cites None of the others: a general technique to distinguish reasoning from memorization in multiple-choice llm evaluation benchmarks,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation None of the others: a general technique to distinguish reasoning from memorization in multiple-choice llm evaluation benchmarks,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.902906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.621749Z digest=sha256:ff1e1928831f57017b1ccaf6ef16f1c89e6705beabb71c2cf546ae271585a7b2

Observation 8ebe5e9b-80b4-4daa-8e32-3a5ab1817e8a · outbound

This paper cites SPIQA: A dataset for multimodal ques- tion answering on scientific papers,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation SPIQA: A dataset for multimodal ques- tion answering on scientific papers,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.889708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.625246Z digest=sha256:f446ee1732ff55e4260b0b47828111b2e40780981000fc05a56c592958b00aac

Observation a24f8483-3c3b-40c5-b657-688e93adfc29 · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Introducing gemini 2.0: our new ai model for the agentic era,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.877164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.629230Z digest=sha256:c9cd0b61c2ccfe7097fdcde9cf0ee2540f060e0b6184dc2b0277fec1c7b97d3e

Observation 4d7baa99-52d9-4371-ae1e-c4f89d245b0c · outbound

This paper cites Gemini 2.0: Flash, flash-lite and pro,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Gemini 2.0: Flash, flash-lite and pro,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.863503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.633286Z digest=sha256:3cecc64ce75a2df7a3f05c453889b8b154cfba9881147ee783316a9269620546

Observation cc371727-9f41-49ed-a630-d8f3116adfcb · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Bleu: a method for automatic evaluation of machine translation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.849362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.637542Z digest=sha256:515a99f4fec979e817b105ac4c67f8179f2fd5ffe5f416fcc855e89c0c0ef324

Observation 0d321d6a-e496-4656-b662-0f001065d8ff · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation ROUGE: A package for automatic evaluation of summaries,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:06:59.641490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:06:59.641490Z digest=sha256:670d21c9ad34e9641c6c7840efd7f0f4aba152e177806b84e535a2a61e874ac7

Observation b22831e4-866f-41ec-9ef3-4e51fefe0a03 · outbound

This paper cites LIMA: Less is more for alignment,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation LIMA: Less is more for alignment,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.823380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.644955Z digest=sha256:99ddc868a5af943501d4a51de20865c4b6f9d57b4a49131f8d382036e8aee79e

Observation 2e9c4103-e71e-4883-a8a9-a4efed9b5f14 · outbound

This paper cites Bag of Tricks for Efficient Text Classification.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Bag of Tricks for Efficient Text Classification

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:06:59.648469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:06:59.648469Z digest=sha256:197eaa45e5d030e4bb776eeb92ee02cfc7e099cfe4061c247c3b1dc5bbe61d88

Observation df0fc659-4e43-45cf-8ddb-97d2566cdb01 · outbound

This paper cites FastText.zip: Compressing text classification models.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation FastText.zip: Compressing text classification models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:06:59.652350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:06:59.652350Z digest=sha256:c476679ba3809ae22d3e619650e6e4754767888d806313f498953ffa6f59387a

Observation 568b6873-8f3f-49f3-a5d3-90ef9d205ed8 · outbound

This paper cites Introducing openai o3 and o4-mini,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Introducing openai o3 and o4-mini,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.808243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.656580Z digest=sha256:36671de2dbcffdc08c871ce7fd3973e082c62a9c604a7c8160f48841d2e26c3b

Observation 9052b625-9441-44f0-80b0-3dd0ed6fdc19 · outbound

This paper cites Introducing gpt-4.1 in the api,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Introducing gpt-4.1 in the api,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.792246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.660179Z digest=sha256:3f2c82e9e063de6adb6fd3d100cc7f3d0bbf70f63fdac602bb1556fe21dd4afb

Observation 820c2407-65b8-493a-bbf0-4fd9e43c1d6d · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation The llama 4 herd: The beginning of a new era of natively multimodal ai innovation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.777059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.663810Z digest=sha256:bf9ad458ac298e18737ebf117908c23966ba22aefb302ef176e487c74d6d016a

Observation 3abb851c-3f92-4017-a412-687421e7c253 · outbound

This paper cites Aya vision: Advancing the frontier of multilingual multimodality,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Aya vision: Advancing the frontier of multilingual multimodality,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.759638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.667596Z digest=sha256:22ce806066f0bc8d145b38496729acfb218612188e0dd22eaac015ae99a13038

Observation 9aa5105c-3e93-495d-bb01-9b7076063a37 · outbound

This paper cites Measuring short-form factuality in large language models,.

IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Measuring short-form factuality in large language models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:06:59.746725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T21:06:59.670819Z digest=sha256:3187efed2931aebd8cb5178d30151e1b2b8b481394e0b5b601a767452b3cff2f

Pith citing papers

No inbound Pith citation observations are available.