Pith. sign in

Paper Citation Record · LEDGER

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu

As of 8 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2508.01006.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.01006 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T06:00:10.327109Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 548a52aa-2253-4475-81d1-0927df9c063e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.281920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.281920Z digest=sha256:938c2cdbf5bf925aa33d176c22873737196dc1079b1f36ec2ac15178c13362fb

Observation 47537837-aa1d-4b6d-8a6b-7c73a663c655 · outbound

This paper cites MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.296880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.296880Z digest=sha256:4084714a4cf2d4b38890be177e918e54880a69866ae1381823b7e994608a2275

Observation f6ac2c71-d32d-4fac-8f5a-8aa11ec5954a · outbound

This paper cites Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T06:00:10.374786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T06:00:10.300890Z digest=sha256:d48ffccc327a8f9ded0a6bd0823445f41150c85d766d021e76837990aebda8da

Observation 9b21b02a-1d0d-4cfb-b3bb-a103a75aaca7 · outbound

This paper cites In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175–184, Online and Punta Cana, Dominican Republic.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175–184, Online and Punta Cana, Dominican Republic

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:17.008156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T06:00:10.304908Z digest=sha256:b24719f6015d4b0f3b9fed616fed9aa63980a975e2ad5a75939fc44457e5dba4

Observation 325c05ce-bb73-4f46-91eb-d7f90fe50223 · outbound

This paper cites In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 4606–4634, Abu Dhabi, United Arab Emirates.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 4606–4634, Abu Dhabi, United Arab Emirates

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:16.984655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T06:00:10.316058Z digest=sha256:d82825d9742b85a6aea9d8b9faa4a2a91c6c16254413ba187aa285005d393b74

Observation 658c8df3-eb7a-4f0c-90ec-0c39dc780deb · outbound

This paper cites https: //huggingface.co/large-traversaal/ Alif-1.0-8B-Instruct.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu https: //huggingface.co/large-traversaal/ Alif-1.0-8B-Instruct

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:16.973213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T06:00:10.319635Z digest=sha256:93f47ff1ee77bae1640bf0d35356eb38cc16a56f112f7d722a51ad38b1d7b57c

Observation d6c2eb2c-9940-4d45-9194-ce4a9ebbffa2 · outbound

This paper cites In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:16.962267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T06:00:10.323430Z digest=sha256:570e98e2ac50efbcd9c1840038c3981d611d7f0796bd158bca4bed2f693bc93f

Observation 1a95b647-5a3a-400a-b0b8-921d650a2b73 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.285644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.285644Z digest=sha256:a28a13e9b9076adff862809c0272f17650ee08e49ec803332c099e97ec43ff8f

Observation 833c7619-d4b4-4f6c-8b1c-ef0c6e9dde59 · outbound

This paper cites an unresolved cited work.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-06T06:00:16.950764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T06:00:10.327109Z digest=sha256:c334a7005cea22f20c49572adffa284ec3a33384b6113eca4706b711817d86b6

Observation 903cf07f-df61-470c-a178-bc27745dfb5a · outbound

This paper cites In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7038–7051, Online and Punta Cana, Dominican Republic.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7038–7051, Online and Punta Cana, Dominican Republic

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:17.020156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T06:00:10.293383Z digest=sha256:89548b49e6edb920a40bb0812efdbd86e1bf758bf37b9038780a24af2564c730

Observation 62d44409-ffa4-4a70-acb5-b3d8dcc8b068 · outbound

This paper cites Crosslingual Generalization through Multitask Finetuning.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu Crosslingual Generalization through Multitask Finetuning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.308640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.308640Z digest=sha256:1372f34364f9fb159e7baee0ce8799ee4fa7b4a79c7312aa23d430a286422182

Observation 12c83f7b-ce1b-4839-b665-57a31b8abe16 · outbound

This paper cites In Findings of the Association for Computational Linguistics: EACL 2023 , pages 1581–1594, Dubrovnik, Croatia.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu In Findings of the Association for Computational Linguistics: EACL 2023 , pages 1581–1594, Dubrovnik, Croatia

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:16.995724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T06:00:10.312630Z digest=sha256:f007d1603ef7f7d17e98f4812a90852da6fc369a49f48222a6b0fb11b1cbacf7

Observation c3f58f95-83ac-4dde-b071-2c9f97a73365 · outbound

This paper cites The Llama 3 Herd of Models.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.289493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.289493Z digest=sha256:3912b66eaa341fff8ce25a73d968adc97e8043b4584d525f0b6e034f27e4a936

Observation b6b58527-6c82-42f4-836d-554e91e84c54 · outbound

This paper cites Preprint, arXiv:2506.13487.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu Preprint, arXiv:2506.13487

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.277918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.277918Z digest=sha256:323dc2aab36185d11fbf78a0cbf6a1bba3163eeb5dfebcc9c5010bd561c84421

Pith citing papers

No inbound Pith citation observations are available.