Pith. sign in

Paper Citation Record · LEDGER

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean

As of 10 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2506.01237.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01237 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:52:27.102294Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved2
  • parse uncertain0
  • malformed identifier5
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f0316d5-ae1a-4f30-869e-e41bb2a62d79 · outbound

This paper cites It en- hances the connection of knowledge between Korean and English.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It en- hances the connection of knowledge between Korean and English

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.730460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.428609Z digest=sha256:ccfdc74bb46d1632a0640f08b402f5e6bc40e3e4aa0d190c22a6034ba2e526bb

Observation b4fbe011-facb-4ee0-99dd-25166b86b62f · outbound

This paper cites It is trained by utilizing instruction fine-tuning methods, including su- pervised fine-tuning (SFT) and direct pref- erence optimization (DPO) (Rafailov et al., 2023).

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It is trained by utilizing instruction fine-tuning methods, including su- pervised fine-tuning (SFT) and direct pref- erence optimization (DPO) (Rafailov et al., 2023)

Reference 2

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:52:27.723264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.433432Z digest=sha256:14dcda103ce1a72c7a7c82a0e248d6334c59689d39c9e9f4e59172322db8ae39

Observation c812159d-e95b-46b6-981d-bed8c6b667ea · outbound

This paper cites This model is trained using the data, such as Peng et al.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean This model is trained using the data, such as Peng et al

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:52:27.715527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.456655Z digest=sha256:45a0a5d80c38677537a1060d06d2ebdad4fae3cd04b6fff88445ad0d39f146fb

Observation 5f6d9de8-2242-49f1-b406-a8412967b2e5 · outbound

This paper cites Specifically, this model is trained by utilizing DPO.8.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean Specifically, this model is trained by utilizing DPO.8

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:52:27.708136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.490491Z digest=sha256:6a3d22a518e5479ef84b536253a0cc41e5ed52b86b16ad0248ef713798d5cc51

Observation 518ea0b2-6e91-46cc-bc74-a466d4c6a52c · outbound

This paper cites This model is trained to utilize the system prompt.9.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean This model is trained to utilize the system prompt.9

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.699090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.538486Z digest=sha256:46d3b23101eccc647fdde42e014a50bc0469114ce770a6e91939915f3cfce197

Observation 8479455c-00a3-433e-b1c0-e2e6f97e257a · outbound

This paper cites 7https://huggingface.co/yanolja/ EEVE-Korean-10.8B-v1.0 8https://github.com/axolotl-ai-cloud/ axolotl 9https://huggingface.co/LGAI-EXAONE/ EXAONE-3.5-7.8B-Instruct.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean 7https://huggingface.co/yanolja/ EEVE-Korean-10.8B-v1.0 8https://github.com/axolotl-ai-cloud/ axolotl 9https://huggingface.co/LGAI-EXAONE/ EXAONE-3.5-7.8B-Instruct

Reference 6

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:52:27.689699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.562866Z digest=sha256:25b11bde9bedac92c9eee5b681a5eec821517cb46ecb85691a7d14ea2c90e192

Observation 051974d3-4825-47bb-9768-a22ea594bd1e · outbound

This paper cites an unresolved cited work.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:52:27.679639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.597483Z digest=sha256:1f389af12c9e08c84df17118587fbd176465e159533bae1727051e89e800bcd9

Observation 168d8a11-122d-4271-9d8d-79a1ef3631c0 · outbound

This paper cites • English-centric LLMs.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean • English-centric LLMs

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.671608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.640732Z digest=sha256:faabdd47073dcf374aa18744946269e3c2fcc387034a49c8b891316cc97218c5

Observation 4095267a-bd6a-4cb6-9048-c97d9402afe2 · outbound

This paper cites It has been instruction-tuned to enhance its performance in various natural language understanding and generation tasks.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It has been instruction-tuned to enhance its performance in various natural language understanding and generation tasks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.500318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.950777Z digest=sha256:9478c6074fbf3b3088d1257c71c081c580463ef27c54229e43e476e76450c449

Observation b4a64fad-c225-4583-8b65-914b827f0f53 · outbound

This paper cites It has been pre-trained on approximately 15 trillion tokens from publicly available sources.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It has been pre-trained on approximately 15 trillion tokens from publicly available sources

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.427842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.976613Z digest=sha256:580a73c28547999d7531706681175cac5fb96144eed006036b38edad32a8abd5

Observation 47328389-96c7-4866-ab9e-491ccb661216 · outbound

This paper cites It demonstrates superior performance in general knowledge and reasoning tasks, achieving high scores on benchmarks such as MMLU-Pro and MMLU- redux.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It demonstrates superior performance in general knowledge and reasoning tasks, achieving high scores on benchmarks such as MMLU-Pro and MMLU- redux

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.360007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:27.017017Z digest=sha256:d16a67b7c8289a80a54d58ad765ab3f6ff1c3218a2bad374bc3c945f786e1898

Observation 85b91125-aa42-4fd5-87c1-53a1bd672ee1 · outbound

This paper cites It is a text-to-text, decoder-only large language model, with open weights for both pre-trained and instruction-tuned vari- ants.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It is a text-to-text, decoder-only large language model, with open weights for both pre-trained and instruction-tuned vari- ants

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.664304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.678712Z digest=sha256:5425a9bae4da7dab1ce7827c34579b3d500b9cc2a2b44f6f934653ceacb7d49b

Observation 7418b453-a3cd-454e-bcd5-d6ef8f578d43 · outbound

This paper cites It is de- signed to deliver high performance across var- ious natural language processing tasks, bene- fiting from its extensive parameter count and advanced training methodologies.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It is de- signed to deliver high performance across var- ious natural language processing tasks, bene- fiting from its extensive parameter count and advanced training methodologies

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.656851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.719055Z digest=sha256:4340158ea54352fcf9fea191b322d790e4c05a0a00609fbc8b1e03673711f598

Observation c98ca859-b053-451f-816d-de4da0096b49 · outbound

This paper cites It is designed to handle various nat- ural language understanding and generation tasks, supporting multiple languages, includ- ing English and Chinese.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It is designed to handle various nat- ural language understanding and generation tasks, supporting multiple languages, includ- ing English and Chinese

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.647836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.747330Z digest=sha256:b6590943af08fa186ed8499a01e58b6515d2f90025de0bb9eacdec2f35d46151

Observation 04c0b046-ae23-4dec-9e32-89630f5abdd0 · outbound

This paper cites It offers enhanced performance in lan- guage understanding and generation tasks, with support for multiple languages and a con- text length of up to 128,000 tokens.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It offers enhanced performance in lan- guage understanding and generation tasks, with support for multiple languages and a con- text length of up to 128,000 tokens

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.640040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.783387Z digest=sha256:1e28c37ac1a39832147d614e943f28dd59e3039ed740d39b96c3653c83dd3abe

Observation 800f19c3-7416-43bf-bd88-f70be22a9ea8 · outbound

This paper cites It sup- ports a context length of up to 128,000 to- kens and is designed to handle complex tasks across multiple languages, including English and Chinese.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It sup- ports a context length of up to 128,000 to- kens and is designed to handle complex tasks across multiple languages, including English and Chinese

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.632240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.821689Z digest=sha256:d5b26d962b780ce394850766fea8453dbdd5ad197f9cf2612937df156fdf5227

Observation 09094eab-8992-4983-8325-a1867a8cc242 · outbound

This paper cites It has been fine-tuned using reasoning data generated by DeepSeek-R1, resulting in enhanced performance in reasoning tasks.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It has been fine-tuned using reasoning data generated by DeepSeek-R1, resulting in enhanced performance in reasoning tasks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.623776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.854506Z digest=sha256:8e4081ba9aef796795b4bee969fc71ad2d3f337129a339137359306746c018d4

Observation 59fabfb4-b184-42fd-a180-dd4aef9fb392 · outbound

This paper cites It has been fine- tuned with reasoning data from DeepSeek-R1, achieving state-of-the-art results in various benchmarks.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It has been fine- tuned with reasoning data from DeepSeek-R1, achieving state-of-the-art results in various benchmarks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.615741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.875686Z digest=sha256:57896e759a8fa3cb3db72ff639c12a7f7811f72105d448d7c5108daaefaaa070

Observation 4b8f039f-1511-4e09-81c5-8e60edabc237 · outbound

This paper cites budget forcing.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean budget forcing

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.561205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.915996Z digest=sha256:32fd4feee28ae978129308d47dedf7972f9d3192f13794b233e950acf5d7035b

Observation 9fe0ef9f-69e3-4443-8611-ccc82dc63ae1 · outbound

This paper cites It is designed for real-time interactions and has been integrated into various Google prod- ucts, including Bard and Pixel smartphones.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It is designed for real-time interactions and has been integrated into various Google prod- ucts, including Bard and Pixel smartphones

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.340272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:27.046501Z digest=sha256:db74aaf91d49691b933f1d3711d964705ce1b447955fe4bb2ec1c11310839512

Observation 38240ac7-0be4-4170-8aa5-efc88dfaaa20 · outbound

This paper cites It introduces fea- tures such as a Multimodal Live API for real- time audio and video interactions, enhanced spatial understanding, and integrated tool use, including Google Search.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It introduces fea- tures such as a Multimodal Live API for real- time audio and video interactions, enhanced spatial understanding, and integrated tool use, including Google Search

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.324463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:27.057472Z digest=sha256:4583d01f3b0ad7690ac86a6aace385fca0b29ae9b3c55ff5b0453710be934104

Observation 1de16268-a40c-4b2c-a3c9-9076e88de300 · outbound

This paper cites It features a context window of up to 200,000 tokens, allowing it to process extensive text sequences effectively.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It features a context window of up to 200,000 tokens, allowing it to process extensive text sequences effectively

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.309300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:27.074131Z digest=sha256:1bdae19d2daede198a00937b83abcfaf04ccb768b4e24210e3cd84452fcadac1

Observation 3b4953cd-dcdb-45c9-b83b-c27fb5feed45 · outbound

This paper cites It maintains a large context window and has been fine-tuned for better alignment with hu- man preferences.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It maintains a large context window and has been fine-tuned for better alignment with hu- man preferences

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.296196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:27.081256Z digest=sha256:81641bf53f331a723e5243ecaa71470e973b2f94791b2198d992c0403c3b63c3

Observation 0b144209-f3ef-40e4-9f9e-0e64933b5e0c · outbound

This paper cites It has been widely used in applications requiring natural language understanding and generation.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It has been widely used in applications requiring natural language understanding and generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.280688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:27.086333Z digest=sha256:57c8cb2e52ce9c848d2399778233ae3238f2ac7fd2288c1baa175c42a43c1a68

Observation 7dc489f5-39c2-4c42-ba18-830aaa88d5ef · outbound

This paper cites It offers rapid response times and has been integrated into various applications for real-time interactions.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It offers rapid response times and has been integrated into various applications for real-time interactions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.258217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:27.092680Z digest=sha256:d78026059f17d05f433f6965dbbe1eda021c05ea52cd31e164d8b9400cae28c0

Observation 48ef840a-2471-4799-a338-842f4de48ef2 · outbound

This paper cites It ex- hibits rapid response times comparable to hu- man reactions and has enhanced performance in non-English languages.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It ex- hibits rapid response times comparable to hu- man reactions and has enhanced performance in non-English languages

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.243231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:27.096971Z digest=sha256:a80b180d443da44d264920cae581973b10286558d88f8356ced2226176fb2c9a

Observation 13635027-7000-4c17-a852-4bcea4e4b444 · outbound

This paper cites think- ing.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean think- ing

Reference 30

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:52:27.226564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:27.102294Z digest=sha256:044ba3542564ddedff6d1a4c728a382e874e53178c1907b16dfae291f1716829

Observation 2e1985cb-ce9c-4add-8f49-8ba247f9b960 · outbound

This paper cites In Proceedings of the 58th Annual Meet- ing of the Association for Computational Linguistics , pages 4609–4622, Online.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean In Proceedings of the 58th Annual Meet- ing of the Association for Computational Linguistics , pages 4609–4622, Online

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.747529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.374728Z digest=sha256:28bb9b47c77a6cc908f360d993d3f3306544b370abcb073bdf00fe571ff75f7e

Observation a52da19e-6583-4709-9fa5-683c456159db · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:26.406804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:26.406804Z digest=sha256:9cc34f1b9519cab17e9ff05bd11cc4cc652833a28af92778a79391c58d2b4a63

Observation b35ffa67-b76c-471a-ac2b-53c49fff9a1e · outbound

This paper cites In Findings of the As- sociation for Computational Linguistics: ACL 2024 , pages 12159–12173, Bangkok, Thailand.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean In Findings of the As- sociation for Computational Linguistics: ACL 2024 , pages 12159–12173, Bangkok, Thailand

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.739596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:52:26.419737Z digest=sha256:5f16a9453792eae1beca01d1166bb06e44557e8b9c5ad323baab28938533ead0

Pith citing papers

No inbound Pith citation observations are available.