Pith. sign in

Paper Citation Record · LEDGER

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean

As of 8 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2506.01237.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01237 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:52:27.102294Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved2
  • parse uncertain0
  • malformed identifier5
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f0316d5-ae1a-4f30-869e-e41bb2a62d79 · outbound

This paper cites It en- hances the connection of knowledge between Korean and English.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It en- hances the connection of knowledge between Korean and English

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.730460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.428609Z digest=sha256:614c4fa43f270d90ac976ed9edbbf5376deb797e420fdde7f5f84ced4310b8b9

Observation b4fbe011-facb-4ee0-99dd-25166b86b62f · outbound

This paper cites It is trained by utilizing instruction fine-tuning methods, including su- pervised fine-tuning (SFT) and direct pref- erence optimization (DPO) (Rafailov et al., 2023).

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It is trained by utilizing instruction fine-tuning methods, including su- pervised fine-tuning (SFT) and direct pref- erence optimization (DPO) (Rafailov et al., 2023)

Reference 2

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:52:27.723264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.433432Z digest=sha256:c2053944e60057c25f56eec96f710438b4281187832694d96bc86f7946077737

Observation c812159d-e95b-46b6-981d-bed8c6b667ea · outbound

This paper cites This model is trained using the data, such as Peng et al.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean This model is trained using the data, such as Peng et al

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:52:27.715527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.456655Z digest=sha256:e7eac1d86560b0daac1def07ae8231444127333466846236916b83e844d42964

Observation 5f6d9de8-2242-49f1-b406-a8412967b2e5 · outbound

This paper cites Specifically, this model is trained by utilizing DPO.8.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean Specifically, this model is trained by utilizing DPO.8

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:52:27.708136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.490491Z digest=sha256:debaa1f081f3754852b968034111d56dd1e8089d4533fbf57ca89c103672a828

Observation 518ea0b2-6e91-46cc-bc74-a466d4c6a52c · outbound

This paper cites This model is trained to utilize the system prompt.9.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean This model is trained to utilize the system prompt.9

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.699090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.538486Z digest=sha256:aa71c4a9c4d96000dd2c912b8f360e1f23d53becf825d9dcc1eafe9d4176acd3

Observation 8479455c-00a3-433e-b1c0-e2e6f97e257a · outbound

This paper cites 7https://huggingface.co/yanolja/ EEVE-Korean-10.8B-v1.0 8https://github.com/axolotl-ai-cloud/ axolotl 9https://huggingface.co/LGAI-EXAONE/ EXAONE-3.5-7.8B-Instruct.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean 7https://huggingface.co/yanolja/ EEVE-Korean-10.8B-v1.0 8https://github.com/axolotl-ai-cloud/ axolotl 9https://huggingface.co/LGAI-EXAONE/ EXAONE-3.5-7.8B-Instruct

Reference 6

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:52:27.689699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.562866Z digest=sha256:98bbde80685bb37f9ea273681230dfba07835f08f08dd5ac7989b67d9ac13d2d

Observation 051974d3-4825-47bb-9768-a22ea594bd1e · outbound

This paper cites an unresolved cited work.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:52:27.679639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.597483Z digest=sha256:eb11b65f00b9cfbaab2df6197a1647f329501359282f404a38513ffaf0620d00

Observation 168d8a11-122d-4271-9d8d-79a1ef3631c0 · outbound

This paper cites • English-centric LLMs.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean • English-centric LLMs

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.671608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.640732Z digest=sha256:0fb4d192998d5de56ee3e014ae9416adb922a386e3575c888d70b67875a15401

Observation 4095267a-bd6a-4cb6-9048-c97d9402afe2 · outbound

This paper cites It has been instruction-tuned to enhance its performance in various natural language understanding and generation tasks.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It has been instruction-tuned to enhance its performance in various natural language understanding and generation tasks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.500318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.950777Z digest=sha256:10b1fb1c695a105b108beabbb3643f35e0a5037725ee44870c94a872d69cad02

Observation b4a64fad-c225-4583-8b65-914b827f0f53 · outbound

This paper cites It has been pre-trained on approximately 15 trillion tokens from publicly available sources.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It has been pre-trained on approximately 15 trillion tokens from publicly available sources

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.427842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.976613Z digest=sha256:d4d2282668f52beb86edd2292cc8ef4d44cb27032765ea1c392b8a7c0b8dd226

Observation 47328389-96c7-4866-ab9e-491ccb661216 · outbound

This paper cites It demonstrates superior performance in general knowledge and reasoning tasks, achieving high scores on benchmarks such as MMLU-Pro and MMLU- redux.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It demonstrates superior performance in general knowledge and reasoning tasks, achieving high scores on benchmarks such as MMLU-Pro and MMLU- redux

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.360007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:27.017017Z digest=sha256:8101c22f1cad67327bb4e092b710e04518b01061015fe40ed81d95c39a255a6e

Observation 85b91125-aa42-4fd5-87c1-53a1bd672ee1 · outbound

This paper cites It is a text-to-text, decoder-only large language model, with open weights for both pre-trained and instruction-tuned vari- ants.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It is a text-to-text, decoder-only large language model, with open weights for both pre-trained and instruction-tuned vari- ants

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.664304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.678712Z digest=sha256:bea0ec0339a5c46faaf216f14b9f816fc17bdc5d0e796ea5ddaa73236e4b293e

Observation 7418b453-a3cd-454e-bcd5-d6ef8f578d43 · outbound

This paper cites It is de- signed to deliver high performance across var- ious natural language processing tasks, bene- fiting from its extensive parameter count and advanced training methodologies.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It is de- signed to deliver high performance across var- ious natural language processing tasks, bene- fiting from its extensive parameter count and advanced training methodologies

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.656851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.719055Z digest=sha256:d22a8bdfe22b88c1d6bd0a030db78d0169aed959dbf8ee6bfdb488d1cb0427e8

Observation c98ca859-b053-451f-816d-de4da0096b49 · outbound

This paper cites It is designed to handle various nat- ural language understanding and generation tasks, supporting multiple languages, includ- ing English and Chinese.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It is designed to handle various nat- ural language understanding and generation tasks, supporting multiple languages, includ- ing English and Chinese

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.647836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.747330Z digest=sha256:5dd87c216e82ba3cc05e423c465d8b9886d5d5260240a3cd5dc33c151204c5c3

Observation 04c0b046-ae23-4dec-9e32-89630f5abdd0 · outbound

This paper cites It offers enhanced performance in lan- guage understanding and generation tasks, with support for multiple languages and a con- text length of up to 128,000 tokens.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It offers enhanced performance in lan- guage understanding and generation tasks, with support for multiple languages and a con- text length of up to 128,000 tokens

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.640040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.783387Z digest=sha256:d4de4978335a8d849741298b787503f6638f9c3b2b54a0a86d301531cd66787d

Observation 800f19c3-7416-43bf-bd88-f70be22a9ea8 · outbound

This paper cites It sup- ports a context length of up to 128,000 to- kens and is designed to handle complex tasks across multiple languages, including English and Chinese.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It sup- ports a context length of up to 128,000 to- kens and is designed to handle complex tasks across multiple languages, including English and Chinese

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.632240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.821689Z digest=sha256:695b4f23f08053d54fe69b38972c3fc22c4cc4211238587917cb55fc28270e1a

Observation 09094eab-8992-4983-8325-a1867a8cc242 · outbound

This paper cites It has been fine-tuned using reasoning data generated by DeepSeek-R1, resulting in enhanced performance in reasoning tasks.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It has been fine-tuned using reasoning data generated by DeepSeek-R1, resulting in enhanced performance in reasoning tasks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.623776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.854506Z digest=sha256:25b04a37622fd4c8f76d7ac79c1722ead735c7901b17c85b86cb393b6015d6d9

Observation 59fabfb4-b184-42fd-a180-dd4aef9fb392 · outbound

This paper cites It has been fine- tuned with reasoning data from DeepSeek-R1, achieving state-of-the-art results in various benchmarks.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It has been fine- tuned with reasoning data from DeepSeek-R1, achieving state-of-the-art results in various benchmarks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.615741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.875686Z digest=sha256:0113bb328960a72ccae17868872d102c5f0750c81b53ad0739024915c4c4b700

Observation 4b8f039f-1511-4e09-81c5-8e60edabc237 · outbound

This paper cites budget forcing.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean budget forcing

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.561205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.915996Z digest=sha256:90990618fedc1bfb8f0c171520cc58ccc5c2fd7989f986ebc33dbdf6e37599e5

Observation 9fe0ef9f-69e3-4443-8611-ccc82dc63ae1 · outbound

This paper cites It is designed for real-time interactions and has been integrated into various Google prod- ucts, including Bard and Pixel smartphones.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It is designed for real-time interactions and has been integrated into various Google prod- ucts, including Bard and Pixel smartphones

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.340272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:27.046501Z digest=sha256:265634f2d9cfe0c49dea6537d2ed3514c63871336926c309d800e7591c7332f3

Observation 38240ac7-0be4-4170-8aa5-efc88dfaaa20 · outbound

This paper cites It introduces fea- tures such as a Multimodal Live API for real- time audio and video interactions, enhanced spatial understanding, and integrated tool use, including Google Search.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It introduces fea- tures such as a Multimodal Live API for real- time audio and video interactions, enhanced spatial understanding, and integrated tool use, including Google Search

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.324463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:27.057472Z digest=sha256:4b9d8b6ea3327656b12fbe44cdc0e6c78249dce9f0a8dfe44f039fc3e040279b

Observation 1de16268-a40c-4b2c-a3c9-9076e88de300 · outbound

This paper cites It features a context window of up to 200,000 tokens, allowing it to process extensive text sequences effectively.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It features a context window of up to 200,000 tokens, allowing it to process extensive text sequences effectively

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.309300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:27.074131Z digest=sha256:173864a52d9a4941469e67ba8ba81a070902988347af26ac7075d7589f2629c6

Observation 3b4953cd-dcdb-45c9-b83b-c27fb5feed45 · outbound

This paper cites It maintains a large context window and has been fine-tuned for better alignment with hu- man preferences.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It maintains a large context window and has been fine-tuned for better alignment with hu- man preferences

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.296196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:27.081256Z digest=sha256:53167f1d3736e87c30240910630f54a2e3a17c8db7d86a7ee55135720a355dd2

Observation 0b144209-f3ef-40e4-9f9e-0e64933b5e0c · outbound

This paper cites It has been widely used in applications requiring natural language understanding and generation.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It has been widely used in applications requiring natural language understanding and generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.280688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:27.086333Z digest=sha256:0407f9023bd2960b5f3f43366eb36fa227ce7315883aa2fe9060a2382787e2ef

Observation 7dc489f5-39c2-4c42-ba18-830aaa88d5ef · outbound

This paper cites It offers rapid response times and has been integrated into various applications for real-time interactions.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It offers rapid response times and has been integrated into various applications for real-time interactions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.258217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:27.092680Z digest=sha256:b6babd176db7e68256cc9919982c36ee857126d7eb650452a572ae840882a195

Observation 48ef840a-2471-4799-a338-842f4de48ef2 · outbound

This paper cites It ex- hibits rapid response times comparable to hu- man reactions and has enhanced performance in non-English languages.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean It ex- hibits rapid response times comparable to hu- man reactions and has enhanced performance in non-English languages

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.243231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:27.096971Z digest=sha256:4e60d1248f800f65350f052416efbb0888a392a6be7bc03a7c30e1de97d737fe

Observation 13635027-7000-4c17-a852-4bcea4e4b444 · outbound

This paper cites think- ing.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean think- ing

Reference 30

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:52:27.226564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:27.102294Z digest=sha256:6e73a28d347b4edb33db8c9433a44417fd229bd7648bd5dbd0a9d358072be474

Observation 2e1985cb-ce9c-4add-8f49-8ba247f9b960 · outbound

This paper cites In Proceedings of the 58th Annual Meet- ing of the Association for Computational Linguistics , pages 4609–4622, Online.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean In Proceedings of the 58th Annual Meet- ing of the Association for Computational Linguistics , pages 4609–4622, Online

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.747529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.374728Z digest=sha256:b174b79fb8b417c0ecd6bc7c13da8f8d517e79558e743499e2e3e4ad93ef3bd5

Observation a52da19e-6583-4709-9fa5-683c456159db · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:26.406804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:26.406804Z digest=sha256:9cc34f1b9519cab17e9ff05bd11cc4cc652833a28af92778a79391c58d2b4a63

Observation b35ffa67-b76c-471a-ac2b-53c49fff9a1e · outbound

This paper cites In Findings of the As- sociation for Computational Linguistics: ACL 2024 , pages 12159–12173, Bangkok, Thailand.

Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean In Findings of the As- sociation for Computational Linguistics: ACL 2024 , pages 12159–12173, Bangkok, Thailand

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:27.739596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:26.419737Z digest=sha256:de9a554e79795615adffc2967d220f1b2a04e9b97bea7711c836d47fe739330f

Pith citing papers

No inbound Pith citation observations are available.