Pith. sign in

Paper Citation Record · LEDGER

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training

As of 19 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2412.02775.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02775 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:11:07.285672Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7c600d3a-bd23-43b9-9d08-2cdb6880a323 · outbound

This paper cites Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.182797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.182797Z digest=sha256:0f2df637248abc9b65077975f2e65d4b22c6e6198bdb7ad1043290edb671a229

Observation 969dd770-d33f-459b-b51c-751c26cf78f1 · outbound

This paper cites LLaMA Beyond English: An Empirical Study on Language Capability Transfer.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training LLaMA Beyond English: An Empirical Study on Language Capability Transfer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.188523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.188523Z digest=sha256:9dad86e64c5c6275b02b2e412642fed7f3143122d7cca78e8b22e8e699951b85

Observation 555a95fd-a86a-4e66-bfdc-43134ddf8ae0 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.193687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.193687Z digest=sha256:768b15b26b70d067bc6f3913c3382f1deb85017d5dbe9258c18f319447ae52f5

Observation 1bd3f824-e827-4b38-805e-4e22681fc76a · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.199483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.199483Z digest=sha256:b3b57f40d83dd9f155d9478599ffd000834575e18bd8b573f35a4645120ba452

Observation 4a9c4d81-3d81-45ea-b1ad-db1a22728878 · outbound

This paper cites Cosmopedia,.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Cosmopedia,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.204845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.204845Z digest=sha256:253ca89469abd499477952c663e0863b525fdac8d877eed2855b8d63f89c56f1

Observation 040194b0-2f73-40b0-abde-8aab779ecf15 · outbound

This paper cites Mistral 7B.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Mistral 7B

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.209519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.209519Z digest=sha256:a5c2a26b0807eda44e4ca1f31f4f3b3834f4bb51095e56182d0aa1bcfa850b8c

Observation cd2ad884-b6ff-4a2e-b055-b31ff11014d0 · outbound

This paper cites Orca: Progressive learning from complex explanation traces of gpt-4,.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Orca: Progressive learning from complex explanation traces of gpt-4,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.214883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.214883Z digest=sha256:00dfbf04021d7bb0141193e0fd22e5a537e7d894b7908dcbf092b5e477294832

Observation 9f38ca34-3f9f-490d-a92b-4c916dd6a136 · outbound

This paper cites Introducing cosmosGPT: Monolingual Training for Turkish Language Models.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Introducing cosmosGPT: Monolingual Training for Turkish Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:11:07.406036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:11:07.219600Z digest=sha256:7d7a4bec6064ad16b8f30fad78f46fd83eb5bdc80c6c52d5100f002053b6fe3c

Observation ad16f25f-7633-4136-8cb7-d48dec37d9d6 · outbound

This paper cites XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.224458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.224458Z digest=sha256:5bb4b660a83983a27d2e10c11297fa3855d7bf602d390f125febe1245da03e28

Observation c4dcce90-dc06-46b0-b83d-87a788ed7bf6 · outbound

This paper cites Few-shot learning with multilingual generative language models,.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Few-shot learning with multilingual generative language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:07.596464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:11:07.229238Z digest=sha256:0f9dee3e8260787331ed06f1aff7430fc7f32234877c0d77eaaba989a02878f5

Observation 964648b7-08ab-4299-9956-a51a9792f818 · outbound

This paper cites Llama 3 model card,.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Llama 3 model card,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.234045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.234045Z digest=sha256:e368a58568160aeb12a916e0d71c5f5cf0915434608737fece3f1757fd130e46

Observation bdf210dc-e9b0-4bfd-ba09-c396e245db50 · outbound

This paper cites Decoupled Weight Decay Regularization.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Decoupled Weight Decay Regularization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.238494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.238494Z digest=sha256:667a8c96afc91f1610c5852a1c1873a55d3b58fad32aa82ff804fb9f6ca82730

Observation 10f653ad-879e-42a6-b987-a12d412b65d9 · outbound

This paper cites Arcee's MergeKit: A Toolkit for Merging Large Language Models.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Arcee's MergeKit: A Toolkit for Merging Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.244151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.244151Z digest=sha256:54ddfed9a84819872051ae07b9f01c4e4b185e61b58e7b7edc68788cd63e82bb

Observation 83bf0cf7-ffb4-488a-8aba-3fc9837e3aa4 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.248678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.248678Z digest=sha256:c0e420834b5c3f8151d36caa60866432d764a8a64e3a58fe97a0185d81905109

Observation 9d5fe4d4-5350-4589-b74c-0f393762d95d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Training Verifiers to Solve Math Word Problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.253350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.253350Z digest=sha256:ab41f99bf06df52b3ac56b27980b9681d04b03813520a10414905edab264dbb7

Observation 6b5c4746-2c08-412c-8ab0-463712d873cf · outbound

This paper cites Measuring massive multitask language understanding,.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Measuring massive multitask language understanding,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.258163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.258163Z digest=sha256:42a39744b9a975e521e2be62eebb77f79553f78d94b35acb59e049407923ee36

Observation d263ef71-830b-47ae-88d9-61562e9dfe88 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods,.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Truthfulqa: Measuring how models mimic human falsehoods,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.262394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.262394Z digest=sha256:703fb147f3fe1ed3f8e1e6373307cd2faf25c541f38c0535bafe2a0aa20b8aa5

Observation f763531d-4cad-4a07-a097-3ead4e23d2f5 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale,.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Winogrande: An adversarial winograd schema challenge at scale,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.267519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.267519Z digest=sha256:fa1686d77d0134b1abc671b402406b42fff7e28ade99c77ff7656c3c524bec6f

Observation 8238fc11-b7fd-4985-b66a-636a10ccb8bb · outbound

This paper cites T ¨urkc ¸e dil mod- ellerinin performans kars ¸ılas ¸tırması performance comparison of turkish language models,.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training T ¨urkc ¸e dil mod- ellerinin performans kars ¸ılas ¸tırması performance comparison of turkish language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:07.542127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:11:07.271826Z digest=sha256:2bac497f279c96503429e5c1de1f0aaaa3a765bd7b96d0b47e6f14a8c79fbf35

Observation 54fbfdef-081b-4004-9995-bbefbe19edc7 · outbound

This paper cites Trendyol/trendyol-llm-7b-chat-v0.1.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Trendyol/trendyol-llm-7b-chat-v0.1

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:07.526525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:11:07.276689Z digest=sha256:5928ffba10f7dbc401fbe52af56628910b94cb035237d0c398d8c2d9ec2b3852

Observation 7848dcca-3cbc-45e6-a3dc-e57cdb9e0f0a · outbound

This paper cites Turkcell/turkcell-llm-7b-v1.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Turkcell/turkcell-llm-7b-v1

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:07.511117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:11:07.281299Z digest=sha256:fbc052d0b03fd5f04075bd0e30c2c54ff2c246d9bfc33c864aac24c2265a5bb2

Observation e221c10c-c261-414b-99cd-f47f6973a8a8 · outbound

This paper cites sambanovasystems/sambalingo-turkish-chat.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training sambanovasystems/sambalingo-turkish-chat

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:07.495637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:11:07.285672Z digest=sha256:2e2551566082448e8dfec56e41129e7ea8c48fbfa9dcc437f28e887678bd9dc2

Pith citing papers

No inbound Pith citation observations are available.