Pith. sign in

Paper Citation Record · LEDGER

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training

As of 19 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2412.02775.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02775 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:11:07.285672Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7c600d3a-bd23-43b9-9d08-2cdb6880a323 · outbound

This paper cites Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.182797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.182797Z digest=sha256:3922e4f4d9718e31cd60148aec18f28ce271055b2fbc4ed0f720a5bfb1b61ba5

Observation 969dd770-d33f-459b-b51c-751c26cf78f1 · outbound

This paper cites LLaMA Beyond English: An Empirical Study on Language Capability Transfer.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training LLaMA Beyond English: An Empirical Study on Language Capability Transfer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.188523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.188523Z digest=sha256:1d39078cc4b62b69999428b65ba06dfc27278a2e692bd4e7fe5f13bcc5d17782

Observation 555a95fd-a86a-4e66-bfdc-43134ddf8ae0 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.193687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.193687Z digest=sha256:4a8819c0473714a23eda37725f31b3f199175e2a12dc145f7137dc6962c819be

Observation 1bd3f824-e827-4b38-805e-4e22681fc76a · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.199483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.199483Z digest=sha256:9e3083e0a679c0b80eef7fe45e6fad2b0c4d9b2fffd500386724b67b993134b8

Observation 4a9c4d81-3d81-45ea-b1ad-db1a22728878 · outbound

This paper cites Cosmopedia,.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Cosmopedia,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.204845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.204845Z digest=sha256:258bb7daa7e4ed633cdc01f0b1acf7a9c3894d06e642eb15e3a288811984da2a

Observation 040194b0-2f73-40b0-abde-8aab779ecf15 · outbound

This paper cites Mistral 7B.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Mistral 7B

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.209519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.209519Z digest=sha256:515a4a679737d2c4b7c7502aad02bc8a78457bc2bc258a280b43f90c8430966f

Observation cd2ad884-b6ff-4a2e-b055-b31ff11014d0 · outbound

This paper cites Orca: Progressive learning from complex explanation traces of gpt-4,.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Orca: Progressive learning from complex explanation traces of gpt-4,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.214883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.214883Z digest=sha256:81313d4ef9a55741ac2db7943285dcb733171e7250608737c37a0a94b89a876b

Observation 9f38ca34-3f9f-490d-a92b-4c916dd6a136 · outbound

This paper cites Introducing cosmosGPT: Monolingual Training for Turkish Language Models.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Introducing cosmosGPT: Monolingual Training for Turkish Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:11:07.406036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:11:07.219600Z digest=sha256:ba3134cb306987afbf68ef10b263aaf3995f8809ef9a903ebe93d90372de9c27

Observation ad16f25f-7633-4136-8cb7-d48dec37d9d6 · outbound

This paper cites XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.224458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.224458Z digest=sha256:621ac44317b7558798053fb38c83955fd42759178804d04f0a6209486546a5ed

Observation c4dcce90-dc06-46b0-b83d-87a788ed7bf6 · outbound

This paper cites Few-shot learning with multilingual generative language models,.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Few-shot learning with multilingual generative language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:07.596464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:11:07.229238Z digest=sha256:e6a1de2166fdb41bc605aa5dd774aaa4868752b0401acd79b86ecccac50c53ae

Observation 964648b7-08ab-4299-9956-a51a9792f818 · outbound

This paper cites Llama 3 model card,.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Llama 3 model card,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.234045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.234045Z digest=sha256:d94aa2002f55cdec88904dc727523bdfc566653e2149e9a4aaae427dbb0d2f26

Observation bdf210dc-e9b0-4bfd-ba09-c396e245db50 · outbound

This paper cites Decoupled Weight Decay Regularization.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Decoupled Weight Decay Regularization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.238494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.238494Z digest=sha256:401d6cba054812f5ae441bb59e45ae8b9d9bf0cb8c68cda7fac45c97adc829a7

Observation 10f653ad-879e-42a6-b987-a12d412b65d9 · outbound

This paper cites Arcee's MergeKit: A Toolkit for Merging Large Language Models.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Arcee's MergeKit: A Toolkit for Merging Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.244151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.244151Z digest=sha256:fa5cafdfe2431763f9787c09eba22ce0c4386e2d21f1d07a4ce10194875fd5dd

Observation 83bf0cf7-ffb4-488a-8aba-3fc9837e3aa4 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.248678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.248678Z digest=sha256:75b8a97af404a8a5596b791d7fcb3f010b701bf4f4b262efca6503b11ed18ecb

Observation 9d5fe4d4-5350-4589-b74c-0f393762d95d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Training Verifiers to Solve Math Word Problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.253350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.253350Z digest=sha256:9bdabc95cc7f330a2dd0e9e4d3c4ece8bf4f437f2f93521cbd229a3bbeabc468

Observation 6b5c4746-2c08-412c-8ab0-463712d873cf · outbound

This paper cites Measuring massive multitask language understanding,.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Measuring massive multitask language understanding,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.258163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.258163Z digest=sha256:dad2a8357e2fa8ba113e17a6b8ec341943560911d35d144dd6ccedce8b23ddaa

Observation d263ef71-830b-47ae-88d9-61562e9dfe88 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods,.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Truthfulqa: Measuring how models mimic human falsehoods,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.262394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.262394Z digest=sha256:119da58958608e1c01d4b18e711d5a6bc5d9104aecb0fc1f3587407921fafba0

Observation f763531d-4cad-4a07-a097-3ead4e23d2f5 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale,.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Winogrande: An adversarial winograd schema challenge at scale,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:07.267519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:07.267519Z digest=sha256:685225797a976d43935ddbde393222617965db9de8198f4b21776b6fb662d20f

Observation 8238fc11-b7fd-4985-b66a-636a10ccb8bb · outbound

This paper cites T ¨urkc ¸e dil mod- ellerinin performans kars ¸ılas ¸tırması performance comparison of turkish language models,.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training T ¨urkc ¸e dil mod- ellerinin performans kars ¸ılas ¸tırması performance comparison of turkish language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:07.542127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:11:07.271826Z digest=sha256:27bf24d7ec9e82b2c77abfc867c1eb9db2743c29a63b04937dc2d784cef169ce

Observation 54fbfdef-081b-4004-9995-bbefbe19edc7 · outbound

This paper cites Trendyol/trendyol-llm-7b-chat-v0.1.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Trendyol/trendyol-llm-7b-chat-v0.1

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:07.526525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:11:07.276689Z digest=sha256:2487b9faa62d75b5bff0ff8ecd0ef4d89e12658dd759250e5c8d7bcb1ff38842

Observation 7848dcca-3cbc-45e6-a3dc-e57cdb9e0f0a · outbound

This paper cites Turkcell/turkcell-llm-7b-v1.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training Turkcell/turkcell-llm-7b-v1

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:07.511117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:11:07.281299Z digest=sha256:04a5518a8b78dd3f0f4190824fea2b10a6b7abf74f30aadc26af7b882bb25f9b

Observation e221c10c-c261-414b-99cd-f47f6973a8a8 · outbound

This paper cites sambanovasystems/sambalingo-turkish-chat.

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training sambanovasystems/sambalingo-turkish-chat

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:07.495637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:11:07.285672Z digest=sha256:d8f74408335c7b958cdc7d30d590a278869524c58c5b8bd0bc1c092380c36d71

Pith citing papers

No inbound Pith citation observations are available.