Pith. sign in

REVIEW 4 major objections 4 minor 43 references

IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems

T0 review · 4 major / 4 minor · reviewed 2026-07-31 · deepseek-v4-flash

Pith's one-line read A new 24,795-example multilingual dataset can teach a compact 7B language model to deliver Indian Knowledge Systems pedagogy nearly on par with a strong general-purpose model, while the unfine-tuned base model scores near zero on the domain

desk verdict The dataset is a genuine new resource; the 'fine-tuning is necessary' claim is weaker than the prose, because the base-model number comes from a different, smaller evaluation. read the letter →

arxiv 2607.23322 v1 pith:UXZIFRB3 submitted 2026-07-25 cs.CL cs.CYcs.ETcs.LG

classification cs.CLcs.CYcs.ETcs.LG
keywords IndianKnowledgeSystemsinstructiontuningmultilingualdatasetpedagogicalAICBSEcurriculumalignmentVedicmathematicsmulti-judgeevaluationtechniquefidelity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that Indian Knowledge Systems (IKS) pedagogy can be taught to language models through a purpose-built instruction dataset, and that such fine-tuning is necessary for the capability. It introduces IKS-Instruct, 24,795 instruction-response pairs spanning seven Indian languages, 41 pedagogical techniques, and CBSE classes 6-12, built from classical texts, Vedic mathematics sutras, curriculum templates, bilingual pairs, dialogues, and cross-tradition comparisons. Under a fixed panel of five external language-model judges scoring 1,201 stratified items on 12 dimensions, the best 7B fine-tune reaches a median score of 6.39, within 0.15 of the 6.54 reference model, while the base model scores near zero on IKS-specific axes. The authors also report that model quality does not rise monotonically with dataset size or curation: the moderate 13,882-pair version produced the strongest model, and both adding noisy pairs and aggressive pruning lowered judge scores. A sympathetic reader would care because it suggests inexpensive, locally deployable 7B models could deliver curriculum-aligned IKS instruction in multiple languages if given the right training data.

What carries the argument

The load-bearing artifact is the IKS-Instruct dataset itself, with its structured metadata (source provenance, technique, language, subject, class level) and its six source categories. Inside the construction pipeline, the 'mechanism-first' generation approach is the key device: instead of generating a response to an instruction, it builds the technique demonstration first and reverse-engineers an instruction that elicits it, which guarantees named techniques are actually shown. Quality is carried by a two-stage evaluation: a Teacher Fidelity filter (a 7B model scoring technique correctness, with human review in the ambiguous band) followed by a 12-dimension multi-judge framework using a uni

What would settle it

Have a panel of human IKS educators score the same 1,201 items on the same 12 dimensions and compare their rankings with the five-judge panel. If the human experts do not place the fine-tuned model close to the reference and far above the base model, the paper's central claims of necessity and near-parity are not supported.

Watch

Extended reading notes

Core claim

The paper's central claim is that IKS-Instruct fine-tuning is necessary for IKS pedagogical competence in compact open models, and sufficient to bring them close to a much stronger general-purpose model. Specifically, a LoRA fine-tune of a 7B Llama-based model on the 13,882-pair v1.1 split reaches a median judge score of 6.39 under a uniform five-judge external panel, compared with 6.54 for the reference model Nemotron-Nano and near-zero scores for the unfine-tuned base on Technique Fidelity, IKS Cultural Depth, and IKS Pedagogical Authenticity. The paper also claims a non-monotonic quality-volume relationship: adding roughly 6,000 noisy pairs (v1.7) slightly reduced scores to 6.33, and aggr

Load-bearing premise

The central comparison rests on the assumption that the five AI judges' 0-10 scores measure real teaching quality rather than reward longer, formally written, or family-similar responses; if that assumption fails, the near-parity claim and the base model's near-zero scores lose their grounding.

Editorial extensions

If this is right

  • Compact 7B models fine-tuned on IKS-Instruct can deliver IKS-grounded educational content across English, Hindi, Sanskrit, Tamil, Telugu, Kannada, and Malayalam at near-reference quality, making local, low-cost classroom deployment realistic.
  • Fine-tuning on IKS-Instruct is necessary for IKS capability: the base model scores near zero on technique fidelity, cultural depth, and pedagogical authenticity, so general-purpose instruction data alone does not confer this competence.
  • Bilingual pairs transfer technique knowledge across languages: models fine-tuned on seven languages can demonstrate Vedic math sutras in unseen languages such as Marathi and Bengali at scores far above the base model, pointing toward coverage of all 22 scheduled Indian languages.
  • Source-anchored pairs reduce hallucination and raise Teacher Fidelity, implying that grounding instruction data in citable classical passages is a concrete quality lever.
  • Dataset quality is not monotonic in size or curation: moderate, coherent versions can outperform both larger noisy versions and aggressively pruned versions, so dataset builders should benchmark against balanced baselines.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the 0.15-point gap to the reference model is only meaningful if the judge panel's scores track genuine pedagogical quality; a human-expert calibration study on the same 1,201 items would test whether the near-parity survives.
  • Editorial inference: the non-monotonic result implies that for domains with expensive annotation, the practical optimum may be a moderately sized, source-anchored dataset rather than maximally curated or maximally scaled one, since the curated v2.1 trains at lower cost with judge scores within the reported noise band.
  • Editorial inference: the technique-schema-transfer finding suggests a cheap expansion path: generate a small number of high-fidelity bilingual exemplars per technique for each new language and test whether technique fidelity transfers, which could extend IKS instruction to the remaining scheduled languages without full-scale dataset construction.
  • Editorial inference: the 12-dimension multi-judge protocol could serve as a template for other culturally grounded pedagogical domains, but its absolute scores should not be compared across different judge panels unless calibrated.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces IKS-Instruct, a 24,795-pair multilingual instruction dataset for Indian Knowledge Systems (IKS) pedagogy, spanning seven languages, 41 techniques, and CBSE curriculum alignment across classes 6–12. It describes a multi-stage construction pipeline (hybrid teacher generation, mechanism-first generation, task-abstraction generation) with quality filtering based on Teacher Fidelity, provenance checks, and multi-judge evaluation. The headline empirical claim is that fine-tuning on IKS-Instruct is necessary for IKS pedagogical capability: under a uniform five-judge external panel, a compact 7B model reaches a median score of 6.39, within 0.15 of a general-purpose reference model (Nemotron-Nano, 6.54), while the base model scores near zero on IKS-specific dimensions. The paper also reports a non-monotonic quality–volume relationship across dataset versions and releases the data, evaluation artifacts, and re-aggregation script.

Significance. If the headline claims hold, IKS-Instruct is a significant resource: it is the first instruction dataset targeting IKS pedagogy across multiple Indian languages with explicit CBSE alignment, and the construction methodology—especially the mechanism-first generation and the Teacher Fidelity filter—is a useful template for domain-specific instruction data. The authors are commendably transparent about several limitations: they disclose the residual Llama-lineage circularity in the judge panel, the exclusion of Nemotron-family judges, the incomparability of the base-model pilot, and the small version-to-version score differences. However, because the central 'necessity' claim depends on a near-zero base-model result from a different protocol, and because all comparative claims rely on uncalibrated LLM-judge scores without human anchoring, the evidence is not yet load-bearing. Strengths include the release of the v1.1 split, per-judge evaluation artifact, and re-aggregation script, which would enable independent verification.

major comments (4)
  1. [§5.3, Table 6] The paper's central claim that IKS-Instruct fine-tuning is 'necessary' rests on the base model scoring near zero, but Table 6 explicitly states that base-model scores come from an earlier small-scale pilot (N=146) and the base was not re-evaluated under the uniform five-judge panel. No per-dimension scores, judge composition, confidence intervals, or representativeness analysis of the pilot sample are provided. Section 8.3 also concedes that LLM judges reward longer, more formal responses, and the paper describes the base model's outputs as 'grammatically fluent and topically related to India.' The near-zero result may therefore be a protocol artifact. The base model must be evaluated under the identical panel on the same 1,201 stratified items, with per-dimension scores reported, before the necessity claim can be supported.
  2. [§5.1, §8.3] The entire comparative evaluation relies on LLM judges' 0–10 scores across 12 dimensions, but there is no human expert calibration or inter-judge reliability analysis. The paper itself hedges in §8.3 that 'absolute scores should be read as comparative under a fixed panel rather than as calibrated quality ratings,' and §8.2 treats version differences of 0.1–0.3 median points as 'indicative rather than definitive.' Yet the abstract and conclusion treat the 0.15-point gap to Nemotron-Nano and the near-zero base result as definitive. The authors should provide per-item variance, confidence intervals, a human-annotated calibration subset, and a sensitivity analysis showing that the conclusions hold under judge subsets without Llama-family models.
  3. [Conclusion (§9)] The conclusion claims that 'the improvement is observed across all evaluation dimensions' and 'the consistency of this pattern across all seven languages' confirms genuine multilingual IKS capability. However, no per-dimension or per-language result tables are reported anywhere in the paper; Table 6 gives only median aggregate scores. Without these breakdowns, the claims are unsupported. Please add the per-dimension and per-language tables, or explicitly restrict the claims to what is actually measured.
  4. [§8.3] Residual self-preference is acknowledged but not addressed quantitatively: the fine-tuned models are LoRA adapters over the Llama-based Airavata backbone, while two of the five judges (Llama-3.1-8B, Llama-3.3-70B) are Llama-family. This could systematically inflate fine-tuned model scores relative to the Nemotron reference and especially relative to the base model. The paper should report the comparison restricted to non-Llama judges (e.g., Qwen2.5-7B, Dracarys-70B, GPT-4o-mini only) or otherwise demonstrate that the headline gaps persist after removing Llama-lineage judges.
minor comments (4)
  1. [References] Several references are incomplete or inconsistent: e.g., [10] gives only 'Lee U. LLaV A-Docent-V2', [16] lists only three authors for an ACL anthology entry, and some entries lack page numbers or publisher details. Please normalize all references to a consistent format.
  2. [§6.2] The claim that bilingual pairs transfer technique knowledge across languages is supported only by preliminary TF-score differences (ΔTF = +3.2 vs +1.8) without sample sizes or significance tests. This is presented as a 'finding' in the conclusion; it would be better framed as a hypothesis pending a controlled cross-lingual transfer experiment.
  3. [§3.2, §5.2] The Teacher Fidelity filter threshold (TF ≥ 6) and the multi-judge median threshold (≥ 7.0 for v2.1 curation) are described as ad hoc but no sensitivity analysis is reported. A short analysis showing how model-level scores vary with these thresholds would strengthen the methodology.
  4. [General] The abstract says 'Model quality does not increase monotonically with data curation,' but Section 8.2 explicitly cautions that the version ranking is 'indicative rather than definitive.' The abstract and conclusion should use the same cautious language.

Circularity Check

1 steps flagged · score 4.0 of 10

Residual Llama-lineage judge self-preference is admitted; base-model necessity claim rests on a different pilot protocol.

  1. other [Section 8.3 (Limitations), discussing the model comparison of Table 6]
    "A residual circularity remains: the panel includes Llama-family judges (Llama-3.1-8B, Llama-3.3-70B), while the fine-tuned models are LoRA adapters over the Llama-based Airavata backbone, so some self-preference toward Llama-lineage outputs cannot be excluded."

    The paper's central quantitative claims (fine-tuned 7B at 6.39 vs. Nemotron-Nano at 6.54; base near zero) are produced by a five-judge panel, two of whose members share the Llama lineage of the LoRA-fine-tuned models. Part of the measured score can therefore reflect judge self-similarity to the model family rather than independent IKS pedagogical capability. The paper explicitly labels this a residual circularity, so the evaluation loop is admitted rather than merely suspected.

full rationale

The core derivation is not, by construction, circular: the dataset is assembled from six source categories, generated with an external Nemotron teacher, filtered by a Teacher-Fidelity model calibrated on 500 expert triples, and evaluated under a uniform five-judge panel that excludes Nemotron-family judges. The version-to-version comparisons (v1.1 6.39, v1.7 6.33, v2.1 6.12) are made under a common protocol and have independent empirical content. The main weakness is the 'necessity' claim: Table 6 explicitly states that the base-model near-zero scores come from an earlier N=146 pilot and that the base was not re-evaluated under the uniform panel, so the base-vs-fine-tuned comparison is an evidential gap rather than a same-protocol measurement. The one genuine circularity is the one the paper itself concedes in Section 8.3: two Llama-family judges evaluate Llama-LoRA models, leaving residual judge self-preference. This is a partial evaluation loop, not a full definitional collapse, because the dataset construction and the reference-model comparison retain independent substance. Score 4 reflects that admitted partial circularity while recognizing the independent content of the dataset and version comparisons.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The central claims rest on the integrity of the constructed dataset and on the validity of LLM-judge measurements. The paper introduces no theoretical entities, but relies on hand-set quality thresholds and a domain taxonomy, and assumes judge scores measure teaching quality; the judge-validity assumption is the most load-bearing.

free parameters (2)
  • Teacher Fidelity (TF) inclusion threshold = 6 (on 0-10 scale)
    Section 5.2: pairs with TF >= 6 enter the training split; TF 4-6 need human review; <4 are excluded. This hand-set cutoff controls dataset composition and therefore model outcomes.
  • Curated-pair judge median threshold = median > 7.0
    Section 6.1: v2.1 retains pairs with multi-judge median above 7.0. This hand-set cutoff shapes the curated dataset.
assumptions (3)
  • domain assumption LLM judge scores on the 12 dimensions are a valid proxy for educational and technical quality.
    Section 5.1/8.3: no human-expert calibration; judges are known to favor length, formal style, and stylistic similarity; residual Llama self-preference admitted.
  • domain assumption The 41 IKS techniques and their pedagogical steps are represented accurately in the retained pairs.
    Section 3.3/5.2: technique verification rests on an automated 7B TF model with r=0.82 on 100 human examples; auto-generated Vedic math had 67% error before filtering, so subtle errors may persist.
  • domain assumption CBSE curriculum outcome codes and subject mappings are correct.
    Section 7.2: mappings are compiled from official CBSE documents by curators; no independent audit is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems." pith.science (2026). https://pith.science/paper/UXZIFRB3

@misc{pith2026260723322,
  author       = {Pith},
  title        = {Pith review of: IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UXZIFRB3}},
  note         = {Machine review of arXiv:2607.23322}
}
read the original abstract

Instruction tuning has become the standard method for adapting large language models to follow human intent, yet existing instruction datasets are dominated by English-language general-knowledge tasks and lack coverage of specialized pedagogical domains. This paper presents IKS-Instruct, a dataset of 24,795 instruction-response pairs for teaching language models to deliver educational content grounded in Indian Knowledge Systems (IKS). The dataset spans seven languages (English, Hindi, Sanskrit, Tamil, Telugu, Kannada, and Malayalam), covers 41 pedagogical techniques from the Vedic oral and mathematical traditions, and is aligned with the Central Board of Secondary Education (CBSE) curriculum for classes 6 through 12. The pairs are derived from six source types: classical text corpora (Bhagavad Gita, Thirukkural, Sangam literature, Vedic texts), curriculum-aligned pedagogical templates, Vedic mathematical sutra demonstrations, bilingual instruction pairs, technique-grounded multi-turn dialogues, and cross-tradition comparative analyses. Quality is assessed through a multi-judge evaluation framework in which independent language models score responses on 12 dimensions including technique fidelity, pedagogical quality, factual accuracy, and IKS cultural depth. Under a uniform five-judge external panel (median aggregation over 1,201 stratified items), the strongest IKS-Instruct fine-tune of a compact 7B model reaches a median judge score of 6.39, within 0.15 of a strong general-purpose reference model (Nemotron-Nano at 6.54) at a fraction of its deployment cost, while the base model without IKS fine-tuning scores near zero on the IKS-specific dimensions. Model quality does not increase monotonically with data curation, a result we report together with the corresponding data-quality gains.

Figures

Figures reproduced from arXiv: 2607.23322 by the authors.

Figure 1
Figure 1. Source category distribution in IKS-Instruct v1.9. Classical text corpus pairs constitute the largest [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Top-level overview of the IKS-Instruct quality filtering pipeline. Generated pairs from the six [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Language distribution of the 24,795 instruction pairs in IKS-Instruct v1.9. English (36.1%) and [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Model quality (median of the uniform five-judge external panel, 12 dimensions, N=1,201) for [PITH_FULL_IMAGE:figures/full_fig_p018_4.png]
Figure 5
Figure 5. Figure 5: CBSE subject by class level distribution in IKS-Instruct (heatmap). Darker cells indicate higher [PITH_FULL_IMAGE:figures/full_fig_p021_5.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

43 extracted references · 3 linked inside Pith

  1. [1]

    Dai, and Quoc V

    Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V . Le. Finetuned language models are zero-shot learners. InInternational Conference on Learning Representations, 2022

  2. [2]

    Smith, Daniel Khashabi, and Hannaneh Hajishirzi

    Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. Self-instruct: Aligning language models with self-generated instructions. In Proceedings of the 61st Annual Meeting of the ACL, pages 13484–13508. Association for Computational Linguistics, 2023

  3. [3]

    Hashimoto

    Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following LLaMA model. https: //github.com/tatsu-lab/stanford_alpaca, 2023. Accessed: 2024-11-15

  4. [4]

    Le, Barret Zoph, Jason Wei, and Adam Roberts

    Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V . Le, Barret Zoph, Jason Wei, and Adam Roberts. The Flan collection: Designing data and methods for 27 effective instruction tuning.Proceedings of the 40th International Conference on Machine Learning, pages 22631–22648, 2023

  5. [5]

    Free dolly: Introducing the world’s first truly open instruction-tuned LLM

    Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. Free dolly: Introducing the world’s first truly open instruction-tuned LLM. https://www.databricks.com/blog/2023/04/12/ dolly-first-open-commercially-viable-instruction-tuned-llm , 2023. Accessed: 2024- 11-15

  6. [6]

    Bactrian-x: Multilingual replicable instruction-following models with low-rank adaptation.arXiv preprint arXiv:2305.15011, 2024

    Haonan Li, Fajri Koto, Minghao Wu, Alham Fikri Aji, and Timothy Baldwin. Bactrian-x: Multilingual replicable instruction-following models with low-rank adaptation.arXiv preprint arXiv:2305.15011, 2024

  7. [7]

    National education policy 2020

    Ministry of Education, Government of India. National education policy 2020. https://www. education.gov.in/sites/upload_files/mhrd/files/NEP_Final_English_0.pdf, 2020. Ac- cessed: 2024-12-15

  8. [8]

    Indian Knowledge Systems in the Curriculum of Higher Education: A Proposed Model of a PG Course in IKS.International Journal of Research Publication and Reviews, 2023

    Poonam Thapliyal. Indian Knowledge Systems in the Curriculum of Higher Education: A Proposed Model of a PG Course in IKS.International Journal of Research Publication and Reviews, 2023

Show all 43 references
  1. [9]

    Kiranmayee Kar. Integrating Indian Knowledge Systems (IKS) into Teacher Education Curriculum: Framework Design, Challenges, and Strategic Interventions in the Context Of NEP.International Journal of Creative and Open Research in Engineering and Management, 2026

  2. [10]

    LLaV A-Docent-V2: Improving Data Quality and Pedagogical Data Generation to Train Large Multimodal Models for Art Appreciation Education.Lecture Notes in Computer Science, 2026

    Lee U. LLaV A-Docent-V2: Improving Data Quality and Pedagogical Data Generation to Train Large Multimodal Models for Art Appreciation Education.Lecture Notes in Computer Science, 2026

  3. [11]

    Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Quality

    Yuto Harada, Yusuke Yamauchi, Yusuke Oda, Yohei Oseki, Yusuke Miyao, and Yu Takagi. Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Quality. InProceedings of the 2025 Conference on Empirical Methods in Natural Languag...

  4. [12]

    Efficient Dataset Construction for Llm Fine-Tuning: A Text-to-Table Generation.SSRN Electronic Journal, 2025

    Sekjin Hwang. Efficient Dataset Construction for Llm Fine-Tuning: A Text-to-Table Generation.SSRN Electronic Journal, 2025. 28

  5. [13]

    Crosslingual generalization through multitask finetuning

    Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al. Crosslingual generalization through multitask finetuning. InProceedings of the 61st Annual Meeting of the ACL, ...

  6. [14]

    Zero-shot cross-lingual transfer in instruction tuning of large language models

    Nadezhda Chirkova. Zero-shot cross-lingual transfer in instruction tuning of large language models. Proceedings of the 17th International Natural Language Generation Conference, 2024

  7. [15]

    mCoT: Multilingual Instruction Tuning for Reasoning Consistency in Language Models

    Huiyuan Lai. mCoT: Multilingual Instruction Tuning for Reasoning Consistency in Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024

  8. [16]

    Indic-TunedLens: Interpreting Multilingual Models in Indian Languages.Proceedings of the 13th Workshop on NLP for Similar Languages, Varieties and Dialects, 2026

    Mihir Panchal, Deeksha Varshney, Mamta, and Asif Ekbal. Indic-TunedLens: Interpreting Multilingual Models in Indian Languages.Proceedings of the 13th Workshop on NLP for Similar Languages, Varieties and Dialects, 2026

  9. [17]

    NE-BERT: A Multilingual Language Model for 9 Northeast Indian Languages.Pro- ceedings of the Second Workshop on Language Models for Low-Resource Languages (LoResLM 2026), 2026

    Badal Nyalang. NE-BERT: A Multilingual Language Model for 9 Northeast Indian Languages.Pro- ceedings of the Second Workshop on Language Models for Low-Resource Languages (LoResLM 2026), 2026

  10. [18]

    Pallapu Pallapu. A Weighted Ensemble of FLAN-T5, MBART and NLLB-200 for Robust Multilingual NLP.Proceedings of the 1st International Conference on Intelligent Methods and Advanced Computer Scientific Innovations, 2025

  11. [19]

    IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian Languages

    Saiful Haq. IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian Languages. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2024

  12. [20]

    Multilingual Tokenization Efficiency in Large Language Models: A Study on Indian Languages.SSRN Electronic Journal, 2025

    Azhar Mohamed. Multilingual Tokenization Efficiency in Large Language Models: A Study on Indian Languages.SSRN Electronic Journal, 2025

  13. [21]

    Divya V Sharma. IndicSynth: A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages.Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025. 29

  14. [22]

    Murthy J.S. A Cross-Cultural Multilingual Text-to-Text Generation Framework Using Novel Sanskriti- GenAI Model for Preserving and Digitizing Cultural Heritage.Lecture Notes in Networks and Systems, 2026

  15. [23]

    Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al

    Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023. doi: 10.1038/s41586-023-06291-2

  16. [24]

    ChatLaw: Open-source legal large language model with integrated external knowledge bases.arXiv preprint arXiv:2306.16092, 2023

    Jiaxi Cui, Zongjian Li, Yang Yan, Bohua Chen, and Li Yuan. ChatLaw: Open-source legal large language model with integrated external knowledge bases.arXiv preprint arXiv:2306.16092, 2023

  17. [25]

    BloombergGPT: A large language model for finance

    Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann. BloombergGPT: A large language model for finance. arXiv preprint arXiv:2303.17564, 2023

  18. [26]

    Jaesung Hwang. Diabetes-Specialized Large Language Models for Clinical Reasoning and Dietary Recommendation: Reflection- and Curriculum-Based Instruction Tuning Study (Preprint).JMIR Preprints, 2026

  19. [27]

    TIFD: Tibetan Instruction-Following Dataset for Large Language Models Supervised Fine-Tuning.Data Intelligence, 2025

    Wenhao Zhuang. TIFD: Tibetan Instruction-Following Dataset for Large Language Models Supervised Fine-Tuning.Data Intelligence, 2025

  20. [28]

    The BEA 2023 shared task on generating AI teacher responses in educational dialogues

    Anaïs Tack and Chris Piech. The BEA 2023 shared task on generating AI teacher responses in educational dialogues. InProceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications, pages 785–795. Association for Computational Linguistics, 2023

  21. [29]

    Jakub Macina. MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems.Findings of the Association for Computational Linguistics: EMNLP 2023, 2023

  22. [30]

    Leveraging Large Language Models for Adaptive Tutoring System via Pedagogical Knowledge-Augmented Prompting.SSRN Electronic Journal, 2026

    Haein Jeon. Leveraging Large Language Models for Adaptive Tutoring System via Pedagogical Knowledge-Augmented Prompting.SSRN Electronic Journal, 2026

  23. [31]

    Using LLMs to Identify Indicators of Persistence from Students’ Dialogues with a Pedagogi- cal Agent.Journal of Educational Data Mining, 2026

    Ober T.M. Using LLMs to Identify Indicators of Persistence from Students’ Dialogues with a Pedagogi- cal Agent.Journal of Educational Data Mining, 2026. 30

  24. [32]

    Rooted in Heritage Collective. Reclaiming Indigenous Pedagogy: Integrating Indian Knowledge Systems (IKS) into Curriculum Frameworks Under NEP 2020.Rooted in Heritage: Integrating Education with Indian Knowledge Systems and Community Support, 2025

  25. [33]

    Panchali Moitra. Stakeholder Perceptions of Integration of Indian Indigenous Knowledge Systems in Mainstream Higher Education Curriculum: An Exploratory Qualitative Study.SSRN Electronic Journal, 2024

  26. [34]

    Fathima Shajeena

    C.P. Fathima Shajeena. Integrating Indian Knowledge System and Self-Regulated Learning: A Lin- guistic Exploration in Malayalam Medium Curriculum Resources.Research Review Journal of Indian Knowledge Systems, 2025

  27. [35]

    Reconstructing Curriculum Through Indian Knowledge Systems: A Policy Review in Light of NEP 2020.Mindcraft: Reimagining Education Through Indian Knowledge Systems, 2026

    Krishnendu Roy. Reconstructing Curriculum Through Indian Knowledge Systems: A Policy Review in Light of NEP 2020.Mindcraft: Reimagining Education Through Indian Knowledge Systems, 2026

  28. [36]

    Indian Knowledge Systems.Knowledge, Curriculum and Learning, 2025

    Mythili Ramchand. Indian Knowledge Systems.Knowledge, Curriculum and Learning, 2025

  29. [37]

    Integrating Vedic Mathematics with Large Language Model: A Novel Approach to Computa- tional Efficiency and Educational Innovation.Lecture Notes in Networks and Systems, 2026

    Naik N. Integrating Vedic Mathematics with Large Language Model: A Novel Approach to Computa- tional Efficiency and Educational Innovation.Lecture Notes in Networks and Systems, 2026

  30. [38]

    Curriculum for class nine CBSE poetry teaching with the integration of REBT based on NEP 2020.Journal of Poetry Therapy, 2025

    Jiny Kuriakose. Curriculum for class nine CBSE poetry teaching with the integration of REBT based on NEP 2020.Journal of Poetry Therapy, 2025

  31. [39]

    Language Diversity: Perceptions of Students towards Multilingual Instruction in Higher Education Pedagogy under NEP 2020.Indian Journal of Language and Linguistics, 2024

    Mohit Saini. Language Diversity: Perceptions of Students towards Multilingual Instruction in Higher Education Pedagogy under NEP 2020.Indian Journal of Language and Linguistics, 2024

  32. [40]

    Construction and Pedagogical Validation of an AI-Empowered Multidimensional Assessment System for English Speech Course.Journal of Educational Research and Policies, 2025

    Hongfang Xiao. Construction and Pedagogical Validation of an AI-Empowered Multidimensional Assessment System for English Speech Course.Journal of Educational Research and Policies, 2025

  33. [41]

    Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena. InAdvances in Neural Information Processi...

  34. [42]

    31 Retrieval-augmented generation for knowledge-intensive NLP tasks

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 31 Retrieval-augmented generation for knowledge-intensive NLP tasks. InAdvances in Neur...

  35. [43]

    Curriculum learning

    Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. Curriculum learning. In Proceedings of the 26th International Conference on Machine Learning (ICML), pages 41–48, 2009. 32

Pith tools

Reviewed July 31, 2026 · model on record in the stance chip above.