REVIEW 46 cited by
Challenges and Applications of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) went from non-existent to ubiquitous in the machine learning discourse within a few years. Due to the fast pace of the field, it is difficult to identify the remaining challenges and already fruitful application areas. In this paper, we aim to establish a systematic set of open problems and application successes so that ML researchers can comprehend the field's current state more quickly and become productive.
Forward citations
Cited by 46 Pith papers
-
CXXCrafter: An LLM-Based Agent for Automated C/C++ Open Source Software Building
An LLM-driven agent, CXXCrafter, automatically builds 587 of 752 C/C++ open-source projects (78%), beating default build commands (39%) and bare LLMs (32 to 38%).
-
Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
Jailbreak attacks push LLM activations outside a safety boundary, mostly in low and middle layers, and a tanh-based penalty that pulls activations back inside this boundary blocks most tested attacks with under 2% uti...
-
FastTPS: An Optimized Method for LLM Token Phase for AI accelerators
FastTPS accelerates LLM token-phase inference via reloading-free static KV-cache management, tiled fused RoPE attention, and interlaced-weight MLP fusion, yielding up to 6× speedup at 93% bandwidth on AMD NPUs.
-
Interpreting learning dynamics of autoencoders: Transient scaling and emerging concepts of the Ising model
Unsupervised autoencoders on Ising configurations form magnetization then energy representations in two dynamical regimes, with recursive error flow fields sharing topology across layers.
-
Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions
Across 74 OSINT/CTI AI studies, hallucination is widely named but end-to-end measured in only one non-reproducible system, so a human–AI co-pilot is the most defensible near-term architecture.
-
Real Faults in Model Context Protocol (MCP) Software: a Comprehensive Taxonomy
MCP server faults form five empirical categories—server setting, server/tool configuration, server/host configuration, documentation, and general programming—confirmed by a 41-practitioner survey.
-
Enhancing Robustness of Autoregressive Language Models against Orthographic Attacks via Pixel-based Approach
A word-as-image pixel language model trained with next-token prediction reports lower perplexity than a token-embedding LLaMA on noisy and non-Latin-script text, though its noise evaluation holds tokenization fixed.
-
PhantomHunter: Detecting Unseen Privately-Tuned LLM-Generated Text via Family-Aware Learning
PhantomHunter detects text from privately fine-tuned LLMs by learning shared token-probability traits within LLaMA, Gemma and Mistral families, reporting F1 above 96% on held-out derivatives.
-
InFact: Informativeness Alignment for Improved LLM Factuality
InFACT trains LLMs with hierarchical informativeness rewards plus abstention, improving factual precision on QA benchmarks while largely preserving recall.
-
Mind the Gap! Choice Independence in Using Multilingual LLMs for Persuasive Co-Writing Tasks in Different Languages
Users who first used a Spanish AI writing assistant subsequently used the English AI writing assistant less, suggesting a spillover that violates choice independence.
-
SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence
SelfElicit uses deep-layer attention to automatically highlight relevant evidence sentences in the input context, yielding consistent QA accuracy gains across six instruction-tuned LLMs.
-
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning
Merging adjacent reasoning steps into coarser training steps for process reward models improves best-of-n accuracy on GSM-Plus and MATH500 by about 0.5 to 3.4 percentage points.
-
Consolidating TinyML Lifecycle with Large Language Models: Reality, Illusion, or Opportunity?
An LLM-based framework automates TinyML data processing and model conversion reliably, but automated Arduino sketch generation fails in 63.3% of runs, making full automation an open challenge.
-
Semantic Drift and the Stability of Operator Control in Reasoning-Class Decision Support Systems
Reasoning LLMs in ultra-long sessions exhibit latent semantic drift that inverts operator control; a fitted stability coefficient Ks detects the bifurcation and a latent-steering arbitrator is proposed to restore it.
-
Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details
For Other-Play in Yokai, agents trained with different implementation details coordinate across implementations about as well as across seeds, supporting inter-seed cross-play as a proxy for cross-implementation evaluation.
-
Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models
Backdoor text samples show smaller log-probability changes under mask-filling perturbations than clean samples, which enables zero-shot backdoor detection without the poisoned model.
-
What Language(s) Does Aya-23 Think In? How Multilinguality Affects Internal Language Representations
Aya-23-8B appears to activate multiple related languages internally and concentrate code-mixing neurons in final layers, but the paper's own limitations undercut the claim that these are language-specific neurons.
-
The Impact of Fine-tuning Large Language Models on Automated Program Repair
On three Java APR benchmarks, LoRA and IA3 adapters match or beat full-model fine-tuning for most tested code LLMs while training less than one percent of parameters.
-
Hallucination Detection with Small Language Models
A multi-small-model ensemble with sentence splitting, z-score normalization, and harmonic mean detects hallucinations in RAG answers with a reported 10% F1 gain over single-model baselines.
-
DLM-One: Diffusion Language Models for One-Step Sequence Generation
DLM-One distills a continuous diffusion language model into a one-step student, achieving roughly 500x inference speedup while staying within a few percent of the teacher on BLEU, ROUGE, and BERTScore, with substantia...
-
AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity
AnchorAttention uses the maximum attention score from initial and local tokens as an anchor to threshold-select important key-value positions at stripe granularity, achieving faster prefill with comparable accuracy.
-
To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization
An EM-style training loop lets 7B math LLMs learn when to invoke code, improving MATH500 by 11 points and AIME by 9.4 points.
-
Generative AI Uses and Risks for Knowledge Workers in a Science Organization
At Argonne National Lab, early adopters of generative AI reported copilot and workflow agent use cases, small but growing usage, and concerns about reliability, privacy, academic publishing, and jobs.
-
PromptShield: Deployable Detection for Prompt Injection Attacks
PromptShield reports a 65.3% true positive rate at 0.1% false positive rate for prompt injection detection, more than six times the best prior model, on its own out-of-distribution evaluation split.
-
Psychologically Enhanced AI Agents
MBTI personality prompts measurably change how LLM agents write stories and play strategic games, with self-reflection before communication supporting cooperative behavior.
-
Insights into User Interface Innovations from a Design Thinking Workshop at deRSE25
A workshop at deRSE25 produced seven user-interface sketches for LLMs that emphasize branching, context management, and user weighting, which the authors map onto their whiteboard-based interface concept.
-
LOCOFY Large Design Models -- Design to code conversion solution
A proprietary design-to-code pipeline is described with claimed high fidelity and LLM outperformance, but the evaluation is self-referential, unquantified, and unreproducible.
-
Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques
A survey that maps LLM applications, vulnerabilities, and defenses across eight cybersecurity domains, but with significant citation and rigor problems.
-
Exploring the Limits of Model Compression in LLMs: A Knowledge Distillation Study on QA Tasks
Distilled students at 43% to 50% of teacher size keep over 90% of teacher Exact Match on SQuAD and MLQA, though one-shot gains reverse on SQuAD test for Pythia.
-
Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications
The paper defines urban LLM agents, surveys their sensing, memory, reasoning, execution, and learning workflows, and organizes their applications across planning, transportation, environment, safety, and society.
-
Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead
A literature review organizes LLM development into a six-phase software engineering lifecycle and identifies challenges and research directions for each phase.
-
LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models
A block-localizing fine-tuning method for gender debiasing is presented, but its stated loss is inconsistent with its reported behavior and the evaluation tables contain duplicate rows.
-
Boosting Self-Efficacy and Performance of Large Language Models via Verbal Efficacy Stimulations
Emotionally styled verbal prompts (encouraging, provocative, critical) modestly improve zero-shot LLM accuracy on many tasks, with the best style varying by model and task zone.
-
Dynamic benchmarking framework for LLM-based conversational data capture
An LLM-based framework that benchmarks conversational data capture using synthetic users, applied to loan applications, shows adaptive follow-up questions improve extraction accuracy.
-
AI Governance through Markets
Market governance mechanisms, supported by standardized AI disclosures, can create financial incentives for responsible AI development, according to this policy paper.
-
Towards Advancing Code Generation with Large Language Models: A Research Roadmap
A roadmap paper that organizes LLM code generation into a six-layer architecture and a four-phase human-in-the-loop workflow, and lists open challenges and recommendations.
-
Visual RAG: Expanding MLLM visual knowledge without fine-tuning
Retrieval-selected demonstration examples let a multimodal LLM classify images as accurately as random many-shot prompting with far fewer examples.
-
Adaptive Parameter-Efficient Federated Fine-Tuning on Heterogeneous Devices
Assigning federated fine-tuning devices different numbers of LoRA layers near the output, with ranks increasing toward the output, reaches target accuracy 1.5-2.8x faster and with up to 42.3% less communication than e...
-
Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems
Multi-stage prompt inference attacks against enterprise LLMs are formalized and defenses are proposed, but the preprint gives no reproducible evidence for its central claims.
-
Is It Time To Treat Prompts As Code? A Multi-Use Case Study For Prompt Optimization Using DSPy
A five-task case study shows DSPy prompt optimization can improve LLM accuracy on some tasks, notably contradiction detection (46.2% to 64.0%), but results vary and no code or data are released.
-
From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs
LLMs can assist both attackers and defenders in cybersecurity, but context limits, hallucinations, and weak reasoning make them unsafe to deploy without human oversight and real-world evaluation.
-
Evaluation of LLMs for mathematical problem solving
A three-model, three-dataset LLM math evaluation using a multi-dimensional reasoning rubric, undermined by contradictory accuracy tables.
-
IntelliChain: An Integrated Framework for Enhanced Socratic Method Dialogue with LLMs and Knowledge Graphs
IntelliChain is an LLM plus knowledge graph tutoring framework for Socratic math teaching, but its claimed benefit rests on one qualitative example, not measured results.
-
Why Are Positional Encodings Nonessential for Deep Autoregressive Transformers? Revisiting a Petroglyph
A didactic review showing that multi-layer autoregressive Transformers can infer position from the causal mask and context alone, so explicit positional encodings are unnecessary beyond one layer.
-
Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
A survey applying the Kaddour et al. challenge taxonomy to GPT-4o and DeepSeek-V3-0324, concluding closed models favor safety while open models favor cost and customization.
-
A Comprehensive Review of Human Error in Risk-Informed Decision Making: Integrating Human Reliability Assessment, Artificial Intelligence, and Human Performance Models
A review of human error research concluding that integrating AI and cognitive models into human reliability assessment can markedly improve predictive fidelity, but data scarcity and opacity remain barriers.
Discussion (0). Continue with ORCID to comment.