REVIEW 74 cited by
Challenges and Applications of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large Language Models (LLMs) went from non-existent to ubiquitous in the machine learning discourse within a few years. Due to the fast pace of the field, it is difficult to identify the remaining challenges and already fruitful application areas. In this paper, we aim to establish a systematic set of open problems and application successes so that ML researchers can comprehend the field's current state more quickly and become productive.
Forward citations
Showing 60 of 74 Pith papers that cite this
-
CXXCrafter: An LLM-Based Agent for Automated C/C++ Open Source Software Building
An LLM-driven agent, CXXCrafter, automatically builds 587 of 752 C/C++ open-source projects (78%), beating default build commands (39%) and bare LLMs (32 to 38%).
-
RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models
RAG can make language models less safe than their non-RAG equivalents, even with safe documents, and current jailbreak methods transfer poorly to RAG.
-
Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
Jailbreak attacks push LLM activations outside a safety boundary, mostly in low and middle layers, and a tanh-based penalty that pulls activations back inside this boundary blocks most tested attacks with under 2% uti...
-
Hidden Language Consistency Phenomena in Reasoning LLMs
Reasoning models often stop using the requested language as problems get harder, and this language breakdown can make accuracy look better than it is.
-
FastTPS: An Optimized Method for LLM Token Phase for AI accelerators
FastTPS accelerates LLM token-phase inference via reloading-free static KV-cache management, tiled fused RoPE attention, and interlaced-weight MLP fusion, yielding up to 6× speedup at 93% bandwidth on AMD NPUs.
-
Interpreting learning dynamics of autoencoders: Transient scaling and emerging concepts of the Ising model
Unsupervised autoencoders on Ising configurations form magnetization then energy representations in two dynamical regimes, with recursive error flow fields sharing topology across layers.
-
Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions
Across 74 OSINT/CTI AI studies, hallucination is widely named but end-to-end measured in only one non-reproducible system, so a human–AI co-pilot is the most defensible near-term architecture.
-
Real Faults in Model Context Protocol (MCP) Software: a Comprehensive Taxonomy
MCP server faults form five empirical categories—server setting, server/tool configuration, server/host configuration, documentation, and general programming—confirmed by a 41-practitioner survey.
-
Enhancing Robustness of Autoregressive Language Models against Orthographic Attacks via Pixel-based Approach
A word-as-image pixel language model trained with next-token prediction reports lower perplexity than a token-embedding LLaMA on noisy and non-Latin-script text, though its noise evaluation holds tokenization fixed.
-
PhantomHunter: Detecting Unseen Privately-Tuned LLM-Generated Text via Family-Aware Learning
PhantomHunter detects text from privately fine-tuned LLMs by learning shared token-probability traits within LLaMA, Gemma and Mistral families, reporting F1 above 96% on held-out derivatives.
-
InFact: Informativeness Alignment for Improved LLM Factuality
InFACT trains LLMs with hierarchical informativeness rewards plus abstention, improving factual precision on QA benchmarks while largely preserving recall.
-
Real-Time Verification of Embodied Reasoning for Generative Skill Acquisition
VERGSA trains a process reward model on MCTS-labeled subtask outcomes and uses it to select scene configurations and subtask supervisions, improving simulated task success rates.
-
MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks
MedArabiQ is a seven-task Arabic medical benchmark showing that closed models generally beat open ones on structured questions, while BERTScore misses serious hallucinations that an LLM judge later reveals.
-
OET: Optimization-based prompt injection Evaluation Toolkit
OET is an optimization-based evaluation toolkit that benchmarks prompt injection attacks and defenses across eight datasets and shows current defenses remain vulnerable in several domains.
-
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
A large-scale comparison of six hallucination detection metric families across 37 models and five decoding methods finds most metrics align poorly with human judgments, with GPT-4-based evaluation performing best.
-
Helping Blind People Grasp: Enhancing a Tactile Bracelet with an Automated Hand Navigation System
An automated vision-to-vibration hand navigation system on a tactile bracelet lets blindfolded and blind users grasp target objects, track one instance among distractors, and avoid obstacles.
-
Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding
DVBench introduces 10,000 expert-annotated questions on crash and near-crash driving videos and reports that no tested vision LLM exceeds 40 percent accuracy under its strict GroupEval scoring.
-
Mind the Gap! Choice Independence in Using Multilingual LLMs for Persuasive Co-Writing Tasks in Different Languages
Users who first used a Spanish AI writing assistant subsequently used the English AI writing assistant less, suggesting a spillover that violates choice independence.
-
SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence
SelfElicit uses deep-layer attention to automatically highlight relevant evidence sentences in the input context, yielding consistent QA accuracy gains across six instruction-tuned LLMs.
-
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning
Merging adjacent reasoning steps into coarser training steps for process reward models improves best-of-n accuracy on GSM-Plus and MATH500 by about 0.5 to 3.4 percentage points.
-
Consolidating TinyML Lifecycle with Large Language Models: Reality, Illusion, or Opportunity?
An LLM-based framework automates TinyML data processing and model conversion reliably, but automated Arduino sketch generation fails in 63.3% of runs, making full automation an open challenge.
-
Context-DPO: Aligning Language Models for Context-Faithfulness
Context-DPO fine-tunes LLMs with direct preference optimization on counterfactual passages, yielding 35-280% context-faithfulness gains on its new ConFiQA benchmark.
-
Feature Coding in the Era of Large Models: Dataset, Test Conditions, and Benchmark
A public benchmark and unified test conditions for compressing intermediate features of large models, with two image-codec baselines evaluated.
-
SMoLoRA: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction Tuning
SMoLoRA uses two separately routed LoRA expert groups, one for visual understanding and one for instruction following, to reduce dual catastrophic forgetting in continual visual instruction tuning.
-
Neon: News Entity-Interaction Extraction for Enhanced Question Answering
Neon builds a timestamped knowledge graph of entity-event tuples extracted from news, and augmenting LLM prompts with these tuples improves temporal entity-centric question answering.
-
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
Near-perfect jailbreak defenses for vision-language models are mostly over-refusal, and the two standard ways of scoring jailbreaks agree only at chance level.
-
Semantic Drift and the Stability of Operator Control in Reasoning-Class Decision Support Systems
Reasoning LLMs in ultra-long sessions exhibit latent semantic drift that inverts operator control; a fitted stability coefficient Ks detects the bifurcation and a latent-steering arbitrator is proposed to restore it.
-
Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details
For Other-Play in Yokai, agents trained with different implementation details coordinate across implementations about as well as across seeds, supporting inter-seed cross-play as a proxy for cross-implementation evaluation.
-
Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models
Backdoor text samples show smaller log-probability changes under mask-filling perturbations than clean samples, which enables zero-shot backdoor detection without the poisoned model.
-
What Language(s) Does Aya-23 Think In? How Multilinguality Affects Internal Language Representations
Aya-23-8B appears to activate multiple related languages internally and concentrate code-mixing neurons in final layers, but the paper's own limitations undercut the claim that these are language-specific neurons.
-
The Impact of Fine-tuning Large Language Models on Automated Program Repair
On three Java APR benchmarks, LoRA and IA3 adapters match or beat full-model fine-tuning for most tested code LLMs while training less than one percent of parameters.
-
Hallucination Detection with Small Language Models
A multi-small-model ensemble with sentence splitting, z-score normalization, and harmonic mean detects hallucinations in RAG answers with a reported 10% F1 gain over single-model baselines.
-
DLM-One: Diffusion Language Models for One-Step Sequence Generation
DLM-One distills a continuous diffusion language model into a one-step student, achieving roughly 500x inference speedup while staying within a few percent of the teacher on BLEU, ROUGE, and BERTScore, with substantia...
-
AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity
AnchorAttention uses the maximum attention score from initial and local tokens as an anchor to threshold-select important key-value positions at stripe granularity, achieving faster prefill with comparable accuracy.
-
Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation
Ordering retrieved documents along the reasoning chain and replacing the causal mask with a prefix mask during LoRA fine-tuning improves multi-hop QA accuracy; peak attention scores can select the best context order.
-
Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models
The paper introduces enigme, a procedurally generated text-puzzle library for benchmarking reasoning in transformer-decoder language models.
-
Understanding and Mitigating Risks of Generative AI in Financial Services
Open-source AI guardrails miss most financial-services content risks that a new domain-specific taxonomy identifies, even when their prompts are expanded to cover the new categories.
-
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
A large collaborative survey organizes LLM and LLM-agent safety issues into a full-stack lifecycle framework from data preparation to deployment.
-
To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization
An EM-style training loop lets 7B math LLMs learn when to invoke code, improving MATH500 by 11 points and AIME by 9.4 points.
-
Generative AI Uses and Risks for Knowledge Workers in a Science Organization
At Argonne National Lab, early adopters of generative AI reported copilot and workflow agent use cases, small but growing usage, and concerns about reliability, privacy, academic publishing, and jobs.
-
PromptShield: Deployable Detection for Prompt Injection Attacks
PromptShield reports a 65.3% true positive rate at 0.1% false positive rate for prompt injection detection, more than six times the best prior model, on its own out-of-distribution evaluation split.
-
LLM Augmentations to support Analytical Reasoning over Multiple Documents
LLMs alone and with dynamic evidence tree augmentation still fail to produce the implicit, speculative reasoning that intelligence analysis requires.
-
Developer Challenges on Large Language Models: A Study of Stack Overflow and OpenAI Developer Forum Posts
A BERTopic analysis of 8,593 Stack Overflow posts and 26,474 OpenAI Developer Forum posts yields 9 and 17 LLM developer challenge topics, with API usage dominant and high unresolved rates.
-
CoCoP: Enhancing Text Classification with LLM through Code Completion Prompt
CoCoP, which formats text classification as code completion, improves LLM accuracy over few-shot prompting and lets small code models approach large general models.
-
Psychologically Enhanced AI Agents
MBTI personality prompts measurably change how LLM agents write stories and play strategic games, with self-reflection before communication supporting cooperative behavior.
-
Insights into User Interface Innovations from a Design Thinking Workshop at deRSE25
A workshop at deRSE25 produced seven user-interface sketches for LLMs that emphasize branching, context management, and user weighting, which the authors map onto their whiteboard-based interface concept.
-
LOCOFY Large Design Models -- Design to code conversion solution
A proprietary design-to-code pipeline is described with claimed high fidelity and LLM outperformance, but the evaluation is self-referential, unquantified, and unreproducible.
-
Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques
A survey that maps LLM applications, vulnerabilities, and defenses across eight cybersecurity domains, but with significant citation and rigor problems.
-
Exploring the Limits of Model Compression in LLMs: A Knowledge Distillation Study on QA Tasks
Distilled students at 43% to 50% of teacher size keep over 90% of teacher Exact Match on SQuAD and MLQA, though one-shot gains reverse on SQuAD test for Pythia.
-
Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications
The paper defines urban LLM agents, surveys their sensing, memory, reasoning, execution, and learning workflows, and organizes their applications across planning, transportation, environment, safety, and society.
-
Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead
A literature review organizes LLM development into a six-phase software engineering lifecycle and identifies challenges and research directions for each phase.
-
LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models
A block-localizing fine-tuning method for gender debiasing is presented, but its stated loss is inconsistent with its reported behavior and the evaluation tables contain duplicate rows.
-
Validating the Effectiveness of a Large Language Model-based Approach for Identifying Children's Development across Various Free Play Settings in Kindergarten
An LLM-based pipeline labeled kindergarten play narratives and achieved high rater agreement, but the claimed validity as a measure of child development is not supported by the evidence.
-
Boosting Self-Efficacy and Performance of Large Language Models via Verbal Efficacy Stimulations
Emotionally styled verbal prompts (encouraging, provocative, critical) modestly improve zero-shot LLM accuracy on many tasks, with the best style varying by model and task zone.
-
Dynamic benchmarking framework for LLM-based conversational data capture
An LLM-based framework that benchmarks conversational data capture using synthetic users, applied to loan applications, shows adaptive follow-up questions improve extraction accuracy.
-
AI Governance through Markets
Market governance mechanisms, supported by standardized AI disclosures, can create financial incentives for responsible AI development, according to this policy paper.
-
Towards Advancing Code Generation with Large Language Models: A Research Roadmap
A roadmap paper that organizes LLM code generation into a six-layer architecture and a four-phase human-in-the-loop workflow, and lists open challenges and recommendations.
-
Visual RAG: Expanding MLLM visual knowledge without fine-tuning
Retrieval-selected demonstration examples let a multimodal LLM classify images as accurately as random many-shot prompting with far fewer examples.
-
Adaptive Parameter-Efficient Federated Fine-Tuning on Heterogeneous Devices
Assigning federated fine-tuning devices different numbers of LoRA layers near the output, with ranks increasing toward the output, reaches target accuracy 1.5-2.8x faster and with up to 42.3% less communication than e...
-
RAG Playground: A Framework for Systematic Evaluation of Retrieval Strategies and Prompt Engineering in RAG Systems
A new RAG evaluation framework reports that hybrid vector-keyword retrieval and structured self-evaluation prompting improve answer quality, reaching a 72.7% pass rate on its own unvalidated metrics.
Discussion (0). Continue with ORCID to comment.