ART optimizes visual pixel inputs to frozen MLLMs to achieve LoRA-competitive accuracy on math and structured tool-use benchmarks without modifying computational graphs.
In Proceedings of the 10th Annual Joint Conference on Digital Libraries, JCDL ’10, pages 11–20, New York, NY , USA
10 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
verdicts
UNVERDICTED 10roles
background 1polarities
support 1representative citing papers
RePrompT uses recurrent prompt tuning to inject prior-visit latent states and cohort-derived population prompt tokens into LLMs, yielding better performance than pure EHR or pure LLM baselines on MIMIC clinical prediction tasks.
LabelPigeon jointly performs translation and label projection via XML tags, improving translation quality in 11 languages and cross-lingual transfer by up to +40.2 F1 on NER across 27 languages.
AnnotateThis lets users improve LLM annotations for climate change mitigation pessimism on social media, yielding 0.15 higher F-Measure and 0.23 higher accuracy than automated prompt refinement when ground truth labels are available.
Targeted prompting and system interventions enable local LLMs such as Llama 3.1 70B to exploit 83% of tested Linux privilege escalation vulnerabilities.
Fine-tuned BERT sentence-pair classifiers reach F1 0.83 while few-shot LLM prompting reaches F1 0.78 on threat and solution framing detection in 440 manually coded German climate news articles.
On a controlled Turkish dataset of 147 examples, few-shot prompting lets some LLMs match or beat a supervised BERT baseline for LVC detection, though results are highly sensitive to prompt design.
Poodle shows that LLMs can be automatically replaced with cheaper models for recurring tasks to save significant cost and energy without extra user effort.
SPG-Layout combines statistical object priors with hierarchical large-object-first placement to produce physically plausible text-driven 3D scenes in non-Manhattan rooms and outperforms baselines on a new 500-scene benchmark.
Finetuning Phi Silica on curated short presentation text improves semantic fidelity, reduces hallucinations, and raises preference win rates over GPT-5-chat rewrites.
citing papers explorer
-
Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training
ART optimizes visual pixel inputs to frozen MLLMs to achieve LoRA-competitive accuracy on math and structured tool-use benchmarks without modifying computational graphs.
-
RePrompT: Recurrent Prompt Tuning for Integrating Structured EHR Encoders with Large Language Models
RePrompT uses recurrent prompt tuning to inject prior-visit latent states and cohort-derived population prompt tokens into LLMs, yielding better performance than pure EHR or pure LLM baselines on MIMIC clinical prediction tasks.
-
Just Use XML: Revisiting Joint Translation and Label Projection
LabelPigeon jointly performs translation and label projection via XML tags, improving translation quality in 11 languages and cross-lingual transfer by up to +40.2 F1 on NER across 27 languages.
-
AnnotateThis: Analyzing a human-LLM system for annotating social media data with the concept of climate change mitigation pessimism
AnnotateThis lets users improve LLM annotations for climate change mitigation pessimism on social media, yielding 0.15 higher F-Measure and 0.23 higher accuracy than automated prompt refinement when ground truth labels are available.
-
Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents
Targeted prompting and system interventions enable local LLMs such as Llama 3.1 70B to exploit 83% of tested Linux privilege escalation vulnerabilities.
-
Comparing BERT Sentence-Pair Classification and Few-Shot LLM Prompting for Detecting Threat and Solution Framing in German Climate News
Fine-tuned BERT sentence-pair classifiers reach F1 0.83 while few-shot LLM prompting reaches F1 0.78 on threat and solution framing detection in 440 manually coded German climate news articles.
-
Supervision versus Demonstration-Based In-Context Learning for Multiword Expression Classification
On a controlled Turkish dataset of 147 examples, few-shot prompting lets some LLMs match or beat a supervised BERT baseline for LVC detection, though results are highly sensitive to prompt design.
-
Poodle: Seamlessly Scaling Down Large Language Models with Just-in-Time Model Replacement
Poodle shows that LLMs can be automatically replaced with cheaper models for recurring tasks to save significant cost and energy without extra user effort.
-
Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments
SPG-Layout combines statistical object priors with hierarchical large-object-first placement to produce physically plausible text-driven 3D scenes in non-Manhattan rooms and outperforms baselines on a new 500-scene benchmark.
-
Short-form Text Rewriting with Phi Silica
Finetuning Phi Silica on curated short presentation text improves semantic fidelity, reduces hallucinations, and raises preference win rates over GPT-5-chat rewrites.