iPOE generates and optimizes annotation guidelines from explanations to produce interpretable prompts, reporting up to 39% gains over baselines on four datasets with LLM explanations substituting for human ones.
Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems , articleno =
10 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Attention heads that decode to task-descriptive words in vocabulary space are shared across prompting styles, and their activation strength causally explains behavioral variance in LLMs.
SLoW selects low-frequency word dictionaries to boost LLM translation quality and efficiency across 100 languages from FLORES.
MLLMs given the same instructions as human participants achieve expert-level performance on perceiving stress in network visualizations and rely on similar visual proxies.
DIP interleaves English word translations into non-English prompts to boost multilingual reasoning on synthetic benchmarks spanning 10-200 languages.
Multimodal LLMs applied to 16 real-world configurators using 18 synthesized criteria can identify usability issues and generate actionable suggestions, with human review confirming reliability.
Neural Computers are introduced as a new machine form where computation, memory, and I/O are unified in a learned runtime state, with initial video-model experiments showing acquisition of basic interface primitives from traces.
LRP-based attention head selection and distributed application improve the efficiency and accuracy of function vectors for steering LLMs compared to prior choices.
Proposes extending preregistration practices to AI agent experiments and supplies a tailored template to limit researcher degrees of freedom.
The study compares MLLM-generated usability evaluations against expert assessments on prioritization of issues and introduces an interactive visualization tool for reviewing model outputs.
citing papers explorer
-
iPOE: Interpretable Prompt Optimization via Explanations
iPOE generates and optimizes annotation guidelines from explanations to produce interpretable prompts, reporting up to 39% gains over baselines on four datasets with LLM explanations substituting for human ones.
-
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
Attention heads that decode to task-descriptive words in vocabulary space are shared across prompting styles, and their activation strength causally explains behavioral variance in LLMs.
-
SLoW: Select Low-frequency Words! Automatic Dictionary Selection for Translation on Large Language Models
SLoW selects low-frequency word dictionaries to boost LLM translation quality and efficiency across 100 languages from FLORES.
-
Exploring MLLMs Perception of Network Visualization Principles
MLLMs given the same instructions as human participants achieve expert-level performance on perceiving stress in network visualizations and rely on similar visual proxies.
-
Dictionary Insertion Prompting for Multilingual Reasoning on Multilingual Large Language Models
DIP interleaves English word translations into non-English prompts to boost multilingual reasoning on synthetic benchmarks spanning 10-200 languages.
-
Usability Analysis of Configurator User Interfaces with Multimodal Large Language Models
Multimodal LLMs applied to 16 real-world configurators using 18 synthesized criteria can identify usability issues and generate actionable suggestions, with human review confirming reliability.
-
Neural Computers
Neural Computers are introduced as a new machine form where computation, memory, and I/O are unified in a learned runtime state, with initial video-model experiments showing acquisition of basic interface primitives from traces.
-
Fast & Faithful Function Vectors
LRP-based attention head selection and distributed application improve the efficiency and accuracy of function vectors for steering LLMs compared to prior choices.
-
Preregistration for Experiments with AI Agents
Proposes extending preregistration practices to AI agent experiments and supplies a tailored template to limit researcher degrees of freedom.
-
Investigating Multimodal Large Language Models to Support Usability Evaluation
The study compares MLLM-generated usability evaluations against expert assessments on prioritization of issues and introduces an interactive visualization tool for reviewing model outputs.