EvoGraph turns linear AI-assisted programming into a manipulable graph of branching histories, reducing cognitive load and enabling better iteration according to a user study with 20 developers.
hub
Ross, Fernando Martinez, Stephanie Houde, Michael Muller, and Justin D
16 Pith papers cite this work, alongside 36 external citations. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
background 3polarities
background 3representative citing papers
Novices performed better and reported lower workload with GitHub Copilot than with human partners, but human partners produced more positive emotions and a smaller drop in retest performance after one week.
Empirical study finds instruction tuning on CodeLLMs improves instruction following at the expense of infilling performance, termed the Instruction-Tuning Tax.
PaperFlow proposes a Profiling-Recommending-Adapting framework for longitudinal scientific paper recommendation and evaluates it on a new user-day benchmark with 24 simulated users, outperforming five baselines in ranking, behavioral alignment, and blind human evaluation.
A new evaluation framework shows that even the best tested LLM only reliably adjusts response complexity in the intended direction 46% of the time across 98 scientific queries.
A design space for reading augmentation that contrasts transmission-oriented 'reading to discard' with transformation-oriented 'creative reading' and identifies two opportunities for future systems.
DoubleAgents shows that a distributed-cognition design with coordination agent, dashboard, and policy module increases user comfort and reliance on AI agents for coordination tasks over time.
Mixed-methods study of 27 developers characterizes five Copilot chat interaction modes and ten needs linked to problem-solving styles and experience levels.
Mixed-methods study adapting UTAUT2 shows individual-level perceptions predict continued GenAI use in Italian SME developers (R²=0.647) while social and organisational factors do not.
Among novice programmers using AI code generators, trust did not predict compliance with suggestions, while performance correlated with both compliance and increased subsequent trust.
Interviews reveal a four-stage vibe coding workflow that accelerates prototyping while introducing tensions between quick efficiency and reflective design intention, plus asymmetries in trust and ownership.
A survey of 87 agents for computer use and 33 datasets that introduces a three-dimensional taxonomy across domain, interaction, and agent perspectives and identifies six research gaps.
Qualitative study of 20 interviews and 24 workshop participants finds AI-driven automation and human-AI collaboration are shifting development roles in SAP BTP and require updates to the existing BTP User Type Matrix.
Novice programmers completed more tasks with lower workload using GitHub Copilot versus a human partner, but reported significantly more positive and arousing emotions with the human teammate.
A survey of user studies on LLM use in programming that identifies interaction behaviors, mixed benefits and weaknesses, and factors influencing human and task performance.
The paper describes ongoing efforts to characterize developer diversity in cognition and context and to use personalization to make LLM-based conversational programming assistants more inclusive.
citing papers explorer
-
Choose Your Own Adventure: Non-Linear AI-Assisted Programming with EvoGraph
EvoGraph turns linear AI-assisted programming into a manipulable graph of branching histories, reducing cognitive load and enabling better iteration according to a user study with 20 developers.
-
Fast and Forgettable: A Controlled Study of Novices' Performance, Learning, Workload, and Emotion in AI-Assisted and Human Pair Programming Paradigms
Novices performed better and reported lower workload with GitHub Copilot than with human partners, but human partners produced more positive emotions and a smaller drop in retest performance after one week.
-
Lost in the Flow with Code Talkers: Unveiling the Instruction-Tuning Tax of Large Language Models in Code Tasks
Empirical study finds instruction tuning on CodeLLMs improves instruction following at the expense of infilling performance, termed the Instruction-Tuning Tax.
-
PaperFlow: Profiling, Recommending, and Adapting Across Daily Paper Streams
PaperFlow proposes a Profiling-Recommending-Adapting framework for longitudinal scientific paper recommendation and evaluates it on a new user-day benchmark with 24 simulated users, outperforming five baselines in ranking, behavioral alignment, and blind human evaluation.
-
Explain Like I'm 5 or Whatever I Choose: Evaluating the Interactive Potential of Language Model Responses
A new evaluation framework shows that even the best tested LLM only reliably adjusts response complexity in the intended direction 46% of the time across 98 scientific queries.
-
Creative Reading: Scaffolding Reading for Transformation
A design space for reading augmentation that contrasts transmission-oriented 'reading to discard' with transformation-oriented 'creative reading' and identifies two opportunities for future systems.
-
DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow
DoubleAgents shows that a distributed-cognition design with coordination agent, dashboard, and policy module increases user comfort and reliance on AI agents for coordination tasks over time.
-
No Two Developers Think Alike: How Problem-Solving Styles and Experience Shape Needs in Conversational Interaction with Copilot
Mixed-methods study of 27 developers characterizes five Copilot chat interaction modes and ten needs linked to problem-solving styles and experience levels.
-
From Early Adoption to Sustained Use: Understanding GenAI Usage Among Software Developers in Italian SMEs
Mixed-methods study adapting UTAUT2 shows individual-level perceptions predict continued GenAI use in Italian SME developers (R²=0.647) while social and organisational factors do not.
-
Relationships Between Trust, Compliance, and Performance for Novice Programmers Using AI Code Generation
Among novice programmers using AI code generators, trust did not predict compliance with suggestions, while performance correlated with both compliance and increased subsequent trust.
-
Vibe Coding in Product Teams: Reconfiguring AI-Assisted Workflows, Prototyping, and Collaboration
Interviews reveal a four-stage vibe coding workflow that accelerates prototyping while introducing tensions between quick efficiency and reflective design intention, plus asymmetries in trust and ownership.
-
A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions
A survey of 87 agents for computer use and 33 datasets that introduces a three-dimensional taxonomy across domain, interaction, and agent perspectives and identifies six research gaps.
-
The impact of artificial intelligence on enterprise software user roles
Qualitative study of 20 interviews and 24 workshop participants finds AI-driven automation and human-AI collaboration are shifting development roles in SAP BTP and require updates to the existing BTP User Type Matrix.
-
OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning
Novice programmers completed more tasks with lower workload using GitHub Copilot versus a human partner, but reported significantly more positive and arousing emotions with the human teammate.
-
Understanding the Human-LLM Dynamic: A Literature Survey of LLM Use in Programming Tasks
A survey of user studies on LLM use in programming that identifies interaction behaviors, mixed benefits and weaknesses, and factors influencing human and task performance.
-
Personalizing LLM-Based Conversational Programming Assistants
The paper describes ongoing efforts to characterize developer diversity in cognition and context and to use personalization to make LLM-based conversational programming assistants more inclusive.