Pith. sign in

REVIEW 2 cited by

Responsible Task Automation: Empowering Large Language Models as Responsible Task Automators

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.01242 v2 pith:SMFN5UXI submitted 2023-06-02 cs.AI cs.CL

classification cs.AIcs.CL
keywords responsibletaskautomationexecutorsllmsmodelstaskscapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent success of Large Language Models (LLMs) signifies an impressive stride towards artificial general intelligence. They have shown a promising prospect in automatically completing tasks upon user instructions, functioning as brain-like coordinators. The associated risks will be revealed as we delegate an increasing number of tasks to machines for automated completion. A big question emerges: how can we make machines behave responsibly when helping humans automate tasks as personal copilots? In this paper, we explore this question in depth from the perspectives of feasibility, completeness and security. In specific, we present Responsible Task Automation (ResponsibleTA) as a fundamental framework to facilitate responsible collaboration between LLM-based coordinators and executors for task automation with three empowered capabilities: 1) predicting the feasibility of the commands for executors; 2) verifying the completeness of executors; 3) enhancing the security (e.g., the protection of users' privacy). We further propose and compare two paradigms for implementing the first two capabilities. One is to leverage the generic knowledge of LLMs themselves via prompt engineering while the other is to adopt domain-specific learnable models. Moreover, we introduce a local memory mechanism for achieving the third capability. We evaluate our proposed ResponsibleTA on UI task automation and hope it could bring more attentions to ensuring LLMs more responsible in diverse scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AutoDroid-V2: Boosting SLM-based GUI Agents via Code Generation

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A document-guided script-based agent lets an on-device 8B language model complete mobile tasks in a single generated program, beating step-wise agents on DroidTask and AitW.

  2. Automated Soap Opera Testing Directed by LLMs and Scenario Knowledge: Feasibility, Challenges, and Road Ahead

    cs.SE 2024-12 conditional novelty 6.0 of 10

    A multi-agent LLM system with a scenario knowledge graph can automate soap opera testing on Android apps, finding real bugs but with more false positives than manual testing.

Pith tools