Pith. sign in

REVIEW 5 cited by

CHOPS: CHat with custOmer Profile Systems for Customer Service with LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.01343 v4 pith:FFVHYIOW submitted 2024-03-31 cs.CL cs.AI

classification cs.CLcs.AI
keywords customerservicechopsexistingllmschatsystemsarchitecture
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Businesses and software platforms are increasingly turning to Large Language Models (LLMs) such as GPT-3.5, GPT-4, GLM-3, and LLaMa-2 for chat assistance with file access or as reasoning agents for customer service. However, current LLM-based customer service models have limited integration with customer profiles and lack the operational capabilities necessary for effective service. Moreover, existing API integrations emphasize diversity over the precision and error avoidance essential in real-world customer service scenarios. To address these issues, we propose an LLM agent named CHOPS (CHat with custOmer Profile in existing System), designed to: (1) efficiently utilize existing databases or systems for accessing user information or interacting with these systems following existing guidelines; (2) provide accurate and reasonable responses or carry out required operations in the system while avoiding harmful operations; and (3) leverage a combination of small and large LLMs to achieve satisfying performance at a reasonable inference cost. We introduce a practical dataset, the CPHOS-dataset, which includes a database, guiding files, and QA pairs collected from CPHOS, an online platform that facilitates the organization of simulated Physics Olympiads for high school teachers and students. We have conducted extensive experiments to validate the performance of our proposed CHOPS architecture using the CPHOS-dataset, with the aim of demonstrating how LLMs can enhance or serve as alternatives to human customer service. Code for our proposed architecture and dataset can be found at {https://github.com/JingzheShi/CHOPS}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PosterMate: Audience-driven Collaborative Persona Agents for Poster Design

    cs.HC 2025-07 conditional novelty 6.0 of 10

    A new design assistant creates audience persona agents from marketing briefs to provide poster feedback and moderated discussion, with user studies showing perceived usefulness and partial evidence for persona-consist...

  2. Effective Red-Teaming of Policy-Adherent Agents

    cs.MA 2025-06 conditional novelty 6.0 of 10

    A policy-aware red-teaming system (CRAFT) induces policy violations in LLM customer service agents at much higher rates than generic jailbreak prompts, using a new security-focused benchmark (tau-break) built from tau-bench.

  3. From Words to Workflows: Automating Business Processes

    cs.AI 2024-12 conditional novelty 5.0 of 10

    Text2Workflow is a multi-prompt LLM system with human feedback that generates JSON workflows from natural language, scoring 71.3% average semantic accuracy on the authors' 60-request Process2JSON dataset, versus 64.2%...

  4. MindFlow+: A Self-Evolving Agent for E-Commerce Customer Service

    cs.CL 2025-07 conditional novelty 4.0 of 10

    MindFlow+ combines tool-augmented demonstrations with reward-token-conditioned SFT to improve an LLM judge's assessment of e-commerce customer service, but only on private data with a self-tuned judge.

  5. LLM Security: Vulnerabilities, Attacks, Defenses, and Countermeasures

    cs.CR 2025-05 conditional novelty 3.0 of 10

    This survey categorizes attacks on large language models by lifecycle phase and maps them to prevention and detection defenses, concluding that only a few defenses are highly effective.

Pith tools