{"work":{"id":"80797370-dc29-417d-83c5-ec1313a64166","openalex_id":null,"doi":null,"arxiv_id":"2401.10019","raw_key":null,"title":"R-Judge: Benchmarking Safety Risk Awareness for LLM Agents","authors":null,"authors_text":"T","year":2024,"venue":"cs.CL","abstract":"Large language models (LLMs) have exhibited great potential in autonomously completing tasks across real-world applications. Despite this, these LLM agents introduce unexpected safety risks when operating in interactive environments. Instead of centering on the harmlessness of LLM-generated content in most prior studies, this work addresses the imperative need for benchmarking the behavioral safety of LLM agents within diverse environments. We introduce R-Judge, a benchmark crafted to evaluate the proficiency of LLMs in judging and identifying safety risks given agent interaction records. R-Judge comprises 569 records of multi-turn agent interaction, encompassing 27 key risk scenarios among 5 application categories and 10 risk types. It is of high-quality curation with annotated safety labels and risk descriptions. Evaluation of 11 LLMs on R-Judge shows considerable room for enhancing the risk awareness of LLMs: The best-performing model, GPT-4o, achieves 74.42% while no other models significantly exceed the random. Moreover, we reveal that risk awareness in open agent scenarios is a multi-dimensional capability involving knowledge and reasoning, thus challenging for LLMs. With further experiments, we find that fine-tuning on safety judgment significantly improve model performance while straightforward prompting mechanisms fail. R-Judge is publicly available at https://github.com/Lordog/R-Judge.","external_url":"https://arxiv.org/abs/2401.10019","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-11T02:47:50.885307+00:00","pith_arxiv_id":"2401.10019","created_at":"2026-05-11T05:51:10.931162+00:00","updated_at":"2026-07-11T02:47:50.885307+00:00","title_quality_ok":true,"display_title":"arXiv preprint arXiv:2401.10019 , year=","render_title":"arXiv preprint arXiv:2401.10019 , year="},"hub":{"state":{"work_id":"80797370-dc29-417d-83c5-ec1313a64166","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":18,"external_cited_by_count":null,"distinct_field_count":7,"first_pith_cited_at":"2024-02-05T23:06:42+00:00","last_pith_cited_at":"2026-07-08T17:34:28+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T17:59:51.935881+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":2}],"polarity_counts":[{"context_polarity":"background","n":2}],"runs":{},"summary":{},"graph":{},"authors":[]}}