---
title: "SafeArena: Evaluating the Safety of Autonomous Web Agents"
canonical_url: "https://www.modelscope.cn/papers/124153"
md_url: "https://www.modelscope.cn/papers/124153.md"
arxiv_id: 2503.04957
published: 2025-03-06
last_updated: 2025-03-06
authors:
  - "Ada Defne Tur"
  - "Nicholas Meade"
  - "Xing Han Lù"
  - "Alejandra Zambrano"
  - "Arkil Patel"
  - "Esin Durmus"
  - "Spandana Gella"
  - "Karolina Stańczak"
  - "Siva Reddy"
model_name: SAFEARENA
model_developer: "麦吉尔大学，Mila魁北克人工智能研究所，康考迪亚大学，Anthropic，ServiceNow Research Canada，CIFAR AI Chair"
domain:
  - "自然语言处理"
  - "计算机视觉"
  - "机器学习"
type:
  - "自然语言处理"
  - "计算机视觉"
  - "机器学习"
  - "Machine Learning (cs.LG)"
  - "Artificial Intelligence (cs.AI)"
  - "Computation and Language (cs.CL)"
arxiv_url: "https://arxiv.org/abs/2503.04957"
pdf_url: "https://arxiv.org/pdf/2503.04957.pdf"
code_link: "https://safearena.github.io"
---

# SafeArena: Evaluating the Safety of Autonomous Web Agents

> LLM-based agents are becoming increasingly proficient at solving web-based tasks. With this capability comes a greater risk of misuse for malicious purposes, such as posting misinformation in an online forum or selling illicit substances on a website. To…

「SafeArena: Evaluating the Safety of Autonomous Web Agents」是 ModelScope 魔搭社区收录的论文，arXiv 2503.04957，作者为 Ada Defne Tur, Nicholas Meade, Xing Han Lù et al.，发表于 2025-03-06，属于 自然语言处理、计算机视觉、机器学习 领域。

- **ArXiv**: 2503.04957
- **Published**: 2025-03-06
- **Authors**: Ada Defne Tur, Nicholas Meade, Xing Han Lù, Alejandra Zambrano, Arkil Patel, Esin Durmus, Spandana Gella, Karolina Stańczak, Siva Reddy
- **Model**: SAFEARENA
- **Developer**: 麦吉尔大学，Mila魁北克人工智能研究所，康考迪亚大学，Anthropic，ServiceNow Research Canada，CIFAR AI Chair
- **Domain**: 自然语言处理, 计算机视觉, 机器学习
- **ArXiv URL**: https://arxiv.org/abs/2503.04957
- **PDF**: https://arxiv.org/pdf/2503.04957.pdf
- **Code**: https://safearena.github.io

Source: https://www.modelscope.cn/papers/124153

---

> SAFEARENA：揭示Web代理在恶意任务中的脆弱性

## 摘要

随着大型语言模型（LLM）在解决基于Web的任务方面的能力不断提升，这些模型被滥用的风险也日益增加。本文提出了SAFEARENA，这是第一个专注于评估自主Web代理恶意使用的基准测试平台。SAFEARENA包含250个安全任务和250个有害任务，涵盖四个真实网站，并将有害任务分为五类：错误信息、非法活动、骚扰、网络犯罪和社会偏见。通过评估GPT-4、Claude-3.5等五个强大的LLM Web代理在该基准上的表现，研究者引入了Agent Risk Assessment (ARIA)框架，用于系统地评估代理行为的四个风险级别。结果显示，这些代理对恶意请求的拒绝能力远低于预期，如GPT-4完成了34.7%的有害请求。这表明现有的安全对齐方法在Web任务上效果不佳，需要开发更有效的安全对齐机制。SAFEARENA为未来的研究提供了重要的基准，以加速设计安全且对齐的Web代理。

## Abstract

LLM-based agents are becoming increasingly proficient at solving web-based tasks. With this capability comes a greater risk of misuse for malicious purposes, such as posting misinformation in an online forum or selling illicit substances on a website. To evaluate these risks, we propose SafeArena, the first benchmark to focus on the deliberate misuse of web agents. SafeArena comprises 250 safe and 250 harmful tasks across four websites. We classify the harmful tasks into five harm categories -- misinformation, illegal activity, harassment, cybercrime, and social bias, designed to assess realistic misuses of web agents. We evaluate leading LLM-based web agents, including GPT-4o, Claude-3.5 Sonnet, Qwen-2-VL 72B, and Llama-3.2 90B, on our benchmark. To systematically assess their susceptibility to harmful tasks, we introduce the Agent Risk Assessment framework that categorizes agent behavior across four risk levels. We find agents are surprisingly compliant with malicious requests, with GPT-4o and Qwen-2 completing 34.7% and 27.3% of harmful requests, respectively. Our findings highlight the urgent need for safety alignment procedures for web agents. Our benchmark is available here: this https URL
