---
title: "API Agents vs. GUI Agents: Divergence and Convergence"
canonical_url: "https://www.modelscope.cn/papers/126771"
md_url: "https://www.modelscope.cn/papers/126771.md"
arxiv_id: 2503.11069
published: 2025-03-14
last_updated: 2025-03-14
authors:
  - "Chaoyun Zhang"
  - "Shilin He"
  - "Liqun Li"
  - "Si Qin"
  - "Yu Kang"
  - "Qingwei Lin"
  - "Dongmei Zhang"
model_developer: "微软"
domain:
  - "自然语言处理"
  - "人工智能"
  - "软件工程"
type:
  - "自然语言处理"
  - "人工智能"
  - "软件工程"
  - "Artificial Intelligence (cs.AI)"
  - "Human-Computer Interaction (cs.HC)"
arxiv_url: "https://arxiv.org/abs/2503.11069"
pdf_url: "https://arxiv.org/pdf/2503.11069.pdf"
---

# API Agents vs. GUI Agents: Divergence and Convergence

> Large language models (LLMs) have evolved beyond simple text generation to power software agents that directly translate natural language commands into tangible actions. While API-based LLM agents initially rose to prominence for their robust automation…

「API Agents vs. GUI Agents: Divergence and Convergence」是 ModelScope 魔搭社区收录的论文，arXiv 2503.11069，作者为 Chaoyun Zhang, Shilin He, Liqun Li et al.，发表于 2025-03-14，属于 自然语言处理、人工智能、软件工程 领域。

- **ArXiv**: 2503.11069
- **Published**: 2025-03-14
- **Authors**: Chaoyun Zhang, Shilin He, Liqun Li, Si Qin, Yu Kang, Qingwei Lin, Dongmei Zhang
- **Developer**: 微软
- **Domain**: 自然语言处理, 人工智能, 软件工程
- **ArXiv URL**: https://arxiv.org/abs/2503.11069
- **PDF**: https://arxiv.org/pdf/2503.11069.pdf

Source: https://www.modelscope.cn/papers/126771

---

> API与GUI之争：大型语言模型代理的分道扬镳与殊途同归

## 摘要

本文探讨了基于API和基于GUI的大型语言模型（LLM）代理在软件自动化中的差异与潜在融合。随着LLM的发展，这些代理不仅能生成文本，还能将自然语言指令转化为实际操作。API-based LLM代理通过调用预定义的程序接口实现高效、可靠的自动化任务，而GUI-based LLM代理则通过模拟人类交互直接操作图形用户界面，提供更灵活但可能较慢的解决方案。研究背景显示，API-based代理最初因高效性和易集成性而受到关注，但随着多模态技术的进步，GUI-based代理逐渐展现出其独特优势，例如更高的灵活性和直观的用户体验。提出的方法中，作者系统分析了两种代理在模态、可靠性、效率、可用性、灵活性、安全性、可维护性和透明度等多个维度上的差异，并讨论了混合方法的可能性。实验结果表明，API-based代理在稳定性、效率和安全性方面表现突出，而GUI-based代理则更适合需要高灵活性和直观交互的场景。该研究的价值贡献在于为从业者和研究人员提供了清晰的选择标准和实际用例，指导如何根据项目需求选择或结合这两种范式，同时预测未来创新将使两者界限更加模糊，推动更灵活的自动化解决方案。

## Abstract

Large language models (LLMs) have evolved beyond simple text generation to power software agents that directly translate natural language commands into tangible actions. While API-based LLM agents initially rose to prominence for their robust automation capabilities and seamless integration with programmatic endpoints, recent progress in multimodal LLM research has enabled GUI-based LLM agents that interact with graphical user interfaces in a human-like manner. Although these two paradigms share the goal of enabling LLM-driven task automation, they diverge significantly in architectural complexity, development workflows, and user interaction models.
This paper presents the first comprehensive comparative study of API-based and GUI-based LLM agents, systematically analyzing their divergence and potential convergence. We examine key dimensions and highlight scenarios in which hybrid approaches can harness their complementary strengths. By proposing clear decision criteria and illustrating practical use cases, we aim to guide practitioners and researchers in selecting, combining, or transitioning between these paradigms. Ultimately, we indicate that continuing innovations in LLM-based automation are poised to blur the lines between API- and GUI-driven agents, paving the way for more flexible, adaptive solutions in a wide range of real-world applications.
