最新发布的论文《DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows》提出了一项新的基准测试方法。该测试基于200个真实用户会话重建任务,涵盖编码、网络调研、文档处理和内容创作四类常见工作流,并刻意加入了缺失依赖、不稳定网络和干扰文件等环境复杂度。
论文指出,Agent框架对任务的拆解方式、工具调用策略以及错误恢复机制,会影响模型在复杂场景下的表现。研究通过这一基准测试,观察了不同框架下同一模型的表现差异。
![]()
特别声明:以上内容(如有图片或视频亦包括在内)为自媒体平台“网易号”用户上传并发布,本平台仅提供信息存储服务。
Notice: The content above (including the pictures and videos if any) is uploaded and posted by a user of NetEase Hao, which is a social media platform and only provides information storage services.