NEBULA：提升VLA评估与诊断的统一生态系统

cs.AI updates on arXiv.org 前天 12:22

NEBULA：提升VLA评估与诊断的统一生态系统

本文提出NEBULA，一个针对单臂操作VLA的统一生态系统，通过细粒度能力测试和系统压力测试，实现精准技能诊断和鲁棒性评估，以支持可复现研究和通用模型开发。

arXiv:2510.16263v1 Announce Type: cross Abstract: The evaluation of Vision-Language-Action (VLA) agents is hindered by the coarse, end-task success metric that fails to provide precise skill diagnosis or measure robustness to real-world perturbations. This challenge is exacerbated by a fragmented data landscape that impedes reproducible research and the development of generalist models. To address these limitations, we introduce \textbf{NEBULA}, a unified ecosystem for single-arm manipulation that enables diagnostic and reproducible evaluation. NEBULA features a novel dual-axis evaluation protocol that combines fine-grained \textit{capability tests} for precise skill diagnosis with systematic \textit{stress tests} that measure robustness. A standardized API and a large-scale, aggregated dataset are provided to reduce fragmentation and support cross-dataset training and fair comparison. Using NEBULA, we demonstrate that top-performing VLAs struggle with key capabilities such as spatial reasoning and dynamic adaptation, which are consistently obscured by conventional end-task success metrics. By measuring both what an agent can do and when it does so reliably, NEBULA provides a practical foundation for robust, general-purpose embodied agents.

Fish AI Reader

AI辅助创作，多种专业模板，深度分析，高质量内容生成。从观点提取到深度思考，FishAI为您提供全方位的创作支持。新版本引入自定义参数，让您的创作更加个性化和精准。

FishAI

鱼阅，AI 时代的下一个智能信息助手，助你摆脱信息焦虑

联系邮箱 441953276@qq.com

相关标签

VLA评估技能诊断鲁棒性评估 NEBULA 可复现研究

相关文章

Layer 1 区块链 Vega 提议关闭 Vega Alpha 主网和代币 VEGA，并启动新项目 Nebula

【Layer 1区块链Vega提议关闭Vega Alpha主网和代币VEGA，并启动新项目Nebula】9月3日消息，专注于衍生品交易的Layer 1区块链Vega Protocol发起提案提议在未来几...

Hierarchical Testing with Rabbit Optimization for Industrial Cyber-Physical Systems

May I have your Attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks

On Barriers to Archival Audio Processing

Metamorphic Testing of Deep Code Models: A Systematic Literature Review

Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics

Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in Large Language Models

MetAdv: A Unified and Interactive Adversarial Testing Platform for Autonomous Driving

ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering