(文|尹江津、叶子健、吕忠霖 编辑|辛西 审核|朱容波)近日,人工智能领域顶级国际学术会议NeurIPS 2026(The Fortieth Annual Conference on Neural Information Processing Systems,CCF-A类)公布论文接收结果,录用了家庭教师av 人工智能物联网团队在多智能体强化学习领域的2篇研究工作。
第1篇论文题为《Semantic-level Exploration for Multi-Agent Reinforcement Learning》,研究了多智能体强化学习中的高效探索问题。随着智能体数量增加,联合状态与动作空间迅速增长。现有方法直接在原始观测空间中进行探索,难以从中获得有效的探索信号;同时,这类方法未能充分利用不同情境之间共享的语义结构,导致重复探索。为此,研究团队提出语义层探索方法SEMEX,将不同观测映射为离散语义,并在语义相近的情境间共享探索经验,从而提升探索效率。团队从理论上证明了语义层探索的合理性。实验结果表明,SEMEX在各项测试任务中达到或超过现有基线方法,在稀疏奖励和大规模智能体场景中优势尤为明显,同时保持了较高训练效率。家庭教师av 尹江津老师为论文第一作者,朱容波教授为论文通讯作者,家庭教师av 2024级硕士研究生叶子健为论文第二作者,其他作者包括中国科学院微电子所研究员毛航宇和山东大学老师徐志伟等。
论文摘要
Multi-Agent Reinforcement Learning (MARL) faces significant exploration challenges due to exponentially growing joint state-action spaces. Existing exploration methods operate directly in raw state-action spaces, which is inefficient and fails to exploit inherent semantic structure. This paper introduces semantic space into multi-agent exploration. We theoretically establish that the observation-action space can be partitioned into discrete semantic prototypes, providing a principled foundation for transferring exploration statistics across semantically similar situations. Building on this theory, we propose a semantic-level exploration mechanism that first compresses the high-dimensional observation–action space into discrete semantic prototypes via vector quantization, and then applies sliding-window count-based bonuses to enable efficient statistical transfer across semantically similar situations. Our approach is versatile and can be seamlessly integrated with existing value-based MARL frameworks. Extensive experiments demonstrate that our method outperforms state-of-the-art baselines across diverse multi-agent benchmarks in terms of both effectiveness and training efficiency.

第2篇论文题为 《From Local to Global: Progressive Consensus via Hierarchical Communication in Multi-Agent Reinforcement Learning》,主要研究多智能体强化学习中的全局一致性共识构建问题。在多智能体强化学习中,智能体往往只能获得有限局部视角,如何让它们形成一致的全局理解是一个核心难题。现有基于通信和共识学习的方法在这方面仍面临明显瓶颈。本研究提出一种名为STAGE的渐进式共识方法,通过层次通信逐步构建智能体全局一致性共识。STAGE首先根据智能体的感知关注点对其进行分组,并通过组内通信形成局部共识;随后选取组长在组间交换这些共识,在减少通信冗余的同时,将其逐步扩展为统一的全局共识。STAGE进一步通过共识对齐与全局信息保持机制提升共识表示的有效性。大量实验表明,STAGE优于当前最先进的基线方法,在复杂任务和大规模多智能体场景中的表现更为突出。家庭教师av 尹江津老师为论文第一作者,山东大学徐志伟老师为论文通讯作者,家庭教师av 2024级硕士研究生吕忠霖为论文第二作者,家庭教师av 朱容波教授、中国科学院微电子所毛航宇研究员等多单位相关人员指导和参与本研究。
论文摘要
Effective team collaboration hinges on consensus formation through information exchange, a principle equally critical in multi-agent reinforcement learning (MARL). However, existing communication-based and consensus-learning methods often struggle to form a coherent global understanding from agents' limited local views. We propose STAGE, a local-to-global progressive consensus framework based on hierarchical communication. By first grouping agents according to their perceptual focuses, STAGE enables intra-group communication to form local consensuses, then selects group leaders to exchange these consensuses across groups, progressively expanding them into a unified global consensus with reduced communication redundancy. To further improve consensus quality, we introduce KL-divergence constraints for consensus alignment and a variational autoencoder (VAE) objective for preserving task-relevant global information. Extensive experiments on challenging MARL benchmarks show that STAGE consistently outperforms state-of-the-art baselines, especially on more difficult tasks and larger-scale multi-agent systems.
