Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate
A study demonstrates that the interaction protocols used in multi-agent debates significantly influence the moral reasoning and judgment outcomes of large language models.
A study demonstrates that the interaction protocols used in multi-agent debates significantly influence the moral reasoning and judgment outcomes of large language models.
A new error-correction method called Experience Memory Graph aims to help LLM agents recover from failures in long-horizon tasks more efficiently than traditional reflection techniques.
A new auditing method uses predicate substitution to test whether large language models genuinely rely on stated premises during chain-of-thought reasoning.
DIVE is a new dimensionality reduction technique for language model embeddings that uses self-limiting gradient updates to improve compression efficiency.
A new framework enables mobile robots to interpret natural language commands for autonomous navigation using RGB-D perception.
MedDiffuseMix provides a saliency-guided diffusion framework to augment medical imaging data while preserving critical diagnostic features.
A new method for test-time learning introduces learnable adaptation policies to help language agents improve their performance through iterative interaction.
The study evaluates the impact of audio separation preprocessing on zero-shot automatic speech recognition performance, finding that cleaner audio does not always improve transcription accuracy.
The Multi-Agent System Process Reward Model provides a method to evaluate and optimize message sequences between agents during inference-time search.
This paper proposes a causal framework to determine when artificial intelligence systems should engage in theory of mind processes during conflict scenarios.
The authors present a method for integrating compiler feedback directly into the autoregressive decoding process to improve the quality of AI-generated code.
Researchers have proposed a training-free approach for detecting human-object interactions in the wild by leveraging the capabilities of multimodal large language models.
A new defense mechanism called Traffic-Aware Randomized Smoothing has been proposed to improve the robustness of LLM-based network intrusion detection systems against traffic manipulation.
A study evaluating deep learning models for echocardiography analysis reveals that current attribution methods often fail to verify temporal faithfulness in medical imaging predictions.
The GHR-VLM framework combines grounded reasoning with vision-language models to enable zero-shot video analytics for transit systems without requiring task-specific training data.
The ExTernD method introduces a post-training factorization technique for large language models that utilizes expanded-rank ternary decomposition to improve quantization efficiency and accuracy.
A new framework for discrete diffusion models explores how tokenization and vocabulary structure influence generative performance.
The proposed ScanFocus framework addresses computational efficiency and precision in spatio-temporal video grounding through a coarse-to-fine processing approach.
A study evaluates the scalability and classroom performance of an AI tutoring agent that integrates retrieval-augmented generation with structured knowledge models.
The authors introduced Monty, a framework designed to improve the accuracy of autoformalization by converting natural language specifications into executable software assertions.
This paper provides a theoretical analysis of the Muon optimizer to clarify the mechanisms behind its performance advantages in large-scale deep learning.
The SteinGate framework introduces a new safety certification method for reinforcement learning to better mitigate rare but catastrophic risks.
The authors introduce STKAN, a spatio-temporal forecasting architecture that leverages Kolmogorov-Arnold Networks to better model complex traffic data.
An analysis of Kaggle contest submissions examines whether the use of AI coding assistants is leading to increased homogenization of software development outputs.