Reassessing Muon for Matrix Factorization
This paper provides a theoretical analysis of the Muon optimizer to clarify the mechanisms behind its performance advantages in large-scale deep learning.
This paper provides a theoretical analysis of the Muon optimizer to clarify the mechanisms behind its performance advantages in large-scale deep learning.
A new framework enables mobile robots to interpret natural language commands for autonomous navigation using RGB-D perception.
A new benchmark based on cognitive psychology tests has been developed to evaluate how AI agents adapt when the reliability of their tools changes during operation.
A study evaluates the scalability and classroom performance of an AI tutoring agent that integrates retrieval-augmented generation with structured knowledge models.
This paper proposes a framework for networked intelligence that uses shared context graphs to facilitate collaboration among multiple AI agents in scientific research.
A study evaluating deep learning models for echocardiography analysis reveals that current attribution methods often fail to verify temporal faithfulness in medical imaging predictions.
The study evaluates the impact of audio separation preprocessing on zero-shot automatic speech recognition performance, finding that cleaner audio does not always improve transcription accuracy.
A new decoupled training strategy for transfer learning aims to reduce computational and energy costs by optimizing classifier heads separately from feature extraction layers.
The paper provides a theoretical framework linking Joint-Embedding Predictive Architectures to active inference principles through variational free energy.
The proposed ScanFocus framework addresses computational efficiency and precision in spatio-temporal video grounding through a coarse-to-fine processing approach.
The ExTernD method introduces a post-training factorization technique for large language models that utilizes expanded-rank ternary decomposition to improve quantization efficiency and accuracy.
The authors introduce STKAN, a spatio-temporal forecasting architecture that leverages Kolmogorov-Arnold Networks to better model complex traffic data.
A study evaluating root cause analysis in microservice failures reveals that current AI and classical methods struggle to effectively process large-scale, multimodal telemetry data.
The GHR-VLM framework combines grounded reasoning with vision-language models to enable zero-shot video analytics for transit systems without requiring task-specific training data.
An analysis of Kaggle contest submissions examines whether the use of AI coding assistants is leading to increased homogenization of software development outputs.
A new defense mechanism called Traffic-Aware Randomized Smoothing has been proposed to improve the robustness of LLM-based network intrusion detection systems against traffic manipulation.
The authors introduced Monty, a framework designed to improve the accuracy of autoformalization by converting natural language specifications into executable software assertions.
MedDiffuseMix provides a saliency-guided diffusion framework to augment medical imaging data while preserving critical diagnostic features.
FixItFlow is an automated system that leverages large language models to generate technical troubleshooting guides from historical cloud incident data.
HealthClaw is a proposed self-evolving agent architecture designed to provide longitudinal health management by maintaining private memory of user routines and medical history.
Researchers have proposed a training-free approach for detecting human-object interactions in the wild by leveraging the capabilities of multimodal large language models.
The Multi-Agent System Process Reward Model provides a method to evaluate and optimize message sequences between agents during inference-time search.
Researchers analyzed nearly 3,000 GitHub projects to understand how the integration of automated bots as active participants influences the organizational structure of open-source software teams.
The authors present a method for integrating compiler feedback directly into the autoregressive decoding process to improve the quality of AI-generated code.