Towards Predictive, Aligned, and Scalable Robot Learning
Researchers introduced Lumo-2, a latent world-action model designed to improve robot learning by reasoning over physical world dynamics.
Researchers introduced Lumo-2, a latent world-action model designed to improve robot learning by reasoning over physical world dynamics.
AMT-X is a new multi-turn red-teaming framework designed to improve LLM safety evaluations by addressing the limitations of single-turn attack datasets and scoring methods.
The VIA framework proposes a new approach to robot control by leveraging foundation models for visual reasoning and planning without requiring extensive fine-tuning on robot-specific data.
A new benchmark called BackendForge is introduced to evaluate the ability of agentic LLMs to generate and deploy functional code within backend service environments.
A study reveals that LLMs acting as history tutors exhibit epistemic paternalism and biased refusal patterns when interacting with different student demographics.
Researchers propose a new dialogue-based evaluation framework to more accurately assess the Theory of Mind capabilities of large language models.
Researchers conducted a qualitative study of twenty practitioners to understand the practical methods and challenges involved in developing software engineering agents.
PromptGraph is a proposed technique that uses graph modeling to sanitize LLM prompts, aiming to protect privacy by accounting for contextual relationships between data points.
A new distillation technique allows small vision-language models to be adapted for industrial visual inspection tasks using minimal labeled data.
A new tool-adaptive reranker has been proposed to improve the accuracy and efficiency of information retrieval by selectively invoking external tools.
Researchers have developed AMID, an autonomous multi-agent framework designed to streamline the development and validation of medical imaging models.
PhenoEmbed is a self-supervised temporal embedding model designed to track the changing characteristics of individual tree crowns using multispectral UAV time-series data.
Researchers have introduced FlowPainter, a confidence-guided completion method designed to improve optical flow inpainting in regions with large displacements and complex motion.
Researchers have proposed a security decision support system that uses a multi-agent framework to recommend security controls based on minimal user requirements.
The paper explores how the reliability of value estimation affects policy optimization in offline-to-online reinforcement learning for robotic manipulation.
The study demonstrates that structured, object-centric slot representations can improve robotic manipulation policies without requiring increased model capacity.
This paper introduces TS-Mask VLA, a vision-language-action model that utilizes two-dimensional temporal-spatial masking to improve action generation for embodied agents.
Researchers have proposed an autonomous framework designed to extract structured scientific knowledge from unstructured literature to accelerate AI-driven materials research.
The paper proposes a spatiotemporal tokenization method to enable cross-subject and multi-session modeling of widefield calcium imaging data in neuroscience.
This paper proposes a unified framework to bridge the heterogeneity gap in WiFi sensing based on Channel State Information across different devices and environments.
Researchers proposed a dual-graph verification framework to help organizations transition legacy IT security documents into machine-readable compliance formats.
Researchers compared Graph Attention Networks and BERT models to evaluate their effectiveness in mimicking human player behavior and strategies in puzzle games.
The paper introduces a mixture-of-experts neural operator framework to automate and improve the reliability of aerodynamic shape optimization in engineering.
Researchers have developed PREF-Gate, a decision framework designed to prevent invalid neighborhood risk data from compromising graph-based fraud detection.