ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples
The ARMOR framework introduces off-policy anchor samples to stabilize reinforcement learning in large language models and prevent over-optimization.
The ARMOR framework introduces off-policy anchor samples to stabilize reinforcement learning in large language models and prevent over-optimization.
Pipette is a new simulation platform and benchmark created to improve the training and data efficiency of robotics systems used in biomedical laboratory environments.
This study performs a comparative analysis of different code-execution environments to determine their impact on the performance of AI coding agents.
The study investigates the limitations of LLMs in collaborative settings, specifically their difficulty in managing pragmatic communication when information is asymmetrically distributed.
Researchers propose a method to improve the interpretability of automated industrial process control recommendations using gradient-based explanation techniques.
A new neural architecture search framework utilizes swarm intelligence and transformer controllers to enable efficient model design on consumer-grade hardware.
Researchers have developed a new simulation testbed designed to evaluate how AI models handle complex, long-horizon business decision-making.
A new multi-agent framework automates the generation of verifiable rules for classifying chemical reactions to improve computer-assisted synthesis planning.
MG2-RAG is a proposed multimodal retrieval-augmented generation framework that uses multi-granularity graphs to improve reasoning and preserve fine-grained visual data.
Researchers have developed a relational database bridge to connect bibliographic metadata with formalized mathematical proof libraries.
A multi-scale vision transformer approach is detailed for identifying multiple plant species in high-resolution photographs using limited training data.
Researchers have developed a deep learning framework using YOLO and explainable AI techniques to automate the taxonomic identification of parasitoid wasps.
A new method using Gaussian Process Regression has been proposed to optimize value estimation in Monte Carlo Tree Search for continuous action spaces.
This review paper examines the evolution of algorithms for the maximum clique problem, comparing classical, AI-based, and quantum computing approaches.
WrAFT is a modular automated writing evaluation system designed to provide scoring and feedback for argumentative essays using various large language models.
A new platform called CrimeNER Demo has been introduced to facilitate named-entity recognition and classification of crime-related information in documents.
This study examines how users perceive their own authorship when collaborating with generative AI tools.
Researchers developed a tool to quantify human versus AI contributions in creative works to address ongoing debates regarding artistic ownership.
Researchers evaluated the performance of five prominent world-model agents in Atari Pong to better understand their isolated capabilities.
This research proposes a model for optimizing power and channel performance in magnetic inductive cellular networks used in underground environments.
A new approach for comparing scientific document versions integrates layout-aware alignment and structural reasoning to handle complex elements like tables and formulas.
Researchers developed a text-centered multimodal system designed to recognize human ambivalence and hesitancy in video data.
A new study explores methods for AI agents to infer personality traits from facial images to improve human-robot interaction.
The HABIB_TAZ system uses synthetic training and multi-objective optimization to help language models separate formal logic from real-world content biases.