IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation
IdeaTrail is a new dataset designed to capture the complete, multi-stage workflow of AI agents performing scientific ideation and research tasks.
IdeaTrail is a new dataset designed to capture the complete, multi-stage workflow of AI agents performing scientific ideation and research tasks.
MemDecay is a new memory management technique that optimizes LLM agent inference by using region-aware KV cache eviction based on semantic structure.
The proposed Gefen optimizer reduces the memory footprint of large-scale model training by sharing second-moment estimates and quantizing first-moment states.
Embodied-R1.5 is a new foundation model designed to unify embodied reasoning and physical intelligence using a large-scale dataset of 15 billion tokens.
Researchers released FindMyText, an open-source Python package designed to detect near-verbatim text containment within large web-crawled datasets using document fingerprinting.
A new scheduling method for heterogeneous edge GPUs aims to improve the efficiency and throughput of vision transformer models in autonomous vehicle applications.
Researchers propose a new speculative decoding method that improves large language model inference speed by utilizing progressive tree drafting to better exploit parallel processing.
GrandCode is a new multi-agent reinforcement learning system that aims to achieve competitive programming performance at the level of human grandmasters.
A new metric called the 50%-task-completion time horizon has been proposed to better compare AI performance against human capabilities in software development tasks.
This study analyzes the internal attention dynamics of vision-language models to explain why visual grounding degrades and proposes scheduling visual relay windows to stabilize reasoning.
The study identifies a critical inefficiency in Chain-of-Thought prompting where models generate redundant but logically valid reasoning steps that current evaluators fail to penalize.
This paper outlines a deployment-focused pathway for transitioning medical AI agents from simple assistants to autonomous clinical systems.
The Mako project introduces a self-evolving agentic operating system capable of autonomously synthesizing and testing new security exploits.
Researchers have introduced SDABench, a new benchmark designed to evaluate the scientific data analysis capabilities of large language models across six distinct areas.
Agentic-DPO is a new training framework designed to optimize AI agent policies on expert trajectories by teaching them to avoid plausible mistakes rather than just imitating sequences.
Researchers propose a formal theory of least autonomy as a security principle to constrain the permissions and workflow capabilities of agentic AI systems.
The authors introduce SWE-MERA, a dynamic benchmark designed to mitigate data contamination and improve the evaluation of LLMs on software engineering tasks.
The paper introduces HCRMap, a mapping framework designed to mitigate compute and communication imbalances caused by uneven expert activation in Mixture-of-Experts models on multi-chiplet systems.
Researchers developed ReflectWorld-MM, a multimedia memory system that organizes long-term video stream data around persistent entities rather than individual frames.
Researchers have introduced a new benchmark dataset of mathematical constant formulas to evaluate the advanced mathematical reasoning capabilities of AI systems.
This paper examines the phenomenon of model collapse from both engineering and creative perspectives, exploring how recursive training on AI-generated data affects model output.
This paper provides a comprehensive survey and a new two-level taxonomy of Graph Neural Network methodologies applied across knowledge graph technologies.
This paper proposes a learnable Dirichlet-process cache that stores only novel inputs to bridge the gap between state-space models and attention mechanisms.
The authors introduce FARS, a fully automated system that enables AI agents to conduct research, generate hypotheses, and write manuscripts at scale.