PM-Bench: Evaluating Prospective Memory in LLM Agents
Researchers introduced PM-Bench, a text-based benchmark designed to evaluate the prospective memory capabilities of large language model agents.
Researchers introduced PM-Bench, a text-based benchmark designed to evaluate the prospective memory capabilities of large language model agents.
A new framework for hardware design uses stepwise refinement to help large language models generate more reliable and verifiable register-transfer level code.
The authors introduce an agentic approach for localizing vulnerability triggers in code by performing interprocedural causal reasoning.
Fin-Analyst is a hybrid trading agent that utilizes a multi-specialist LLM pipeline to process diverse financial data sources for automated equity market analysis.
A new signal-guided optimization technique aims to improve machine unlearning by addressing the varying memorization strengths of individual training samples.
The authors introduce a method called CARE-LoRA designed to reduce memory usage during the fine-tuning of large pre-trained models.
Researchers propose a vendor-neutral metric to evaluate the reconstructability and validity of AI agent safety testing results.
The study introduces DRIFTLENS to measure how personalized memory injection in LLMs can alter the reasoning trajectories used to generate responses.
Researchers identified a unified geometric mechanism in transformer models that explains how conflicting memory sources lead to confident hallucinations.
Research indicates that current state-of-the-art LLMs struggle significantly with bidirectional Korean-Braille translation, highlighting gaps in accessibility-focused capabilities.
Researchers have introduced Visual Access Sweep, a causal intervention method to study how Vision-Language Models utilize image tokens during long Chain-of-Thought reasoning.
Researchers introduced a new method using execution-based semantic interaction graphs to better quantify uncertainty in code-generating large language models.
The authors investigate how the training duration of individual domain experts influences the performance of merged large language models.
Researchers proposed a method for automatic speech recognition using a discrete diffusion language model to transcribe audio in parallel rather than through traditional autoregressive decoding.
This survey examines various inference optimization techniques designed to improve the practical speed of masked diffusion large language models.
Researchers demonstrated that using structured, line-anchored feedback in AI code editing tools significantly reduces token consumption and improves the accuracy of generated changes.
This paper investigates a vulnerability in large language model plan evaluators where strategic plans are rewarded for omitting explicit details.
A biologically inspired learning method called mistake-gated learning is introduced to reduce the energy and memory requirements of training neural networks.
Researchers proposed a compliance-aware federated learning framework that adjusts differential privacy noise to accommodate varying institutional data standards and resource levels.
A comparative study using NLP metrics reveals similarities and differences in how humans and leading large language models navigate conceptual spaces during semantic memory retrieval.
Researchers have developed a method to improve the real-time control of robots by optimizing the asynchronous inference of vision-language-action models on edge hardware.
Researchers have developed a variation-aware entropy scheduling method to improve reinforcement learning performance in environments subject to drift.
PalmClaw is a new framework designed to enable native, multi-step agentic task execution directly on mobile devices.
This research proposes a framework for detecting health misinformation in low-resource languages by combining small language models with culturally sensitive NLP techniques.