Step-Level Preference Learning for Generative Agents in Social Simulations
This research introduces a step-level preference learning method to better align the intermediate decision-making processes of generative agents with human preferences.
This research introduces a step-level preference learning method to better align the intermediate decision-making processes of generative agents with human preferences.
Researchers have developed an adaptive layer-freezing technique for federated learning to reduce the computational burden of training models in resource-constrained healthcare environments.
Researchers introduced a schema-aware grounding method to improve the accuracy of agentic text-to-SPARQL query generation for knowledge base question answering.
The authors introduce a rollout-based training method to improve the performance of constrained diffusion models in satisfying complex feasibility requirements.
This research evaluates the strategic decision-making capabilities of vision-language models by testing their performance in simulated soccer scenarios.
RetroAgent utilizes large language models to navigate structured memory for more effective multi-step retrosynthesis planning in chemical research.
The study examines the limitations and potential misapplications of using item response theory to evaluate artificial intelligence benchmarks.
The authors introduce a visual masked autoencoder combined with normalizing flows to enhance the generalization capabilities of time series anomaly detection models.
A new benchmark called Alipay-PIBench has been introduced to evaluate the performance of AI coding agents in handling complex payment integration tasks.
This research proposes a framework to transition LLM-based agents from stateless interactions to situated models capable of maintaining temporal continuity for psychological support.
The study demonstrates that high transition accuracy in LLM-synthesized world models does not necessarily correlate with effective planning performance.
Researchers have developed SOReL, a Bayesian model-based reinforcement learning method that enables hyperparameter selection without requiring online environment interactions.
A new post-training method for multimodal document question answering focuses on reasoning-free alignment to reduce computational costs and improve visual grounding.
Researchers introduced a new benchmark designed to evaluate how effectively AI agents adapt to the continuous updates and changes within Model Context Protocol servers.
This paper models the value of connectivity within networks of AI agents to determine optimal collaboration protocols.
Researchers investigate the efficacy of specialized agentic systems compared to general-purpose LLMs for automating business process workflows.
A new method called SciDiagramEdit uses natural language instructions to automate the complex process of editing scientific diagrams and figures.
A new contrastive framework called SARA aims to improve the robustness of preference-based reinforcement learning against noisy or inaccurate human labeling.
Researchers have developed a new evaluation framework called JKP to test the epistemic stability of vision-language models when subjected to repetitive adversarial questioning.
A new flow-matching technique for feature-space world modeling addresses the trade-off between reconstruction quality and predictive accuracy in stochastic environments.
Researchers introduced Contrastive Policy Optimization to improve reinforcement learning by using token-level disagreement to better distinguish between useful uncertainty and harmful confusion.
A study evaluated the performance of OpenAI's o4-mini model on undergraduate physics problems to assess the reasoning capabilities of current inference-scaling LLMs.
LIGO-PINN introduces a gated optimization technique to improve the training stability and convergence of physics-informed neural networks in complex domains.
The study demonstrates that using demographically-conditioned synthetic medical images can help detect and mitigate bias in diagnostic classifiers.