✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
4
ResearcharXiv cs.AI·7d agoPrimary

Closed-Loop Knowledge Dynamics: An Operational Framework for Saturation and Escape

Researchers propose a theoretical framework to address performance saturation in closed-loop AI systems by identifying how external information can overcome internal feedback limitations.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Energy-Efficient Federated Learning via Adaptive Encoder Freezing for MRI-to-CT Conversion: A Green AI-Guided Research

Researchers have developed an adaptive layer-freezing technique for federated learning to reduce the computational burden of training models in resource-constrained healthcare environments.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning

A new contrastive framework called SARA aims to improve the robustness of preference-based reinforcement learning against noisy or inaccurate human labeling.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Fully Offline Reinforcement Learning

Researchers have developed SOReL, a Bayesian model-based reinforcement learning method that enables hyperparameter selection without requiring online environment interactions.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Reinforcement Learning in Switching Non-Stationary Markov Decision Processes: Algorithms and Convergence Analysis

This paper presents a mathematical framework for reinforcement learning in environments where the underlying dynamics switch between different states.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon Strategy Game Planning

The SAGA agent framework is proposed to improve long-horizon strategic planning in complex games by better managing relational structures and evolving goals.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

From Stateless to Situated: Building a Psychological World for LLM-Based Agents

This research proposes a framework to transition LLM-based agents from stateless interactions to situated models capable of maintaining temporal continuity for psychological support.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

When Does Belief-Based Agent Memory Help? Reliability-Conditional Updating and Provenance-Capped Poisoning Defense

A new study evaluates the effectiveness of belief-based memory architectures in LLM agents using Bayesian inference to manage long-term information.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

HG-RAG: Hierarchy-Guided Retrieval-Augmented Generation for Structured Knowledge Graphs

Researchers have developed a framework called HG-RAG that improves retrieval-augmented generation by utilizing hierarchical graph traversal for more effective knowledge retrieval.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Global drivers and barriers to the public acceptance of autonomous vehicles: Evidence from 17 countries

A multi-country study identifies the primary global factors influencing public acceptance of Level 3 autonomous vehicles.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability

ToolAnchor is a proposed method to help AI agents overcome behavioral inertia and more effectively integrate new tools into their workflows.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models

The study demonstrates that high transition accuracy in LLM-synthesized world models does not necessarily correlate with effective planning performance.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI

This paper argues for the application of sociotechnical systems analysis to better understand and mitigate risks in automated decision-making technologies.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Assessing AI in Introductory Physics Problem Solving

A study evaluated the performance of OpenAI's o4-mini model on undergraduate physics problems to assess the reasoning capabilities of current inference-scaling LLMs.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

ViPSAM: Visual Prompting Medical Image Segmentation Using Segment Anything Model

ViPSAM is a new visual prompting framework designed to improve medical image segmentation in non-contrast CT scans by leveraging contrast-enhanced MRI data.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution

Researchers investigate the efficacy of specialized agentic systems compared to general-purpose LLMs for automating business process workflows.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Step-Level Preference Learning for Generative Agents in Social Simulations

This research introduces a step-level preference learning method to better align the intermediate decision-making processes of generative agents with human preferences.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation

Researchers introduced a schema-aware grounding method to improve the accuracy of agentic text-to-SPARQL query generation for knowledge base question answering.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

RetroAgent: Harnessing LLMs to Search Over Structured Memory for Agentic Retrosynthesis Planning

RetroAgent utilizes large language models to navigate structured memory for more effective multi-step retrosynthesis planning in chemical research.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Integration Matters: Rollout-Based Training for Constrained Diffusion Models

The authors introduce a rollout-based training method to improve the performance of constrained diffusion models in satisfying complex feasibility requirements.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

MAPS: Modeling Co-Existing Subjective Perspectives and Shared Meaning in Multi-Agent Cognitive Dialogue

The proposed MAPS framework enables multi-agent dialogue systems to model distinct subjective perspectives and cognitive styles during interaction.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows

Researchers introduced a benchmark designed to evaluate the accuracy and traceability of LLM agents performing complex structural engineering workflows.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation

This study investigates common assumptions regarding prompting techniques and dataset usage in the evaluation of large language models for multiple-choice tasks.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Can LLMs Build a MaxSAT Solver from Papers? The CoreForge Experience

The CoreForge project demonstrates an iterative workflow where large language models are used to construct a MaxSAT solver directly from academic literature rather than existing codebases.

Read at arxiv.org ↗
← NewerPage 144Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.