✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
4
ResearcharXiv cs.AI·7d agoPrimary

Step-Level Preference Learning for Generative Agents in Social Simulations

This research introduces a step-level preference learning method to better align the intermediate decision-making processes of generative agents with human preferences.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Energy-Efficient Federated Learning via Adaptive Encoder Freezing for MRI-to-CT Conversion: A Green AI-Guided Research

Researchers have developed an adaptive layer-freezing technique for federated learning to reduce the computational burden of training models in resource-constrained healthcare environments.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation

Researchers introduced a schema-aware grounding method to improve the accuracy of agentic text-to-SPARQL query generation for knowledge base question answering.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Integration Matters: Rollout-Based Training for Constrained Diffusion Models

The authors introduce a rollout-based training method to improve the performance of constrained diffusion models in satisfying complex feasibility requirements.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

SportD: Can VLMs Physically Strategize?

This research evaluates the strategic decision-making capabilities of vision-language models by testing their performance in simulated soccer scenarios.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

RetroAgent: Harnessing LLMs to Search Over Structured Memory for Agentic Retrosynthesis Planning

RetroAgent utilizes large language models to navigate structured memory for more effective multi-step retrosynthesis planning in chemical research.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Can We Trust Item Response Theory for AI Evaluation?

The study examines the limitations and potential misapplications of using item response theory to evaluate artificial intelligence benchmarks.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

VAN-AD: Visual Masked Autoencoder with Normalizing Flow For Time Series Anomaly Detection

The authors introduce a visual masked autoencoder combined with normalizing flows to enhance the generalization capabilities of time series anomaly detection models.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents

A new benchmark called Alipay-PIBench has been introduced to evaluate the performance of AI coding agents in handling complex payment integration tasks.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

From Stateless to Situated: Building a Psychological World for LLM-Based Agents

This research proposes a framework to transition LLM-based agents from stateless interactions to situated models capable of maintaining temporal continuity for psychological support.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models

The study demonstrates that high transition accuracy in LLM-synthesized world models does not necessarily correlate with effective planning performance.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Fully Offline Reinforcement Learning

Researchers have developed SOReL, a Bayesian model-based reinforcement learning method that enables hyperparameter selection without requiring online environment interactions.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment

A new post-training method for multimodal document question answering focuses on reasoning-free alignment to reduce computational costs and improve visual grounding.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers

Researchers introduced a new benchmark designed to evaluate how effectively AI agents adapt to the continuous updates and changes within Model Context Protocol servers.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

ANet Patu-1: The Value of Connection in the Agent Network

This paper models the value of connectivity within networks of AI agents to determine optimal collaboration protocols.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution

Researchers investigate the efficacy of specialized agentic systems compared to general-purpose LLMs for automating business process workflows.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions

A new method called SciDiagramEdit uses natural language instructions to automate the complex process of editing scientific diagrams and figures.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning

A new contrastive framework called SARA aims to improve the robustness of preference-based reinforcement learning against noisy or inaccurate human labeling.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs

Researchers have developed a new evaluation framework called JKP to test the epistemic stability of vision-language models when subjected to repetitive adversarial questioning.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Flow Matching in Feature Space for Stochastic World Modeling

A new flow-matching technique for feature-space world modeling addresses the trade-off between reconstruction quality and predictive accuracy in stochastic environments.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

Researchers introduced Contrastive Policy Optimization to improve reinforcement learning by using token-level disagreement to better distinguish between useful uncertainty and harmful confusion.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Assessing AI in Introductory Physics Problem Solving

A study evaluated the performance of OpenAI's o4-mini model on undergraduate physics problems to assess the reasoning capabilities of current inference-scaling LLMs.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

LIGO-PINN: Learned Initialization via Gated Optimization to Alleviate Convergence Failures in Physics Informed Neural Networks

LIGO-PINN introduces a gated optimization technique to improve the training stability and convergence of physics-informed neural networks in complex domains.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Demographically-Conditioned Synthetic Medical Images for Bias Mitigation and Bias Detection in Disease Classifiers

The study demonstrates that using demographically-conditioned synthetic medical images can help detect and mitigate bias in diagnostic classifiers.

Read at arxiv.org ↗
← NewerPage 147Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.