✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
4
ResearcharXiv cs.AI·7d agoPrimary

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning

A new contrastive framework called SARA aims to improve the robustness of preference-based reinforcement learning against noisy or inaccurate human labeling.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi

This study explores the feasibility of running modern multimodal AI models on legacy hardware by deploying a compact assistant on a 2011-era GPU.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Fully Offline Reinforcement Learning

Researchers have developed SOReL, a Bayesian model-based reinforcement learning method that enables hyperparameter selection without requiring online environment interactions.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving

Researchers have developed a retrieval-augmented generation framework to automate the creation of regulation-compliant test scenarios for autonomous driving simulations.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs

Researchers have developed a new evaluation framework called JKP to test the epistemic stability of vision-language models when subjected to repetitive adversarial questioning.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Reinforcement Learning in Switching Non-Stationary Markov Decision Processes: Algorithms and Convergence Analysis

This paper presents a mathematical framework for reinforcement learning in environments where the underlying dynamics switch between different states.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon Strategy Game Planning

The SAGA agent framework is proposed to improve long-horizon strategic planning in complex games by better managing relational structures and evolving goals.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

When Does Belief-Based Agent Memory Help? Reliability-Conditional Updating and Provenance-Capped Poisoning Defense

A new study evaluates the effectiveness of belief-based memory architectures in LLM agents using Bayesian inference to manage long-term information.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment

A new post-training method for multimodal document question answering focuses on reasoning-free alignment to reduce computational costs and improve visual grounding.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

Researchers introduced Contrastive Policy Optimization to improve reinforcement learning by using token-level disagreement to better distinguish between useful uncertainty and harmful confusion.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

From Stateless to Situated: Building a Psychological World for LLM-Based Agents

This research proposes a framework to transition LLM-based agents from stateless interactions to situated models capable of maintaining temporal continuity for psychological support.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Flow Matching in Feature Space for Stochastic World Modeling

A new flow-matching technique for feature-space world modeling addresses the trade-off between reconstruction quality and predictive accuracy in stochastic environments.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

CXRAgent: Director-Orchestrated Multi-Stage Reasoning for Chest X-Ray Interpretation

Researchers have developed a multi-stage agentic framework designed to improve the accuracy and reasoning capabilities of AI models in interpreting chest X-ray images.

Read at arxiv.org ↗
4
InfrastructurearXiv cs.AI·7d agoPrimary

An Intelligent-Cloud Edge Multimodal Interaction System for Robots

A new cloud-edge framework combines gesture detection with LLM and VLM agents to enable efficient multimodal interaction for robots with limited onboard computing power.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery

This study compares the performance of leading vision-language models against human experts in predicting building typologies from street-level imagery.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions

A new method called SciDiagramEdit uses natural language instructions to automate the complex process of editing scientific diagrams and figures.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

teLLMe Why (Ain't Nothing but a Jam): Exploratory Causal Analysis of Urban Driving Data

The teLLMe system provides a framework for performing exploratory causal analysis on observational urban traffic data derived from dashcam video.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

NORACL: Neurogenesis for Oracle-free Resource-Adaptive Continual Learning

The proposed NORACL framework introduces a neurogenesis-inspired approach to continual learning that allows models to adapt their architecture dynamically to new tasks.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Can LLMs Build a MaxSAT Solver from Papers? The CoreForge Experience

The CoreForge project demonstrates an iterative workflow where large language models are used to construct a MaxSAT solver directly from academic literature rather than existing codebases.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents

A new benchmark called Alipay-PIBench has been introduced to evaluate the performance of AI coding agents in handling complex payment integration tasks.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·10d agoPrimary

GRASP: GRanularity-Aware Search Policy for Agentic RAG

The GRASP framework improves agentic retrieval-augmented generation by optimizing when to retrieve information and controlling the granularity of the retrieved context.

Read at arxiv.org ↗
4
Model ReleasearXiv cs.AI·10d agoPrimary

MorphologyFM: A Foundation Model for Morphology-Aware Representation Learning from ECG and Pulse Oximetry Waveforms

The authors present MorphologyFM, a foundation model designed to learn clinical waveform representations while preserving the structural morphology of ECG and pulse oximetry data.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·10d agoPrimary

Playful AI in Professional Email: A Field Experiment on Tone and Recipient Engagement

A randomized field experiment evaluated how professional and playful emails rewritten by GPT-5 affected recipient engagement and behavior in workplace communications.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

RAD: Retrieval High-quality Demonstrations to Enhance Decision-making

The proposed RAD framework improves offline reinforcement learning by retrieving high-quality demonstrations to help models generalize beyond static training datasets.

Read at arxiv.org ↗
← NewerPage 148Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.