✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
5
ResearcharXiv cs.AI·8d agoPrimary

Traffic-Aware Randomized Smoothing for LLM-Based Network Intrusion Detection

A new defense mechanism called Traffic-Aware Randomized Smoothing has been proposed to improve the robustness of LLM-based network intrusion detection systems against traffic manipulation.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

A Self-Evolving Agent for Longitudinal Personal Health Management

HealthClaw is a proposed self-evolving agent architecture designed to provide longitudinal health management by maintaining private memory of user routines and medical history.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies

A new method for test-time learning introduces learnable adaptation policies to help language agents improve their performance through iterative interaction.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System

A study evaluates the scalability and classroom performance of an AI tutoring agent that integrates retrieval-augmented generation with structured knowledge models.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level

The ExTernD method introduces a post-training factorization technique for large language models that utilizes expanded-rank ternary decomposition to improve quantization efficiency and accuracy.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

IMMNet: Hybrid Fusion of Model-based and Data-driven Approaches for Maneuvering Target Tracking

IMMNet integrates traditional model-based tracking algorithms with neural components to improve the accuracy and interpretability of maneuvering target tracking in 3D environments.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

MASPRM: Multi-Agent System Process Reward Model

The Multi-Agent System Process Reward Model provides a method to evaluate and optimize message sequences between agents during inference-time search.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

Reassessing Muon for Matrix Factorization

This paper provides a theoretical analysis of the Muon optimizer to clarify the mechanisms behind its performance advantages in large-scale deep learning.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild

Researchers have proposed a training-free approach for detecting human-object interactions in the wild by leveraging the capabilities of multimodal large language models.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

Faithful Autoformalization of Natural Language Assertions

The authors introduced Monty, a framework designed to improve the accuracy of autoformalization by converting natural language specifications into executable software assertions.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code

The authors present a method for integrating compiler feedback directly into the autoregressive decoding process to improve the quality of AI-generated code.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

A Causal Model of Theory of Mind in Conflict for Artificial Intelligence

This paper proposes a causal framework to determine when artificial intelligence systems should engage in theory of mind processes during conflict scenarios.

Read at arxiv.org ↗
5
ProductarXiv cs.AI·8d agoPrimary

FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents

FixItFlow is an automated system that leverages large language models to generate technical troubleshooting guides from historical cloud incident data.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

Beyond Backbone Backpropagation: A Decoupled Strategy for Efficient Transfer Learning

A new decoupled training strategy for transfer learning aims to reduce computational and energy costs by optimizing classifier heads separately from feature extraction layers.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

The SIGReg Objective as Variational Free Energy: A Theoretical Active-Inference Account of JEPA World Models

The paper provides a theoretical framework linking Joint-Embedding Predictive Architectures to active inference principles through variational free energy.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

The Hitchhiker's Guide to Monoculture

An analysis of Kaggle contest submissions examines whether the use of AI coding assistants is leading to increased homogenization of software development outputs.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

When Audio Separation Hurts Zero-Shot ASR: Evaluating SAM-Audio with Whisper on Bengali and English Speech

The study evaluates the impact of audio separation preprocessing on zero-shot automatic speech recognition performance, finding that cleaner audio does not always improve transcription accuracy.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

STKAN: Kolmogorov-Arnold Networks for Spatio-Temporal Forecasting

The authors introduce STKAN, a spatio-temporal forecasting architecture that leverages Kolmogorov-Arnold Networks to better model complex traffic data.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy

The SteinGate framework introduces a new safety certification method for reinforcement learning to better mitigate rare but catastrophic risks.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

Experience Memory Graph: One-Shot Error Correction for Agents

A new error-correction method called Experience Memory Graph aims to help LLM agents recover from failures in long-horizon tasks more efficiently than traditional reflection techniques.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle

A new evaluation framework called STOCKTAKE aims to distinguish between perception errors and execution failures in LLM agents during long-term decision-making tasks.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges

This review examines the integration of explainable AI techniques within federated learning architectures to improve transparency in privacy-preserving distributed model training.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate

A study demonstrates that the interaction protocols used in multi-agent debates significantly influence the moral reasoning and judgment outcomes of large language models.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution

A new auditing method uses predicate substitution to test whether large language models genuinely rely on stated premises during chain-of-thought reasoning.

Read at arxiv.org ↗
← NewerPage 136Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.