✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
7
ResearcharXiv cs.AI·4d agoPrimary

An offline approach to fNIRS-guided reinforcement learning for robot behavior

A study explores the use of functional near-infrared spectroscopy to provide brain-signal feedback for training reinforcement learning models in robotics.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Safety from Honesty in a Disinterested AI Predictor

This paper proposes a formal safety framework for an AI predictor designed to provide honest outputs by conditioning on epistemically contextualized data.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

A study into the mechanistic interpretability of LLM-as-a-judge models reveals that scoring biases can be identified and analyzed within the internal hidden states of the models.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·4d agoPrimary

Accounting for Hysteresis and Eddy Currents in Finite Element Simulations of Ferromagnetic Laminated Cores using a Recurrent Neural Network

Researchers developed a recurrent neural network approach to improve the computational efficiency of simulating electromagnetic effects in electrical machine cores.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Agentic Routing: The Harness-Native Data Flywheel

A new approach to agentic routing is proposed to optimize model selection within execution harnesses by leveraging the specialized strengths of different AI models.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Technical Report on the CVPR 2026@AdvML Workshop Challenge

A technical report details a CVPR 2026 challenge focused on testing the robustness of autonomous driving vision-language agents against adversarial multimodal attacks.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Distributed Denial of Science: How Indirect Data Poisoning of AI Systems Can Industrialize Scientific Fraud

This study explores the potential for malicious actors to automate scientific fraud by weaponizing AI systems to generate misleading research data.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·4d agoPrimary

BrainPilot: Automating Brain Discovery with Agentic Research

The authors introduce an agentic AI system designed to automate complex research workflows within the field of neuroscience.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Large Language Models in Misinformation Ecosystems: Misuse, Defense, and Vulnerability

A new study analyzes the systemic risks posed by large language models in misinformation ecosystems, proposing a framework to categorize vulnerabilities and defense strategies.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

ML in a Box: Analyzing Containerization Practices in Open Source ML Projects

This empirical study analyzes containerization practices in open-source machine learning projects to understand how iterative workflows affect build performance and container size.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation

Quantum Circuit Vision is a new cost-aware benchmark designed to evaluate how well multimodal AI agents can interpret quantum circuit diagrams and generate executable code.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Do These Violent Delights Have Violent Ends? Measuring the Post-Merge Fate of Agentic Code

This study evaluates the post-merge stability and maintenance requirements of code generated autonomously by AI agents in real-world repositories.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

JEPA for AI-Native 6G: Predictive Representations and Open Challenges

This paper discusses the application of Joint-Embedding Predictive Architecture as a self-supervised learning paradigm for AI-native 6G wireless networks.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Lifelong Representations: A Survey on Continual Self-Supervised Learning for Vision Models

This paper provides a comprehensive survey of continual self-supervised learning techniques for computer vision, highlighting connections to vision-language models.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns

This study demonstrates that applying controlled perturbations to multimodal language models can replicate the specific picture-naming error patterns observed in human stroke patients with aphasia.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models

Researchers developed MAGIC, a framework that leverages large language models to automate the generation of navigable, multi-scene 3D game environments with consistent transitions.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

From Neural Network Decisions to Training Cases: An Exact Account via Case-Based Decision Theory

This work applies case-based decision theory to map neural network predictions back to specific training examples for improved model auditing.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation

Researchers have introduced AgentFootprint, a benchmark designed to measure and evaluate the persistent storage footprint left on disk by large language model agents.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

NextFund: A Unified Performance Tracking Platform for Agentic Portfolio Management

The authors introduce NextFund, a unified platform designed to track and evaluate the performance, reasoning, and execution steps of LLM-based portfolio management agents.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Laguerre Geometry for Interpreting Large Language Models

This study proposes using Laguerre Geometry to model and analyze how concepts are structured and separated within large language models.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Co4ICF: Co-evolving Physics-Informed Surrogate and RL-based Pulse Optimizer for Inertial Confinement Fusion

The Co4ICF framework couples a physics-informed surrogate model with a reinforcement learning optimizer to prevent out-of-distribution errors in inertial confinement fusion simulations.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Dynamic Agent Skills: A Lifecycle Survey and Taxonomy of Evolving Skill Libraries

This paper presents a taxonomy and survey analyzing how reusable procedures and skill libraries for language model agents evolve over time.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

How Much Does Correctness Cost? Budgeted Placement of Strong Correctors in a Weak Multi-Agent Swarm

The paper analyzes the optimal placement and cost-efficiency of highly accurate oracle agents to guide a swarm of cheaper, unreliable agents toward consensus.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

DiffUE: Enhancing Utility-Unlearnability Trade-off of Unlearnable Examples via Diffusion Autoencoders

Researchers introduced DiffUE, a method using diffusion autoencoders to improve the effectiveness of unlearnable examples in protecting images from unauthorized AI training.

Read at arxiv.org ↗
← NewerPage 93Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.