✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
7
ResearcharXiv cs.AI·7d agoPrimary

Inclusive Federated Learning Through Compliance-Weighted Noise Allocation in Healthcare AI

Researchers proposed a compliance-aware federated learning framework that adjusts differential privacy noise to accommodate varying institutional data standards and resource levels.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning

Researchers have developed a variation-aware entropy scheduling method to improve reinforcement learning performance in environments subject to drift.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

TADPO: Reinforcement Learning Goes Off-road

Researchers have introduced TADPO, a reinforcement learning approach designed to improve autonomous vehicle navigation in complex, unmapped off-road environments.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Comparing Semantic Navigation in Humans and Large Language Models using Natural Language Processing

A comparative study using NLP metrics reveals similarities and differences in how humans and leading large language models navigate conceptual spaces during semantic memory retrieval.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs

The authors investigate how the training duration of individual domain experts influences the performance of merged large language models.

Read at arxiv.org ↗
7
ProductData Center Dynamics·6d ago

Plug Power sells Texas site to Stream Data Centers

Plug Power has sold a Texas industrial site to Stream Data Centers, which intends to repurpose the land for data center infrastructure.

Read at datacenterdynamics.com ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Visual Access Boundaries in Vision-Language Model Reasoning

Researchers have introduced Visual Access Sweep, a causal intervention method to study how Vision-Language Models utilize image tokens during long Chain-of-Thought reasoning.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Cost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems

Researchers introduced Cost-Governed RAG, an architecture designed to attribute both retrieval and generation costs to individual tenants in multi-user language model systems.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

PM-Bench: Evaluating Prospective Memory in LLM Agents

Researchers introduced PM-Bench, a text-based benchmark designed to evaluate the prospective memory capabilities of large language model agents.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Evidence-Grounded AI for Musculoskeletal Care

This work discusses the application of evidence-grounded AI systems to support the longitudinal management and rehabilitation of musculoskeletal diseases.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan Evaluation

This paper investigates a vulnerability in large language model plan evaluators where strategic plans are rewarded for omitting explicit details.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model

Researchers proposed a method for automatic speech recognition using a discrete diffusion language model to transcribe audio in parallel rather than through traditional autoregressive decoding.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

I'm Sorry, but I Can't Help with Braille: Revealing Accessibility Failures in State-of-the-Art LLMs

Research indicates that current state-of-the-art LLMs struggle significantly with bidirectional Korean-Braille translation, highlighting gaps in accessibility-focused capabilities.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

CARE-LoRA: Compressed Activation REconstruction for Memory-Efficient LoRA

The authors introduce a method called CARE-LoRA designed to reduce memory usage during the fine-tuning of large pre-trained models.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Signal-Guided Optimization for Machine Unlearning

A new signal-guided optimization technique aims to improve machine unlearning by addressing the varying memorization strengths of individual training samples.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

AutoTrace: From Patches to Triggers via Agentic Interprocedural Exploration

The authors introduce an agentic approach for localizing vulnerability triggers in code by performing interprocedural causal reasoning.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Sparse Autoencoders for Interpretable Out-of-Distribution Detection

This study introduces a method using sparse autoencoders to better detect out-of-distribution data by analyzing intermediate layers of neural networks.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Self-Consistent Flow: Unifying Velocity and Endpoint Prediction for Rectified Flow Models

Researchers have developed a unified approach for rectified flow models that combines velocity and endpoint prediction to improve generative model training.

Read at arxiv.org ↗
7
ProductarXiv cs.AI·7d agoPrimary

Fin-Analyst at FinMMEval 2026 Task 3: A Live Hybrid Trading Agent with LLM Specialists and Rule-Based Signals

Fin-Analyst is a hybrid trading agent that utilizes a multi-specialist LLM pipeline to process diverse financial data sources for automated equity market analysis.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs

Researchers introduced a new method using execution-based semantic interaction graphs to better quantify uncertainty in code-generating large language models.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Evaluating Health Misinformation in Low-Resource Languages: Integrating Small Language Models with a Culturally-Sensitive Responsible NLP Framework (Bangla as a Case Study)

This research proposes a framework for detecting health misinformation in low-resource languages by combining small language models with culturally sensitive NLP techniques.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric

Researchers propose a vendor-neutral metric to evaluate the reconstructability and validity of AI agent safety testing results.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs

The EG-VAR architecture uses the Lean 4 kernel to verify agentic reasoning and tool usage, aiming to reduce hallucinations in empirical inference.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference

Researchers have developed a method to improve the real-time control of robots by optimizing the asynchronous inference of vision-language-action models on edge hardware.

Read at arxiv.org ↗
← NewerPage 101Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.