✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
4
ResearcharXiv cs.AI·10d agoPrimary

GAE: Graph-Augmented Evolution for Scientific Discovery via Reinforcement Optimization

The Graph-Augmented Evolution framework combines large language models with reinforcement optimization to improve automated scientific discovery by addressing limitations in evolutionary program search.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·10d agoPrimary

Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?

The authors introduce Who and When Pro, a large-scale benchmark designed to evaluate how effectively large language models can attribute failures in agentic workflows.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·10d agoPrimary

Looped State-Space Language Models with Adaptive Exit-State Selection

This study explores whether looped architectures and adaptive exit strategies can improve the computational depth and reasoning of state-space language models.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·10d agoPrimary

Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices

Researchers developed an automated tensor scheduling system that optimizes hybrid CPU-GPU memory offloading to run large language models more efficiently on consumer devices.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·10d agoPrimary

SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL

The ScaleCUA framework aims to improve computer use agents by combining verifiable task synthesis with efficient online reinforcement learning.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·10d agoPrimary

Unified Backbone Refinement for Diffusion Models via Internal-Latent Analysis

Researchers introduce DUNE, a training-free refinement method that reduces artifacts in diffusion models by analyzing and stabilizing early-stage latent fluctuations.

Read at arxiv.org ↗
4
FundingLeiphone 雷锋网·10d ago

独家|把芯片设计交给AI,上海AI Lab李林阳创业获数千万元首轮融资

Chinese startup Novasilicon has secured multi-million yuan in seed funding to develop AI-driven chip design solutions.

Read at leiphone.com ↗
4
ResearchSimon Willison·7d ago

Firefox in WebAssembly

Developers have successfully compiled the Firefox browser into WebAssembly, allowing it to run entirely within another web browser.

Read at simonwillison.net ↗
4
ResearcharXiv cs.AI·9d agoPrimary

The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank

The paper introduces NameRank, a metric designed to measure how well large language models recognize specific researchers and tools within their parametric memory.

Read at arxiv.org ↗
4
OpinionarXiv cs.AI·9d agoPrimary

Optimization Is Not All You Need

This paper critiques the prevailing optimization culture in AI alignment, arguing that evaluating models solely on predefined, measurable axes fails to capture true value.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·9d agoPrimary

On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage

Researchers analyzed the citation faithfulness and coverage of a four-billion parameter research agent running locally on a consumer laptop.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·9d agoPrimary

Internet of Agentic Things: Networked AI Agents for Closed-Loop IoT Orchestration

The paper introduces the Internet of Agentic Things, an architectural framework that connects autonomous AI agents across cloud, edge, and physical IoT layers.

Read at arxiv.org ↗
4
ProductData Center Dynamics·9d ago

Alphabet spin-out Verrus plans three-building data center campus outside Portland, Oregon

Verrus, an Alphabet spin-out, is planning the construction of a new grid-reactive data center campus in Oregon to support infrastructure needs.

Read at datacenterdynamics.com ↗
4
Product36Kr 36氪·10d ago

德明利:预计上半年净利润为57亿元–65亿元

Chinese storage company Demingli projects a first-half net profit of up to 6.5 billion yuan, reversing a previous loss due to surging AI-driven storage demand.

Read at 36kr.com ↗
4
ProductSCMP Tech·12d ago

In China’s electronics hub, a memory chip crisis is hitting consumers hard

The global artificial intelligence boom has driven up memory and solid-state drive prices in Shenzhen's Huaqiangbei electronics market, significantly increasing costs for PC builders.

Read at scmp.com ↗
4
ProductarXiv cs.AI·8d agoPrimary

FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents

FixItFlow is an automated system that leverages large language models to generate technical troubleshooting guides from historical cloud incident data.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·8d agoPrimary

Beyond Backbone Backpropagation: A Decoupled Strategy for Efficient Transfer Learning

A new decoupled training strategy for transfer learning aims to reduce computational and energy costs by optimizing classifier heads separately from feature extraction layers.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·8d agoPrimary

The Hitchhiker's Guide to Monoculture

An analysis of Kaggle contest submissions examines whether the use of AI coding assistants is leading to increased homogenization of software development outputs.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·8d agoPrimary

STKAN: Kolmogorov-Arnold Networks for Spatio-Temporal Forecasting

The authors introduce STKAN, a spatio-temporal forecasting architecture that leverages Kolmogorov-Arnold Networks to better model complex traffic data.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·8d agoPrimary

SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy

The SteinGate framework introduces a new safety certification method for reinforcement learning to better mitigate rare but catastrophic risks.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·8d agoPrimary

Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies

A new method for test-time learning introduces learnable adaptation policies to help language agents improve their performance through iterative interaction.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·8d agoPrimary

Reassessing Muon for Matrix Factorization

This paper provides a theoretical analysis of the Muon optimizer to clarify the mechanisms behind its performance advantages in large-scale deep learning.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·8d agoPrimary

Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science

This paper proposes a framework for networked intelligence that uses shared context graphs to facilitate collaboration among multiple AI agents in scientific research.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·8d agoPrimary

Set-shifting Behavioral Test for Harnessed Agents

A new benchmark based on cognitive psychology tests has been developed to evaluate how AI agents adapt when the reliability of their tools changes during operation.

Read at arxiv.org ↗
← NewerPage 140Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.