✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
5
ResearcharXiv cs.AI·8d agoPrimary

On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage

Researchers analyzed the citation faithfulness and coverage of a four-billion parameter research agent running locally on a consumer laptop.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·8d agoPrimary

The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank

The paper introduces NameRank, a metric designed to measure how well large language models recognize specific researchers and tools within their parametric memory.

Read at arxiv.org ↗
5
FundingLeiphone 雷锋网·9d ago

独家|把芯片设计交给AI,上海AI Lab李林阳创业获数千万元首轮融资

Chinese startup Novasilicon has secured multi-million yuan in seed funding to develop AI-driven chip design solutions.

Read at leiphone.com ↗
5
ResearchSimon Willison·6d ago

Firefox in WebAssembly

Developers have successfully compiled the Firefox browser into WebAssembly, allowing it to run entirely within another web browser.

Read at simonwillison.net ↗
5
ProductData Center Dynamics·8d ago

Alphabet spin-out Verrus plans three-building data center campus outside Portland, Oregon

Verrus, an Alphabet spin-out, is planning the construction of a new grid-reactive data center campus in Oregon to support infrastructure needs.

Read at datacenterdynamics.com ↗
5
Product36Kr 36氪·9d ago

德明利:预计上半年净利润为57亿元–65亿元

Chinese storage company Demingli projects a first-half net profit of up to 6.5 billion yuan, reversing a previous loss due to surging AI-driven storage demand.

Read at 36kr.com ↗
5
ProductSCMP Tech·11d ago

In China’s electronics hub, a memory chip crisis is hitting consumers hard

The global artificial intelligence boom has driven up memory and solid-state drive prices in Shenzhen's Huaqiangbei electronics market, significantly increasing costs for PC builders.

Read at scmp.com ↗
5
ResearcharXiv cs.AI·7d agoPrimary

DIVE: Embedding Compression via Self-Limiting Gradient Updates

DIVE is a new dimensionality reduction technique for language model embeddings that uses self-limiting gradient updates to improve compression efficiency.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·7d agoPrimary

Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges

This review examines the integration of explainable AI techniques within federated learning architectures to improve transparency in privacy-preserving distributed model training.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·7d agoPrimary

Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate

A study demonstrates that the interaction protocols used in multi-agent debates significantly influence the moral reasoning and judgment outcomes of large language models.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·7d agoPrimary

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle

A new evaluation framework called STOCKTAKE aims to distinguish between perception errors and execution failures in LLM agents during long-term decision-making tasks.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·7d agoPrimary

Experience Memory Graph: One-Shot Error Correction for Agents

A new error-correction method called Experience Memory Graph aims to help LLM agents recover from failures in long-horizon tasks more efficiently than traditional reflection techniques.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·7d agoPrimary

Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution

A new auditing method uses predicate substitution to test whether large language models genuinely rely on stated premises during chain-of-thought reasoning.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·7d agoPrimary

Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code

The authors present a method for integrating compiler feedback directly into the autoregressive decoding process to improve the quality of AI-generated code.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·7d agoPrimary

Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild

Researchers have proposed a training-free approach for detecting human-object interactions in the wild by leveraging the capabilities of multimodal large language models.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·7d agoPrimary

MedDiffuseMix: Preserving Diagnostic Evidence with Saliency-Aware Diffusion Medical Image Data Augmentation

MedDiffuseMix provides a saliency-guided diffusion framework to augment medical imaging data while preserving critical diagnostic features.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·7d agoPrimary

Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science

This paper proposes a framework for networked intelligence that uses shared context graphs to facilitate collaboration among multiple AI agents in scientific research.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·7d agoPrimary

Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies

A new method for test-time learning introduces learnable adaptation policies to help language agents improve their performance through iterative interaction.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·7d agoPrimary

Set-shifting Behavioral Test for Harnessed Agents

A new benchmark based on cognitive psychology tests has been developed to evaluate how AI agents adapt when the reliability of their tools changes during operation.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·7d agoPrimary

How Far Can Root Cause Analysis Go on Real-World Telemetry Data?

A study evaluating root cause analysis in microservice failures reveals that current AI and classical methods struggle to effectively process large-scale, multimodal telemetry data.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·7d agoPrimary

When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects

Researchers analyzed nearly 3,000 GitHub projects to understand how the integration of automated bots as active participants influences the organizational structure of open-source software teams.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·7d agoPrimary

A Self-Evolving Agent for Longitudinal Personal Health Management

HealthClaw is a proposed self-evolving agent architecture designed to provide longitudinal health management by maintaining private memory of user routines and medical history.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·7d agoPrimary

When Audio Separation Hurts Zero-Shot ASR: Evaluating SAM-Audio with Whisper on Bengali and English Speech

The study evaluates the impact of audio separation preprocessing on zero-shot automatic speech recognition performance, finding that cleaner audio does not always improve transcription accuracy.

Read at arxiv.org ↗
5
ProductarXiv cs.AI·7d agoPrimary

FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents

FixItFlow is an automated system that leverages large language models to generate technical troubleshooting guides from historical cloud incident data.

Read at arxiv.org ↗
← NewerPage 129Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.