✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
2
ResearcharXiv cs.AI·12d agoPrimary

Towards Autonomous and Auditable Medical Imaging Model Development

Researchers have developed AMID, an autonomous multi-agent framework designed to streamline the development and validation of medical imaging models.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

STEC: Evidence Compression for Deep Search in Open-domain Multi-Hop QA

The STEC framework introduces evidence compression to help search-based language models resolve conflicting information across multiple search trajectories in multi-hop question answering.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

ECG-LDC: A Hardware-Efficient Low-Dimensional Computing Framework for ECG Arrhythmia Classification

The authors present ECG-LDC, a low-dimensional computing framework designed for energy-efficient and accurate arrhythmia classification on resource-constrained wearable devices.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion

The VoxENES 2026 benchmark is introduced to improve the detection of synthetic speech generated by modern LLM-based text-to-speech and voice conversion systems.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks

A new theoretical framework explains the emergence of inductive reasoning capabilities in transformer models by analyzing their learning dynamics across various synthetic tasks.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

SALT-GNN: Handling Dense Neighborhoods in Anti-Money Laundering Graphs via Statistics-Aware Attention

Researchers developed SALT-GNN, a graph neural network framework that uses statistics-aware attention to improve the detection of money laundering in dense transaction networks.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Cross-Layer Misalignment Detection in Agent Skills: A Progressive Loading-Aware Contrastive Learning Approach

Researchers have developed a contrastive learning method to identify discrepancies between the natural-language descriptions of AI agent skills and their actual execution-time behaviors.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents

A new evaluation framework called MM-ToolSandBox has been introduced to assess the performance of visually grounded AI agents across diverse tool-calling tasks.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Characterizing Large Language Model Agentic Workflows: A Study on N8n Ecosystem

Researchers conducted an empirical analysis of how large language model agents are utilized within low-code and no-code automation platforms.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

A Knowledge-Based Multi-Agent Framework for Security Control Recommendation

Researchers have proposed a security decision support system that uses a multi-agent framework to recommend security controls based on minimal user requirements.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

KGCQual: An Interpretable Framework for Evaluating the Knowledge Graph Construction Quality from Text

The KGCQual framework offers an interpretable metric to evaluate the structural and semantic quality of automatically constructed knowledge graphs.

Read at arxiv.org ↗
2
RegulationData Center Dynamics·10d ago

Italian judge blocks Inwit's efforts to stop Telecom Italia, Fastweb+Vodafone tower agreement exits

An Italian court has ruled against Inwit in a dispute regarding telecommunications tower agreements, though the company intends to appeal.

Read at datacenterdynamics.com ↗
2
ResearcharXiv cs.AI·12d agoPrimary

GES-TSP: Graph Edge Sparsification for TSP

Researchers have proposed Graph Edge Sparsification, a machine learning-based approach designed to improve the computational efficiency of solving large-scale Traveling Salesman Problems.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Faithful, Not Corrective: Message-Format Effects in Multi-Hop Agent Relays Are Tier-Dependent

Researchers investigated how message formatting affects information fidelity and generation quality when LLM agents pass information across multiple hops.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

PhenoEmbed: Self-Supervised Multispectral UAV Time-Series Embeddings for Individual Tree Crown Phenology

PhenoEmbed is a self-supervised temporal embedding model designed to track the changing characteristics of individual tree crowns using multispectral UAV time-series data.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Towards Predictive, Aligned, and Scalable Robot Learning

Researchers introduced Lumo-2, a latent world-action model designed to improve robot learning by reasoning over physical world dynamics.

Read at arxiv.org ↗
2
RegulationSCMP Tech·15d ago

Beyond Claude Code: the Chinese AI tools poised to benefit after back-door alert

A cybersecurity warning from Chinese authorities regarding a backdoor in Anthropic's Claude Code is expected to drive local developers toward domestic AI coding alternatives.

Read at scmp.com ↗
2
Product36Kr 36氪·15d ago

9点1氪丨“国产存储第一股”长鑫科技公布承销团阵容;SK海力士登陆美股,上市首日大涨近13%;OpenAI推出ChatGPT智能体

This news roundup highlights SK Hynix's successful debut on the Nasdaq, the upcoming IPO of Chinese memory manufacturer CXMT, and OpenAI's launch of ChatGPT agents.

Read at 36kr.com ↗
2
Opinion36Kr 36氪·15d ago

智谱CEO唐杰发内部信:“GLM 时刻”和万亿俱乐部之后,什么是更重要的事

Zhipu AI founder Tang Jie announced in an internal letter that the company will prioritize long-term AGI goals, autonomous agents, and safety over short-term monetization.

Read at 36kr.com ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts

The paper introduces Nested-ReFT, an efficient reinforcement learning method for fine-tuning large language models on complex reasoning tasks using off-policy rollouts.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

Researchers propose a new dialogue-based evaluation framework to more accurately assess the Theory of Mind capabilities of large language models.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

TENET: One Step Toward Test-Driven Development for Repository-Level Code Generation

The TENET framework aims to facilitate repository-level test-driven development by enabling AI agents to synthesize code based on developer-defined test specifications.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Multi-Agent Routing as Set-Valued Prediction: A WildChat Benchmark and Cost-Aware Evaluation

A new benchmark for multi-agent routing evaluates how effectively models select appropriate agents for tasks while managing execution costs.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students

A study reveals that LLMs acting as history tutors exhibit epistemic paternalism and biased refusal patterns when interacting with different student demographics.

Read at arxiv.org ↗
← NewerPage 181Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.