✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
6
ProductThe Next Platform·8d ago

The Aspirations Of HPE And Dell In The Quantum-Classical HPC Datacenter

Hewlett Packard Enterprise and Dell are outlining their strategic goals for integrating quantum and classical high-performance computing within modern datacenters.

Read at nextplatform.com ↗
6
ProductData Center Dynamics·7d ago

Valfortec plans €300m data center campus in Alicante, Spain

Valfortec is planning to develop a 60MW data center campus in Alicante, Spain, with an investment of 300 million euros.

Read at datacenterdynamics.com ↗
6
Funding36Kr 36氪·8d ago

长飞光纤等在武汉成立智能创投基金,出资额5亿

Yangtze Optical Fibre and Cable has partnered with other entities to establish a 500 million yuan venture capital fund in Wuhan focused on smart technologies.

Read at 36kr.com ↗
6
InfrastructureAWS ML Blog·4d ago

How Smartsheet built a remote MCP server on AWS

Smartsheet has detailed the architecture and AWS-based infrastructure used to deploy a remote Model Context Protocol server.

Read at aws.amazon.com ↗
6
ResearcharXiv cs.AI·8d agoPrimary

Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents

Researchers identify a failure mode where LLM agents successfully complete tasks while violating safety policies, and propose deterministic gates as a mitigation strategy.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation

A new benchmark, SPQR, evaluates the stability of safety alignment in text-to-image models when subjected to common downstream fine-tuning techniques.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

To Answer or to Abstain: Mitigating Search-Agent Hallucinations via Abstention-Aware Reinforcement Learning

Researchers have developed an abstention-aware reinforcement learning approach to reduce hallucinations in search-augmented LLMs by penalizing incorrect answers when retrieval fails.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming

Researchers propose a new red-teaming framework that uses autonomous agents to identify security vulnerabilities in production LLM agent workflows.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

ReflectWorld-MM: An Entity-Oriented Multi-Media Memory System for Open-Ended Video Streams

Researchers developed ReflectWorld-MM, a multimedia memory system that organizes long-term video stream data around persistent entities rather than individual frames.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning

GrandCode is a new multi-agent reinforcement learning system that aims to achieve competitive programming performance at the level of human grandmasters.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

Gefen: Optimized Stochastic Optimizer

The proposed Gefen optimizer reduces the memory footprint of large-scale model training by sharing second-moment estimates and quantizing first-moment states.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning

This study analyzes the internal attention dynamics of vision-language models to explain why visual grounding degrades and proposes scheduling visual relay windows to stabilize reasoning.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

ActiveFly-Bench is a new benchmark designed to evaluate aerial embodied perception by linking high-level task understanding, behavior planning, and low-level control for unmanned aerial vehicles.

Read at arxiv.org ↗
6
Model ReleasearXiv cs.AI·8d agoPrimary

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

Embodied-R1.5 is a new foundation model designed to unify embodied reasoning and physical intelligence using a large-scale dataset of 15 billion tokens.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

Valid $\ne$ Necessary: Diagnosing Latent Inefficiency in Chain-of-Thought

The study identifies a critical inefficiency in Chain-of-Thought prompting where models generate redundant but logically valid reasoning steps that current evaluators fail to penalize.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

Remembering Distinct Items, Not Tokens: A Learnable Dirichlet-Process Cache Between State-Space Models and Attention

This paper proposes a learnable Dirichlet-process cache that stores only novel inputs to bridge the gap between state-space models and attention mechanisms.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

This paper outlines a deployment-focused pathway for transitioning medical AI agents from simple assistants to autonomous clinical systems.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation

IdeaTrail is a new dataset designed to capture the complete, multi-stage workflow of AI agents performing scientific ideation and research tasks.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

Researchers have introduced SDABench, a new benchmark designed to evaluate the scientific data analysis capabilities of large language models across six distinct areas.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks

The authors introduce SWE-MERA, a dynamic benchmark designed to mitigate data contamination and improve the evaluation of LLMs on software engineering tasks.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

Knowledge Graphs Meet Graph Neural Networks: A Comprehensive Survey

This paper provides a comprehensive survey and a new two-level taxonomy of Graph Neural Network methodologies applied across knowledge graph technologies.

Read at arxiv.org ↗
6
InfrastructurearXiv cs.AI·8d agoPrimary

MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference

MemDecay is a new memory management technique that optimizes LLM agent inference by using region-aware KV cache eviction based on semantic structure.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

Model Collapse: On Recursion, Noise, and Uncharted Machine Visions

This paper examines the phenomenon of model collapse from both engineering and creative perspectives, exploring how recursive training on AI-generated data affects model output.

Read at arxiv.org ↗
6
ResearcharXiv cs.AI·8d agoPrimary

Unlocking Parallelism in Autoregressive Language Models via Speculative Decoding with Progressive Tree Drafting

Researchers propose a new speculative decoding method that improves large language model inference speed by utilizing progressive tree drafting to better exploit parallel processing.

Read at arxiv.org ↗
← NewerPage 106Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.