✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
5
ResearcharXiv cs.AI·9d agoPrimary

Remembering Distinct Items, Not Tokens: A Learnable Dirichlet-Process Cache Between State-Space Models and Attention

This paper proposes a learnable Dirichlet-process cache that stores only novel inputs to bridge the gap between state-space models and attention mechanisms.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

ActiveFly-Bench is a new benchmark designed to evaluate aerial embodied perception by linking high-level task understanding, behavior planning, and low-level control for unmanned aerial vehicles.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning

This study analyzes the internal attention dynamics of vision-language models to explain why visual grounding degrades and proposes scheduling visual relay windows to stabilize reasoning.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks

The authors introduce SWE-MERA, a dynamic benchmark designed to mitigate data contamination and improve the evaluation of LLMs on software engineering tasks.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora

Researchers released FindMyText, an open-source Python package designed to detect near-verbatim text containment within large web-crawled datasets using document fingerprinting.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

The Ramanujan Challenge For AI

Researchers have introduced a new benchmark dataset of mathematical constant formulas to evaluate the advanced mathematical reasoning capabilities of AI systems.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

Model Collapse: On Recursion, Noise, and Uncharted Machine Visions

This paper examines the phenomenon of model collapse from both engineering and creative perspectives, exploring how recursive training on AI-generated data affects model output.

Read at arxiv.org ↗
5
Model ReleasearXiv cs.AI·9d agoPrimary

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

Embodied-R1.5 is a new foundation model designed to unify embodied reasoning and physical intelligence using a large-scale dataset of 15 billion tokens.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

Knowledge Graphs Meet Graph Neural Networks: A Comprehensive Survey

This paper provides a comprehensive survey and a new two-level taxonomy of Graph Neural Network methodologies applied across knowledge graph technologies.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

Unlocking Parallelism in Autoregressive Language Models via Speculative Decoding with Progressive Tree Drafting

Researchers propose a new speculative decoding method that improves large language model inference speed by utilizing progressive tree drafting to better exploit parallel processing.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning

GrandCode is a new multi-agent reinforcement learning system that aims to achieve competitive programming performance at the level of human grandmasters.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

Measuring AI Ability to Complete Long Software Tasks

A new metric called the 50%-task-completion time horizon has been proposed to better compare AI performance against human capabilities in software development tasks.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories

Agentic-DPO is a new training framework designed to optimize AI agent policies on expert trajectories by teaching them to avoid plausible mistakes rather than just imitating sequences.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation

IdeaTrail is a new dataset designed to capture the complete, multi-stage workflow of AI agents performing scientific ideation and research tasks.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

Mako: A Self-Evolving Agentic Operating System (SE-AOS) for Autonomous Web Exploitation

The Mako project introduces a self-evolving agentic operating system capable of autonomously synthesizing and testing new security exploits.

Read at arxiv.org ↗
5
InfrastructurearXiv cs.AI·9d agoPrimary

Edge Physical AI Deployment of Vision Transformers on Heterogeneous Edge GPU Targeting Autonomous Vehicles

A new scheduling method for heterogeneous edge GPUs aims to improve the efficiency and throughput of vision transformer models in autonomous vehicle applications.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

ReflectWorld-MM: An Entity-Oriented Multi-Media Memory System for Open-Ended Video Streams

Researchers developed ReflectWorld-MM, a multimedia memory system that organizes long-term video stream data around persistent entities rather than individual frames.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

Valid $\ne$ Necessary: Diagnosing Latent Inefficiency in Chain-of-Thought

The study identifies a critical inefficiency in Chain-of-Thought prompting where models generate redundant but logically valid reasoning steps that current evaluators fail to penalize.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

A Theory of Least Autonomy in AI

Researchers propose a formal theory of least autonomy as a security principle to constrain the permissions and workflow capabilities of agentic AI systems.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

This paper outlines a deployment-focused pathway for transitioning medical AI agents from simple assistants to autonomous clinical systems.

Read at arxiv.org ↗
5
InfrastructurearXiv cs.AI·9d agoPrimary

MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference

MemDecay is a new memory management technique that optimizes LLM agent inference by using region-aware KV cache eviction based on semantic structure.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

Gefen: Optimized Stochastic Optimizer

The proposed Gefen optimizer reduces the memory footprint of large-scale model training by sharing second-moment estimates and quantizing first-moment states.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

FARS: A Fully Automated Research System Deployed at Scale

The authors introduce FARS, a fully automated system that enables AI agents to conduct research, generate hypotheses, and write manuscripts at scale.

Read at arxiv.org ↗
5
ResearcharXiv cs.AI·9d agoPrimary

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

Researchers have introduced SDABench, a new benchmark designed to evaluate the scientific data analysis capabilities of large language models across six distinct areas.

Read at arxiv.org ↗
← NewerPage 120Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.