✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
2
ResearcharXiv cs.AI·12d agoPrimary

Reverse Engineering Compliance: A Dual-Graph Verification Framework for Auditing Legacy IT Security Concepts

Researchers proposed a dual-graph verification framework to help organizations transition legacy IT security documents into machine-readable compliance formats.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

What Does It Mean to Break a Distillation Defense?

Researchers are analyzing the effectiveness of output perturbation defenses against distillation attacks on black-box language models.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability

A new training-free method for diffusion and flow matching models accelerates generation by utilizing x-prediction to reduce the number of required neural function evaluations.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Cross-Subject Modeling for Widefield Calcium Imaging via Atlas-Aligned Spatiotemporal Tokenization

The paper proposes a spatiotemporal tokenization method to enable cross-subject and multi-session modeling of widefield calcium imaging data in neuroscience.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces

The Filtered Reasoning Score is a new evaluation metric designed to assess the quality of reasoning in large language models by focusing on their most confident outputs.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·13d agoPrimary

EHR-MPC: Inference-Time Control for Sepsis Treatment with Generative Patient Digital Twins

EHR-MPC is a new framework that uses generative electronic health record models as patient digital twins to optimize sepsis treatment strategies dynamically during inference.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Tool-Adaptive LLM Reranker

A new tool-adaptive reranker has been proposed to improve the accuracy and efficiency of information retrieval by selectively invoking external tools.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services

A new benchmark called BackendForge is introduced to evaluate the ability of agentic LLMs to generate and deploy functional code within backend service environments.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Faithful, Not Corrective: Message-Format Effects in Multi-Hop Agent Relays Are Tier-Dependent

Researchers investigated how message formatting affects information fidelity and generation quality when LLM agents pass information across multiple hops.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation

Tool-MCoT introduces a small language model optimized for content safety moderation by leveraging external tools to reduce computational costs and latency.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Towards Autonomous and Auditable Medical Imaging Model Development

Researchers have developed AMID, an autonomous multi-agent framework designed to streamline the development and validation of medical imaging models.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

From ambiguous utterances to governed reuse classes: canonicalization, quotient invariance, and conditional decidability

Researchers propose a mathematical framework to replace similarity heuristics in semantic caching with governed conversational demand classes.

Read at arxiv.org ↗
2
OpinionTechCrunch AI·11d ago

Anthropic’s newest ad is creeping people out

Anthropic's latest advertisement has sparked strong emotional reactions and discomfort among viewers.

Read at techcrunch.com ↗
2
ResearcharXiv cs.AI·12d agoPrimary

An Autonomous Scientific Knowledge Generation Framework for AI-Driven Scientific Discovery

Researchers have proposed an autonomous framework designed to extract structured scientific knowledge from unstructured literature to accelerate AI-driven materials research.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging

This paper introduces TS-Mask VLA, a vision-language-action model that utilizes two-dimensional temporal-spatial masking to improve action generation for embodied agents.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation

AMT-X is a new multi-turn red-teaming framework designed to improve LLM safety evaluations by addressing the limitations of single-turn attack datasets and scoring methods.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning

The study demonstrates that structured, object-centric slot representations can improve robotic manipulation policies without requiring increased model capacity.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning

The paper explores how the reliability of value estimation affects policy optimization in offline-to-online reinforcement learning for robotic manipulation.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Comparative Analysis of GAT and BERT for Human-Like Playtesting

Researchers compared Graph Attention Networks and BERT models to evaluate their effectiveness in mimicking human player behavior and strategies in puzzle games.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Researchers introduced Engram, a module that adds conditional memory to Transformers to enable efficient knowledge lookup as an alternative to standard Mixture-of-Experts scaling.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

A Knowledge-Based Multi-Agent Framework for Security Control Recommendation

Researchers have proposed a security decision support system that uses a multi-agent framework to recommend security controls based on minimal user requirements.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·13d agoPrimary

OpenProver: Agentic and Interactive Theorem Proving with Lean 4

The paper presents OpenProver, an open-source, agentic system that integrates a Planner-Worker-Verifier architecture for automated theorem proving using Lean 4.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Toward Autonomous Soft Robotic Endovascular Navigation via Imitation Learning

Researchers are utilizing imitation learning to develop autonomous navigation systems for soft robotic guidewires used in endovascular surgery.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

PREF-Gate: Provenance-Constrained Relational Evidence Fusion with Validation-Gated Selection for Graph Fraud Detection

Researchers have developed PREF-Gate, a decision framework designed to prevent invalid neighborhood risk data from compromising graph-based fraud detection.

Read at arxiv.org ↗
← NewerPage 180Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.