✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
7
ResearcharXiv cs.AI·7d agoPrimary

Towards Predictive, Aligned, and Scalable Robot Learning

Researchers introduced Lumo-2, a latent world-action model designed to improve robot learning by reasoning over physical world dynamics.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation

AMT-X is a new multi-turn red-teaming framework designed to improve LLM safety evaluations by addressing the limitations of single-turn attack datasets and scoring methods.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

VIA: Visual Interface Agent for Robot Control

The VIA framework proposes a new approach to robot control by leveraging foundation models for visual reasoning and planning without requiring extensive fine-tuning on robot-specific data.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services

A new benchmark called BackendForge is introduced to evaluate the ability of agentic LLMs to generate and deploy functional code within backend service environments.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students

A study reveals that LLMs acting as history tutors exhibit epistemic paternalism and biased refusal patterns when interacting with different student demographics.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

Researchers propose a new dialogue-based evaluation framework to more accurately assess the Theory of Mind capabilities of large language models.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study

Researchers conducted a qualitative study of twenty practitioners to understand the practical methods and challenges involved in developing software engineering agents.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

PromptGraph: Graph-Guided Prompt Sanitization for Balancing Privacy and Utility in LLM Inference

PromptGraph is a proposed technique that uses graph modeling to sanitize LLM prompts, aiming to protect privacy by accounting for contextual relationships between data points.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Answer-Conditioned Chain-of-Thought Distillation for Few-Shot Industrial Vision with Small VLMs

A new distillation technique allows small vision-language models to be adapted for industrial visual inspection tasks using minimal labeled data.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Tool-Adaptive LLM Reranker

A new tool-adaptive reranker has been proposed to improve the accuracy and efficiency of information retrieval by selectively invoking external tools.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Towards Autonomous and Auditable Medical Imaging Model Development

Researchers have developed AMID, an autonomous multi-agent framework designed to streamline the development and validation of medical imaging models.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

PhenoEmbed: Self-Supervised Multispectral UAV Time-Series Embeddings for Individual Tree Crown Phenology

PhenoEmbed is a self-supervised temporal embedding model designed to track the changing characteristics of individual tree crowns using multispectral UAV time-series data.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

FlowPainter: Inpainting Optical Flow via Confidence-Guided Completion

Researchers have introduced FlowPainter, a confidence-guided completion method designed to improve optical flow inpainting in regions with large displacements and complex motion.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

A Knowledge-Based Multi-Agent Framework for Security Control Recommendation

Researchers have proposed a security decision support system that uses a multi-agent framework to recommend security controls based on minimal user requirements.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning

The paper explores how the reliability of value estimation affects policy optimization in offline-to-online reinforcement learning for robotic manipulation.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning

The study demonstrates that structured, object-centric slot representations can improve robotic manipulation policies without requiring increased model capacity.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging

This paper introduces TS-Mask VLA, a vision-language-action model that utilizes two-dimensional temporal-spatial masking to improve action generation for embodied agents.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

An Autonomous Scientific Knowledge Generation Framework for AI-Driven Scientific Discovery

Researchers have proposed an autonomous framework designed to extract structured scientific knowledge from unstructured literature to accelerate AI-driven materials research.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Cross-Subject Modeling for Widefield Calcium Imaging via Atlas-Aligned Spatiotemporal Tokenization

The paper proposes a spatiotemporal tokenization method to enable cross-subject and multi-session modeling of widefield calcium imaging data in neuroscience.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

The Universal Language of CSI:Unifying Wireless Sensing Across Devices and Environments

This paper proposes a unified framework to bridge the heterogeneity gap in WiFi sensing based on Channel State Information across different devices and environments.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Reverse Engineering Compliance: A Dual-Graph Verification Framework for Auditing Legacy IT Security Concepts

Researchers proposed a dual-graph verification framework to help organizations transition legacy IT security documents into machine-readable compliance formats.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Comparative Analysis of GAT and BERT for Human-Like Playtesting

Researchers compared Graph Attention Networks and BERT models to evaluate their effectiveness in mimicking human player behavior and strategies in puzzle games.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Knowledge-Constrained Shape Optimization with a Mixture-of-Experts Neural Operator for High-Confidence Design

The paper introduces a mixture-of-experts neural operator framework to automate and improve the reliability of aerodynamic shape optimization in engineering.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

PREF-Gate: Provenance-Constrained Relational Evidence Fusion with Validation-Gated Selection for Graph Fraud Detection

Researchers have developed PREF-Gate, a decision framework designed to prevent invalid neighborhood risk data from compromising graph-based fraud detection.

Read at arxiv.org ↗
← NewerPage 90Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.