✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
2
ResearcharXiv cs.AI·12d agoPrimary

STAMP: Provenance-Guided Credit Assignment for Deep Search Agents

The proposed STAMP framework addresses the reward-credit mismatch in deep-search agents by using a reference-based verifier to evaluate whether cited documents support specific claims.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground Truth

This study introduces RouteCast, an evaluation framework designed to assess model-generated strategic routes when ground truth feedback is delayed or private.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·13d agoPrimary

ARCANA: A Reflective Multi-Agent Program Synthesis Framework for ARC-AGI-2 Reasoning

The authors present ARCANA, a collaborative multi-agent framework designed to solve ARC-AGI-2 reasoning tasks under strict hardware and test-time constraints.

Read at arxiv.org ↗
2
FundingQbitAI 量子位·14d ago

近百名玩家涌入具身数据:一年融资44.7亿,谁能真靠“卖数据”赚钱?

The domestic embodied AI data industry has seen rapid growth, with 15 independent data service providers securing around 4.47 billion yuan in funding over the past year.

Read at qbitai.com ↗
2
Opinion36Kr 36氪·13d ago

中信建投:计算机板块半年报预告陆续披露,AI算力硬件持续高景气

CSC Financial reports that semi-annual earnings previews confirm high demand for AI servers and intelligent computing infrastructure, while AI software applications are beginning to show revenue recovery.

Read at 36kr.com ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders

This paper introduces new sparsity regularizers for Top-k sparse autoencoders to improve the interpretability of features in vision foundation models.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Self-Regulated Reading with AI Support: An Eight-Week Study with Students

A longitudinal study of college students reveals how AI chatbot interactions influence cognitive engagement and reading habits during academic tasks.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

AgentLens is a new benchmark designed to evaluate coding agents by assessing the quality of their entire interaction trajectory rather than just final task outcomes.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Efficiently Learning Branching Networks for Multitask Algorithmic Reasoning

This study presents a method for training branching neural networks to perform multiple algorithmic reasoning tasks simultaneously.

Read at arxiv.org ↗
2
OpinionarXiv cs.AI·11d agoPrimary

Agentic Service-Oriented Computing: A Manifesto for the Next Frontier of Service-Oriented Computing

This manifesto outlines how LLM-powered autonomous agents are transforming software systems and proposes integrating them with service-oriented computing principles.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents

A novel reinforcement learning technique called Sibling-Guided Credit Distillation improves how agents learn to use tools by better attributing rewards to specific actions.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

When Does Personality Composition Matter for Multi-Agent LLM Teams?

Researchers investigated how personality-based prompting affects the task performance of multi-agent LLM teams, finding that communication styles significantly influence collaborative outcomes.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Learning When to Trust in Contextual Social Bandits

A new study identifies a failure mode called contextual sycophancy, where reinforcement learning agents receive biased feedback from evaluators in specific critical scenarios.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

TerraZero provides a procedural simulation environment and training stack to support the development of autonomous driving agents through large-scale self-play.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast

A study investigates how anatomical and contrast factors in brain MRI scans contribute to demographic predictability, highlighting potential biases in medical AI.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes

Researchers have proposed a reinforcement learning framework for Parametrized Action Markov Decision Processes that integrates incomplete domain knowledge and gradient guidance to improve sample efficiency.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes

The authors developed LakeQuest, a new benchmark designed to evaluate question-answering systems across complex, heterogeneous enterprise data lakes.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

The One-Word Census: Answer-Choice Conformity Across 44 Language Models

An analysis of 44 language models reveals a strong tendency toward convergent behavior when prompted to select single words from open-ended categories.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation

The proposed CARE-PPO framework integrates uncertainty estimation into reinforcement learning fine-tuning to reduce hallucinations and improve confidence in language-based quantitative predictions.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

TRAIL: A Platform for Configurable Human--AI Teaming Experiments

The Team Research and AI Integration Lab (TRAIL) platform provides a new infrastructure for conducting reproducible experiments on human-AI collaboration.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Sparse Inter-Layer Dependencies of Transformer FFN Neurons

A new training-free attribution method is introduced to analyze the sparse dependencies between neurons in Transformer feedforward network blocks.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study

The paper explores ontology-amplified distillation and contextuality auditing to adapt sovereign enterprise language models for regulated financial institutions.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning

A new framework combines elastic regularization and synthetic replay to mitigate catastrophic forgetting during the federated fine-tuning of multimodal models.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Graph Feedback Controls Consensus and Clique Formation in Open-Weight Language-Model Populations

Researchers investigated how interaction graphs and feedback mechanisms influence convention formation and consensus in populations of open-weight language models.

Read at arxiv.org ↗
← NewerPage 175Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.