✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
2
ResearcharXiv cs.AI·11d agoPrimary

Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution

A new study explores the need for AI agents to develop task-aware execution capabilities to better estimate the complexity of workflows and avoid inefficient resource usage.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Solution of the Hempel's statistical ambiguity problem and Causal AI

The paper addresses Carl Hempel's statistical ambiguity problem in inductive inference by applying concepts from causal AI.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

AgentLens is a new benchmark designed to evaluate coding agents by assessing the quality of their entire interaction trajectory rather than just final task outcomes.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

PFAdapter: Hierarchical LoRA Decomposition for Personalized Federated MLLMs

The authors present a hierarchical LoRA decomposition technique designed to facilitate personalized federated fine-tuning of multimodal large language models on edge devices.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Good Benchmarks

This paper defines the characteristics of high-quality benchmark tasks, emphasizing correctness, solvability, verifiability, and alignment with real-world practitioner problems.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Research Novelty in Information Systems Journals After ChatGPT: Differences Across Institutional Language Contexts

An analysis of academic publications indicates that the adoption of large language models has influenced the novelty of research output in information systems journals.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

The One-Word Census: Answer-Choice Conformity Across 44 Language Models

An analysis of 44 language models reveals a strong tendency toward convergent behavior when prompted to select single words from open-ended categories.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

AAAI-26 Dual Submissions: Novel Challenges

AAAI-26 organizers are addressing the rising issue of dual submissions in AI research to maintain the integrity of the peer-review process.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Sparse Portrait Images

FFAvatar uses a Transformer-based 3D Gaussian approach to enable the rapid, incremental construction of animatable 4D head avatars from sparse portrait images.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

TRAIL: A Platform for Configurable Human--AI Teaming Experiments

The Team Research and AI Integration Lab (TRAIL) platform provides a new infrastructure for conducting reproducible experiments on human-AI collaboration.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Accepted Prefixes Are Not All You Need: A Negative Result on PEFT-Based Block-Diffusion Drafting

The study presents a negative result regarding the effectiveness of using parameter-efficient fine-tuning adapters as block-diffusion drafters for speculative decoding.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Self-Regulated Reading with AI Support: An Eight-Week Study with Students

A longitudinal study of college students reveals how AI chatbot interactions influence cognitive engagement and reading habits during academic tasks.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes

Researchers have proposed a reinforcement learning framework for Parametrized Action Markov Decision Processes that integrates incomplete domain knowledge and gradient guidance to improve sample efficiency.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes

The authors developed LakeQuest, a new benchmark designed to evaluate question-answering systems across complex, heterogeneous enterprise data lakes.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Sparse Inter-Layer Dependencies of Transformer FFN Neurons

A new training-free attribution method is introduced to analyze the sparse dependencies between neurons in Transformer feedforward network blocks.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation

The proposed CARE-PPO framework integrates uncertainty estimation into reinforcement learning fine-tuning to reduce hallucinations and improve confidence in language-based quantitative predictions.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning

A new framework combines elastic regularization and synthetic replay to mitigate catastrophic forgetting during the federated fine-tuning of multimodal models.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Efficiently Learning Branching Networks for Multitask Algorithmic Reasoning

This study presents a method for training branching neural networks to perform multiple algorithmic reasoning tasks simultaneously.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Gene Expression-Informed Jointly Controlled Generative Modeling for Precision Molecular Design

The authors introduce JoPMol, a generative modeling framework that integrates gene expression data to assist in personalized drug discovery and molecular design.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Evaluating Reliability in Machine Learning Models for Early Chronic Kidney Disease Prediction: A Systematic Review of Data Leakage and Predictor Stability

A systematic review highlights significant methodological inconsistencies and data leakage issues in machine learning models used for chronic kidney disease detection.

Read at arxiv.org ↗
2
Regulation36Kr 36氪·11d ago

国家统计局:经济新动能不断成长壮大,日益挑起中国经济的大梁

China's National Bureau of Statistics reported that high-tech manufacturing and AI-related industries, such as integrated circuits, experienced significant growth in the first half of the year.

Read at 36kr.com ↗
2
Product36Kr 36氪·14d ago

三星电子计划将龙仁首座芯片工厂的投产时间提前至2029年

Samsung Electronics plans to accelerate the production timeline of its first semiconductor factory in the Yongin chip cluster to 2029 to meet the rapidly growing global demand for AI chips.

Read at 36kr.com ↗
2
ProductQbitAI 量子位·8d ago

B站成WAIC官方AI科技视频平台,月均超1.9亿用户消费AI内容

Bilibili has been designated as the official video platform for the World Artificial Intelligence Conference, reporting high user engagement with AI-related content.

Read at qbitai.com ↗
2
Regulation36Kr 36氪·13d ago

民政部等14部门联合发文,推动我国康复辅具产业扩能提质

Fourteen Chinese government departments have jointly launched a three-year action plan to boost the rehabilitation assistive devices industry, targeting advancements in brain-computer interfaces and rehabilitation robotics.

Read at 36kr.com ↗
← NewerPage 176Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.