✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
2
ResearcharXiv cs.AI·12d agoPrimary

Transfer Learning Across Policy Regimes in Adaptive Multi-Agent Systems

This paper frames institutional and regulatory changes in adaptive socio-technical systems as a transfer-learning problem within multi-agent environments.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Efficiently Adapting Spoken Language Models for the Singaporean Context

Researchers adapted an open-source spoken language model to the multilingual Singaporean context using parameter-efficient fine-tuning and synthetic datasets.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Large Multimodal Model-Based Environment-Aware Mobility Management

This paper explores the integration of large multimodal models to improve mobility management in wireless networks by predicting user trajectories and making real-time decisions.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·13d agoPrimary

Multimodal Reward Hacking in Reinforcement Learning

This paper investigates the phenomenon of reward hacking in multimodal large language models aligned via reinforcement learning across various tasks and model scales.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection

A study on the ASVspoof5 dataset reveals how gender composition in training data affects the performance and bias of audio deepfake detection models.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Workload-Driven Optimization for On-Device Real-Time Subtitle Translation

The study explores optimization techniques for low-latency, on-device English-to-Traditional-Chinese subtitle translation under strict resource constraints.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Exploring Agentic Workflows for Generating High Quality Math Visual Aids

The paper examines how agentic workflows can be used to generate accurate and pedagogically effective mathematical diagrams for middle school education.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·13d agoPrimary

The Patchwork Problem in LLM-Generated Code

This paper identifies the patchwork problem in LLM-generated code, where locally correct code snippets fail to integrate coherently into larger software projects.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground Truth

This study introduces RouteCast, an evaluation framework designed to assess model-generated strategic routes when ground truth feedback is delayed or private.

Read at arxiv.org ↗
2
Regulation36Kr 36氪·10d ago

中信建投:关注能源结构的绿色低碳转型、新能源及氢基燃料的发展

New government guidelines in China emphasize the transition toward green energy and low-carbon industrial development for the upcoming five-year plan.

Read at 36kr.com ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Listen to the Features: Voice Anonymization Driven by Content Embedding Matching over Signal Reconstruction

Researchers developed a voice anonymization model that prioritizes content preservation over realistic speech generation by decoding content embeddings without waveform reconstruction loss.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

STAMP: Provenance-Guided Credit Assignment for Deep Search Agents

The proposed STAMP framework addresses the reward-credit mismatch in deep-search agents by using a reference-based verifier to evaluate whether cited documents support specific claims.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

CRiT-QA: Evaluating Multi-hop Reasoning with Counterfactual Chains and Distractor Traps

The CRiT-QA benchmark evaluates the multi-hop reasoning capabilities of large language models by testing their reliance on context versus internal knowledge using counterfactual chains and distractor traps.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT Systems

This paper proposes an affordable physical benchmark platform to evaluate the transferability of reinforcement learning models from simulation to real-world AIoT systems.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·12d agoPrimary

AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation

The paper proposes AdvNav, a behavior-guided black-box adversarial attack method targeting the vulnerabilities of Vision-and-Language Navigation systems.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Self-Regulated Reading with AI Support: An Eight-Week Study with Students

A longitudinal study of college students reveals how AI chatbot interactions influence cognitive engagement and reading habits during academic tasks.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

AgentLens is a new benchmark designed to evaluate coding agents by assessing the quality of their entire interaction trajectory rather than just final task outcomes.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Efficiently Learning Branching Networks for Multitask Algorithmic Reasoning

This study presents a method for training branching neural networks to perform multiple algorithmic reasoning tasks simultaneously.

Read at arxiv.org ↗
2
OpinionarXiv cs.AI·11d agoPrimary

Agentic Service-Oriented Computing: A Manifesto for the Next Frontier of Service-Oriented Computing

This manifesto outlines how LLM-powered autonomous agents are transforming software systems and proposes integrating them with service-oriented computing principles.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents

A novel reinforcement learning technique called Sibling-Guided Credit Distillation improves how agents learn to use tools by better attributing rewards to specific actions.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

When Does Personality Composition Matter for Multi-Agent LLM Teams?

Researchers investigated how personality-based prompting affects the task performance of multi-agent LLM teams, finding that communication styles significantly influence collaborative outcomes.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Learning When to Trust in Contextual Social Bandits

A new study identifies a failure mode called contextual sycophancy, where reinforcement learning agents receive biased feedback from evaluators in specific critical scenarios.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

TerraZero provides a procedural simulation environment and training stack to support the development of autonomous driving agents through large-scale self-play.

Read at arxiv.org ↗
2
ResearcharXiv cs.AI·11d agoPrimary

Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast

A study investigates how anatomical and contrast factors in brain MRI scans contribute to demographic predictability, highlighting potential biases in medical AI.

Read at arxiv.org ↗
← NewerPage 174Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.