✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
7
ResearcharXiv cs.AI·6d agoPrimary

Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model

Researchers proposed a method for automatic speech recognition using a discrete diffusion language model to transcribe audio in parallel rather than through traditional autoregressive decoding.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

Evidence-Grounded AI for Musculoskeletal Care

This work discusses the application of evidence-grounded AI systems to support the longitudinal management and rehabilitation of musculoskeletal diseases.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

PM-Bench: Evaluating Prospective Memory in LLM Agents

Researchers introduced PM-Bench, a text-based benchmark designed to evaluate the prospective memory capabilities of large language model agents.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

Cost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems

Researchers introduced Cost-Governed RAG, an architecture designed to attribute both retrieval and generation costs to individual tenants in multi-user language model systems.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources

RESOURCE2SKILL is a new framework designed to extract procedural agent skills from diverse multimodal human resources like tutorial videos and technical documentation.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

Pigeonholing: how bad prompts hurt models, causing collapse and mistakes

Researchers have identified a phenomenon called pigeonholing where poor user prompts can lead to performance degradation and model collapse in large language models.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

Comparing Semantic Navigation in Humans and Large Language Models using Natural Language Processing

A comparative study using NLP metrics reveals similarities and differences in how humans and leading large language models navigate conceptual spaces during semantic memory retrieval.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels

A new protocol called JADR has been proposed to measure the underlying safety fragility of language models beyond standard jailbreak testing.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs

The authors investigate how the training duration of individual domain experts influences the performance of merged large language models.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

When Data Imbalance Helps: Robust Generalization Through Shortcut Saturation

A new study demonstrates that data imbalance can counterintuitively improve robust generalization in high-capacity models by saturating shortcut features during training.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Behavioural Signatures of Risk-Sensitive Decision-Making in Large Language Models

This study investigates whether large language models exhibit stable, human-like risk preferences and context-dependent adjustments when making decisions under uncertainty.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning

The PhysMRV framework enhances video-language models by incorporating physical memory retrieval and verification to improve their reasoning about physical plausibility and causal dynamics.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure

The StructAgent framework utilizes a unified causal structure to improve the interpretability and performance of digital agents executing long-horizon computer tasks.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Heterogeneous Agent Cohorts for Safe Open-Ended Exploration with Runtime Constraint Memory

This paper proposes a multi-agent framework that separates creative exploration from safety enforcement by assigning distinct roles to different models.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Filtering Harmful Actions Isn't Enough: Phantom Transfer in Agentic SDF

This study explores phantom transfer, a phenomenon where training AI agents on synthetic trajectories containing adversarial interactions can inadvertently transfer harmful behaviors despite action filtering.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

A Symbolic Neural CPU for Quantization-Simulated Writeback and Interpretable Program Execution

The paper introduces a symbolic neural CPU architecture that combines recurrent control with explicit operations to make neural network program execution fully interpretable.

Read at arxiv.org ↗
7
Product36Kr 36氪·9d ago

今年1-5月,我国云计算设备、半导体设备出口金额累计同比分别大增114.4%、91.5%

Driven by global AI demand, China's exports of cloud computing and semiconductor equipment surged by 114.4% and 91.5% respectively in the first five months of the year.

Read at 36kr.com ↗
7
OpinionTechCrunch AI·5d ago

Why AMI Labs’ Alexandre LeBrun won’t call his AI ‘AGI’ or ‘superintelligence’

AMI Labs CEO Alexandre LeBrun explains his decision to avoid using terms like AGI or superintelligence when describing his company's AI models.

Read at techcrunch.com ↗
7
FundingData Center Dynamics·7d ago

Trinidad and Tobago inks MoUs with American companies for AI data centers

Trinidad and Tobago has signed memorandums of understanding with American firms to explore the development of AI-focused data center infrastructure.

Read at datacenterdynamics.com ↗
7
Regulation36Kr 36氪·11d ago

苹果在重磅诉讼中起诉OpenAI窃取商业机密‌

Apple has filed a lawsuit against OpenAI, accusing the AI startup of systematically encouraging Apple employees to leak trade secrets and proprietary designs to develop its own hardware products.

Read at 36kr.com ↗
7
ResearchQbitAI 量子位·8d ago

机器狗指挥人类用天平称重!清华现场演示:无脚本,任务随机,观众即兴出题

Tsinghua University researchers demonstrated an unscripted physical AI system where a robotic dog autonomously directed humans to perform weighing tasks based on real-time audience prompts.

Read at qbitai.com ↗
7
Funding36Kr 36氪·6d ago

长鑫科技注册资本10年增超6000倍

Semiconductor firm Changxin Technology has finalized its IPO pricing and significantly expanded its registered capital and workforce over the past decade.

Read at 36kr.com ↗
7
RegulationMIT News AI·8d ago

New method aims to keep kids safe from illegal AI-generated content

Researchers have developed an auditing technique designed to identify malicious capabilities in generative AI models without requiring the generation of prohibited content.

Read at news.mit.edu ↗
7
Model ReleasearXiv cs.AI·8d agoPrimary

A Sovereign, Open-Source Foundation Model for German and English

The release of Soofi S 30B-A3B introduces an open-source, sovereign mixture-of-experts hybrid model optimized for German and English language tasks with high throughput efficiency.

Read at arxiv.org ↗
← NewerPage 97Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.