✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
7
ResearcharXiv cs.AI·6d agoPrimary

Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs

Researchers introduced a new method using execution-based semantic interaction graphs to better quantify uncertainty in code-generating large language models.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

Sparse Autoencoders for Interpretable Out-of-Distribution Detection

This study introduces a method using sparse autoencoders to better detect out-of-distribution data by analyzing intermediate layers of neural networks.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

Signal-Guided Optimization for Machine Unlearning

A new signal-guided optimization technique aims to improve machine unlearning by addressing the varying memorization strengths of individual training samples.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

AutoTrace: From Patches to Triggers via Agentic Interprocedural Exploration

The authors introduce an agentic approach for localizing vulnerability triggers in code by performing interprocedural causal reasoning.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

Self-Consistent Flow: Unifying Velocity and Endpoint Prediction for Rectified Flow Models

Researchers have developed a unified approach for rectified flow models that combines velocity and endpoint prediction to improve generative model training.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

CARE-LoRA: Compressed Activation REconstruction for Memory-Efficient LoRA

The authors introduce a method called CARE-LoRA designed to reduce memory usage during the fine-tuning of large pre-trained models.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

I'm Sorry, but I Can't Help with Braille: Revealing Accessibility Failures in State-of-the-Art LLMs

Research indicates that current state-of-the-art LLMs struggle significantly with bidirectional Korean-Braille translation, highlighting gaps in accessibility-focused capabilities.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model

Researchers proposed a method for automatic speech recognition using a discrete diffusion language model to transcribe audio in parallel rather than through traditional autoregressive decoding.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan Evaluation

This paper investigates a vulnerability in large language model plan evaluators where strategic plans are rewarded for omitting explicit details.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

Evidence-Grounded AI for Musculoskeletal Care

This work discusses the application of evidence-grounded AI systems to support the longitudinal management and rehabilitation of musculoskeletal diseases.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

PM-Bench: Evaluating Prospective Memory in LLM Agents

Researchers introduced PM-Bench, a text-based benchmark designed to evaluate the prospective memory capabilities of large language model agents.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

Cost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems

Researchers introduced Cost-Governed RAG, an architecture designed to attribute both retrieval and generation costs to individual tenants in multi-user language model systems.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

Comparing Semantic Navigation in Humans and Large Language Models using Natural Language Processing

A comparative study using NLP metrics reveals similarities and differences in how humans and leading large language models navigate conceptual spaces during semantic memory retrieval.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·6d agoPrimary

Pigeonholing: how bad prompts hurt models, causing collapse and mistakes

Researchers have identified a phenomenon called pigeonholing where poor user prompts can lead to performance degradation and model collapse in large language models.

Read at arxiv.org ↗
7
ProductData Center Dynamics·6d ago

Plug Power sells Texas site to Stream Data Centers

Plug Power has sold a Texas industrial site to Stream Data Centers, which intends to repurpose the land for data center infrastructure.

Read at datacenterdynamics.com ↗
7
ResearcharXiv cs.AI·6d agoPrimary

Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels

A new protocol called JADR has been proposed to measure the underlying safety fragility of language models beyond standard jailbreak testing.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

When Data Imbalance Helps: Robust Generalization Through Shortcut Saturation

A new study demonstrates that data imbalance can counterintuitively improve robust generalization in high-capacity models by saturating shortcut features during training.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Behavioural Signatures of Risk-Sensitive Decision-Making in Large Language Models

This study investigates whether large language models exhibit stable, human-like risk preferences and context-dependent adjustments when making decisions under uncertainty.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning

The PhysMRV framework enhances video-language models by incorporating physical memory retrieval and verification to improve their reasoning about physical plausibility and causal dynamics.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure

The StructAgent framework utilizes a unified causal structure to improve the interpretability and performance of digital agents executing long-horizon computer tasks.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Heterogeneous Agent Cohorts for Safe Open-Ended Exploration with Runtime Constraint Memory

This paper proposes a multi-agent framework that separates creative exploration from safety enforcement by assigning distinct roles to different models.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

Filtering Harmful Actions Isn't Enough: Phantom Transfer in Agentic SDF

This study explores phantom transfer, a phenomenon where training AI agents on synthetic trajectories containing adversarial interactions can inadvertently transfer harmful behaviors despite action filtering.

Read at arxiv.org ↗
7
ResearcharXiv cs.AI·7d agoPrimary

A Symbolic Neural CPU for Quantization-Simulated Writeback and Interpretable Program Execution

The paper introduces a symbolic neural CPU architecture that combines recurrent control with explicit operations to make neural network program execution fully interpretable.

Read at arxiv.org ↗
7
Product36Kr 36氪·9d ago

今年1-5月,我国云计算设备、半导体设备出口金额累计同比分别大增114.4%、91.5%

Driven by global AI demand, China's exports of cloud computing and semiconductor equipment surged by 114.4% and 91.5% respectively in the first five months of the year.

Read at 36kr.com ↗
← NewerPage 96Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.