✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
3
ResearcharXiv cs.AI·10d agoPrimary

PM-Bench: Evaluating Prospective Memory in LLM Agents

Researchers introduced PM-Bench, a text-based benchmark designed to evaluate the prospective memory capabilities of large language model agents.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement

A new framework for hardware design uses stepwise refinement to help large language models generate more reliable and verifiable register-transfer level code.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

AutoTrace: From Patches to Triggers via Agentic Interprocedural Exploration

The authors introduce an agentic approach for localizing vulnerability triggers in code by performing interprocedural causal reasoning.

Read at arxiv.org ↗
3
ProductarXiv cs.AI·10d agoPrimary

Fin-Analyst at FinMMEval 2026 Task 3: A Live Hybrid Trading Agent with LLM Specialists and Rule-Based Signals

Fin-Analyst is a hybrid trading agent that utilizes a multi-specialist LLM pipeline to process diverse financial data sources for automated equity market analysis.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Signal-Guided Optimization for Machine Unlearning

A new signal-guided optimization technique aims to improve machine unlearning by addressing the varying memorization strengths of individual training samples.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

CARE-LoRA: Compressed Activation REconstruction for Memory-Efficient LoRA

The authors introduce a method called CARE-LoRA designed to reduce memory usage during the fine-tuning of large pre-trained models.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric

Researchers propose a vendor-neutral metric to evaluate the reconstructability and validity of AI agent safety testing results.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models

The study introduces DRIFTLENS to measure how personalized memory injection in LLMs can alter the reasoning trajectories used to generate responses.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination

Researchers identified a unified geometric mechanism in transformer models that explains how conflicting memory sources lead to confident hallucinations.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

I'm Sorry, but I Can't Help with Braille: Revealing Accessibility Failures in State-of-the-Art LLMs

Research indicates that current state-of-the-art LLMs struggle significantly with bidirectional Korean-Braille translation, highlighting gaps in accessibility-focused capabilities.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Visual Access Boundaries in Vision-Language Model Reasoning

Researchers have introduced Visual Access Sweep, a causal intervention method to study how Vision-Language Models utilize image tokens during long Chain-of-Thought reasoning.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs

Researchers introduced a new method using execution-based semantic interaction graphs to better quantify uncertainty in code-generating large language models.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs

The authors investigate how the training duration of individual domain experts influences the performance of merged large language models.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model

Researchers proposed a method for automatic speech recognition using a discrete diffusion language model to transcribe audio in parallel rather than through traditional autoregressive decoding.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques

This survey examines various inference optimization techniques designed to improve the practical speed of masked diffusion large language models.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Line-Anchored Feedback Cuts Token Costs and Improves Correctness in AI Code Editing

Researchers demonstrated that using structured, line-anchored feedback in AI code editing tools significantly reduces token consumption and improves the accuracy of generated changes.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan Evaluation

This paper investigates a vulnerability in large language model plan evaluators where strategic plans are rewarded for omitting explicit details.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Mistake gating leads to energy and memory efficient continual learning

A biologically inspired learning method called mistake-gated learning is introduced to reduce the energy and memory requirements of training neural networks.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Inclusive Federated Learning Through Compliance-Weighted Noise Allocation in Healthcare AI

Researchers proposed a compliance-aware federated learning framework that adjusts differential privacy noise to accommodate varying institutional data standards and resource levels.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Comparing Semantic Navigation in Humans and Large Language Models using Natural Language Processing

A comparative study using NLP metrics reveals similarities and differences in how humans and leading large language models navigate conceptual spaces during semantic memory retrieval.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference

Researchers have developed a method to improve the real-time control of robots by optimizing the asynchronous inference of vision-language-action models on edge hardware.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning

Researchers have developed a variation-aware entropy scheduling method to improve reinforcement learning performance in environments subject to drift.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

PalmClaw: A Native On-Device Agent Framework for Mobile Phones

PalmClaw is a new framework designed to enable native, multi-step agentic task execution directly on mobile devices.

Read at arxiv.org ↗
3
ResearcharXiv cs.AI·10d agoPrimary

Evaluating Health Misinformation in Low-Resource Languages: Integrating Small Language Models with a Culturally-Sensitive Responsible NLP Framework (Bangla as a Case Study)

This research proposes a framework for detecting health misinformation in low-resource languages by combining small language models with culturally sensitive NLP techniques.

Read at arxiv.org ↗
← NewerPage 159Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.