On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage
Researchers analyzed the citation faithfulness and coverage of a four-billion parameter research agent running locally on a consumer laptop.
Researchers analyzed the citation faithfulness and coverage of a four-billion parameter research agent running locally on a consumer laptop.
The paper introduces NameRank, a metric designed to measure how well large language models recognize specific researchers and tools within their parametric memory.
Chinese startup Novasilicon has secured multi-million yuan in seed funding to develop AI-driven chip design solutions.
Developers have successfully compiled the Firefox browser into WebAssembly, allowing it to run entirely within another web browser.
Verrus, an Alphabet spin-out, is planning the construction of a new grid-reactive data center campus in Oregon to support infrastructure needs.
Chinese storage company Demingli projects a first-half net profit of up to 6.5 billion yuan, reversing a previous loss due to surging AI-driven storage demand.
The global artificial intelligence boom has driven up memory and solid-state drive prices in Shenzhen's Huaqiangbei electronics market, significantly increasing costs for PC builders.
DIVE is a new dimensionality reduction technique for language model embeddings that uses self-limiting gradient updates to improve compression efficiency.
This review examines the integration of explainable AI techniques within federated learning architectures to improve transparency in privacy-preserving distributed model training.
A study demonstrates that the interaction protocols used in multi-agent debates significantly influence the moral reasoning and judgment outcomes of large language models.
A new evaluation framework called STOCKTAKE aims to distinguish between perception errors and execution failures in LLM agents during long-term decision-making tasks.
A new error-correction method called Experience Memory Graph aims to help LLM agents recover from failures in long-horizon tasks more efficiently than traditional reflection techniques.
A new auditing method uses predicate substitution to test whether large language models genuinely rely on stated premises during chain-of-thought reasoning.
The authors present a method for integrating compiler feedback directly into the autoregressive decoding process to improve the quality of AI-generated code.
Researchers have proposed a training-free approach for detecting human-object interactions in the wild by leveraging the capabilities of multimodal large language models.
MedDiffuseMix provides a saliency-guided diffusion framework to augment medical imaging data while preserving critical diagnostic features.
This paper proposes a framework for networked intelligence that uses shared context graphs to facilitate collaboration among multiple AI agents in scientific research.
A new method for test-time learning introduces learnable adaptation policies to help language agents improve their performance through iterative interaction.
A new benchmark based on cognitive psychology tests has been developed to evaluate how AI agents adapt when the reliability of their tools changes during operation.
A study evaluating root cause analysis in microservice failures reveals that current AI and classical methods struggle to effectively process large-scale, multimodal telemetry data.
Researchers analyzed nearly 3,000 GitHub projects to understand how the integration of automated bots as active participants influences the organizational structure of open-source software teams.
HealthClaw is a proposed self-evolving agent architecture designed to provide longitudinal health management by maintaining private memory of user routines and medical history.
The study evaluates the impact of audio separation preprocessing on zero-shot automatic speech recognition performance, finding that cleaner audio does not always improve transcription accuracy.
FixItFlow is an automated system that leverages large language models to generate technical troubleshooting guides from historical cloud incident data.