Faithful Autoformalization of Natural Language Assertions
The authors introduced Monty, a framework designed to improve the accuracy of autoformalization by converting natural language specifications into executable software assertions.
The authors introduced Monty, a framework designed to improve the accuracy of autoformalization by converting natural language specifications into executable software assertions.
Researchers have proposed a training-free approach for detecting human-object interactions in the wild by leveraging the capabilities of multimodal large language models.
An analysis of Kaggle contest submissions examines whether the use of AI coding assistants is leading to increased homogenization of software development outputs.
A study evaluating deep learning models for echocardiography analysis reveals that current attribution methods often fail to verify temporal faithfulness in medical imaging predictions.
This paper proposes a causal framework to determine when artificial intelligence systems should engage in theory of mind processes during conflict scenarios.
The proposed ScanFocus framework addresses computational efficiency and precision in spatio-temporal video grounding through a coarse-to-fine processing approach.
The paper provides a theoretical framework linking Joint-Embedding Predictive Architectures to active inference principles through variational free energy.
A new method for test-time learning introduces learnable adaptation policies to help language agents improve their performance through iterative interaction.
A new framework for discrete diffusion models explores how tokenization and vocabulary structure influence generative performance.
A new defense mechanism called Traffic-Aware Randomized Smoothing has been proposed to improve the robustness of LLM-based network intrusion detection systems against traffic manipulation.
A new framework enables mobile robots to interpret natural language commands for autonomous navigation using RGB-D perception.
The ExTernD method introduces a post-training factorization technique for large language models that utilizes expanded-rank ternary decomposition to improve quantization efficiency and accuracy.
This paper proposes a framework for networked intelligence that uses shared context graphs to facilitate collaboration among multiple AI agents in scientific research.
A new decoupled training strategy for transfer learning aims to reduce computational and energy costs by optimizing classifier heads separately from feature extraction layers.
DIVE is a new dimensionality reduction technique for language model embeddings that uses self-limiting gradient updates to improve compression efficiency.
This review examines the integration of explainable AI techniques within federated learning architectures to improve transparency in privacy-preserving distributed model training.
A study demonstrates that the interaction protocols used in multi-agent debates significantly influence the moral reasoning and judgment outcomes of large language models.
A new evaluation framework called STOCKTAKE aims to distinguish between perception errors and execution failures in LLM agents during long-term decision-making tasks.
A new error-correction method called Experience Memory Graph aims to help LLM agents recover from failures in long-horizon tasks more efficiently than traditional reflection techniques.
A new auditing method uses predicate substitution to test whether large language models genuinely rely on stated premises during chain-of-thought reasoning.
CITIC Securities and Optics Valley Financial Holdings have entered a partnership to provide capital market services and investment support for emerging technology companies.
Infineon and LS Electric have signed a non-binding memorandum of understanding to jointly develop next-generation direct current power systems for AI data centers.
Researchers have developed a contract-based regression verification tool that uses large language models to infer partial contracts, ensuring sound software patch verification without requiring manual specifications.
The SPARK framework profiles and steers the latent reasoning states of large language models to diagnose and correct reasoning failures before final outputs are generated.