Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs
Researchers have developed a new evaluation framework called JKP to test the epistemic stability of vision-language models when subjected to repetitive adversarial questioning.
Researchers have developed a new evaluation framework called JKP to test the epistemic stability of vision-language models when subjected to repetitive adversarial questioning.
This study investigates common assumptions regarding prompting techniques and dataset usage in the evaluation of large language models for multiple-choice tasks.
The proposed MAPS framework enables multi-agent dialogue systems to model distinct subjective perspectives and cognitive styles during interaction.
RetroAgent utilizes large language models to navigate structured memory for more effective multi-step retrosynthesis planning in chemical research.
Researchers introduced a schema-aware grounding method to improve the accuracy of agentic text-to-SPARQL query generation for knowledge base question answering.
This research introduces a step-level preference learning method to better align the intermediate decision-making processes of generative agents with human preferences.
The study demonstrates that high transition accuracy in LLM-synthesized world models does not necessarily correlate with effective planning performance.
ToolAnchor is a proposed method to help AI agents overcome behavioral inertia and more effectively integrate new tools into their workflows.
Researchers have developed a framework called HG-RAG that improves retrieval-augmented generation by utilizing hierarchical graph traversal for more effective knowledge retrieval.
A multi-country study identifies the primary global factors influencing public acceptance of Level 3 autonomous vehicles.
Researchers have introduced VideoSEMA, a hybrid architecture combining Mamba-like blocks and temporal attention to improve the efficiency of video classification models.
A study evaluated the performance of OpenAI's o4-mini model on undergraduate physics problems to assess the reasoning capabilities of current inference-scaling LLMs.
ViPSAM is a new visual prompting framework designed to improve medical image segmentation in non-contrast CT scans by leveraging contrast-enhanced MRI data.
This paper argues for the application of sociotechnical systems analysis to better understand and mitigate risks in automated decision-making technologies.
The authors introduce a rollout-based training method to improve the performance of constrained diffusion models in satisfying complex feasibility requirements.
LIGO-PINN introduces a gated optimization technique to improve the training stability and convergence of physics-informed neural networks in complex domains.
Researchers investigate the efficacy of specialized agentic systems compared to general-purpose LLMs for automating business process workflows.
This study explores the feasibility of running modern multimodal AI models on legacy hardware by deploying a compact assistant on a 2011-era GPU.
Researchers introduced Contrastive Policy Optimization to improve reinforcement learning by using token-level disagreement to better distinguish between useful uncertainty and harmful confusion.
A new cloud-edge framework combines gesture detection with LLM and VLM agents to enable efficient multimodal interaction for robots with limited onboard computing power.
Researchers have introduced a new benchmark designed to evaluate how language models handle ambiguous instructions, policy conflicts, and adversarial commands.
A randomized field experiment evaluated how professional and playful emails rewritten by GPT-5 affected recipient engagement and behavior in workplace communications.
Researchers propose a new method for detecting AI-generated text by analyzing the latent trajectories of the autoregressive generation process.
The proposed RAD framework improves offline reinforcement learning by retrieving high-quality demonstrations to help models generalize beyond static training datasets.