✦FireflyAI news, ranked by what matters
✦ Night SkyListLatest
✦ TopLatest
Model ReleaseFundingRegulationResearchInfrastructureProductOpinion
4
ResearcharXiv cs.AI·7d agoPrimary

Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs

Researchers have developed a new evaluation framework called JKP to test the epistemic stability of vision-language models when subjected to repetitive adversarial questioning.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation

This study investigates common assumptions regarding prompting techniques and dataset usage in the evaluation of large language models for multiple-choice tasks.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

MAPS: Modeling Co-Existing Subjective Perspectives and Shared Meaning in Multi-Agent Cognitive Dialogue

The proposed MAPS framework enables multi-agent dialogue systems to model distinct subjective perspectives and cognitive styles during interaction.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

RetroAgent: Harnessing LLMs to Search Over Structured Memory for Agentic Retrosynthesis Planning

RetroAgent utilizes large language models to navigate structured memory for more effective multi-step retrosynthesis planning in chemical research.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation

Researchers introduced a schema-aware grounding method to improve the accuracy of agentic text-to-SPARQL query generation for knowledge base question answering.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Step-Level Preference Learning for Generative Agents in Social Simulations

This research introduces a step-level preference learning method to better align the intermediate decision-making processes of generative agents with human preferences.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models

The study demonstrates that high transition accuracy in LLM-synthesized world models does not necessarily correlate with effective planning performance.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability

ToolAnchor is a proposed method to help AI agents overcome behavioral inertia and more effectively integrate new tools into their workflows.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

HG-RAG: Hierarchy-Guided Retrieval-Augmented Generation for Structured Knowledge Graphs

Researchers have developed a framework called HG-RAG that improves retrieval-augmented generation by utilizing hierarchical graph traversal for more effective knowledge retrieval.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Global drivers and barriers to the public acceptance of autonomous vehicles: Evidence from 17 countries

A multi-country study identifies the primary global factors influencing public acceptance of Level 3 autonomous vehicles.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding

Researchers have introduced VideoSEMA, a hybrid architecture combining Mamba-like blocks and temporal attention to improve the efficiency of video classification models.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Assessing AI in Introductory Physics Problem Solving

A study evaluated the performance of OpenAI's o4-mini model on undergraduate physics problems to assess the reasoning capabilities of current inference-scaling LLMs.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

ViPSAM: Visual Prompting Medical Image Segmentation Using Segment Anything Model

ViPSAM is a new visual prompting framework designed to improve medical image segmentation in non-contrast CT scans by leveraging contrast-enhanced MRI data.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI

This paper argues for the application of sociotechnical systems analysis to better understand and mitigate risks in automated decision-making technologies.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Integration Matters: Rollout-Based Training for Constrained Diffusion Models

The authors introduce a rollout-based training method to improve the performance of constrained diffusion models in satisfying complex feasibility requirements.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

LIGO-PINN: Learned Initialization via Gated Optimization to Alleviate Convergence Failures in Physics Informed Neural Networks

LIGO-PINN introduces a gated optimization technique to improve the training stability and convergence of physics-informed neural networks in complex domains.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution

Researchers investigate the efficacy of specialized agentic systems compared to general-purpose LLMs for automating business process workflows.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi

This study explores the feasibility of running modern multimodal AI models on legacy hardware by deploying a compact assistant on a 2011-era GPU.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

Researchers introduced Contrastive Policy Optimization to improve reinforcement learning by using token-level disagreement to better distinguish between useful uncertainty and harmful confusion.

Read at arxiv.org ↗
4
InfrastructurearXiv cs.AI·7d agoPrimary

An Intelligent-Cloud Edge Multimodal Interaction System for Robots

A new cloud-edge framework combines gesture detection with LLM and VLM agents to enable efficient multimodal interaction for robots with limited onboard computing power.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity

Researchers have introduced a new benchmark designed to evaluate how language models handle ambiguous instructions, policy conflicts, and adversarial commands.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·10d agoPrimary

Playful AI in Professional Email: A Field Experiment on Tone and Recipient Engagement

A randomized field experiment evaluated how professional and playful emails rewritten by GPT-5 affected recipient engagement and behavior in workplace communications.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

Latent Trajectory Discrimination for AI-Generated Text Detection

Researchers propose a new method for detecting AI-generated text by analyzing the latent trajectories of the autoregressive generation process.

Read at arxiv.org ↗
4
ResearcharXiv cs.AI·7d agoPrimary

RAD: Retrieval High-quality Demonstrations to Enhance Decision-making

The proposed RAD framework improves offline reinforcement learning by retrieving high-quality demonstrations to help models generalize beyond static training datasets.

Read at arxiv.org ↗
← NewerPage 142Older →
Firefly aggregates headlines and links to original sources. All content belongs to its respective publishers.