mimile
Back to feed

AI Agents Cut False Positives by Interacting Directly With Software

AI digest

This digest was compiled by AI from multiple sources — links to the originals are below.

AI Agents Cut False Positives by Interacting Directly With Software

AI agents that run commands and observe software behavior can tell real vulnerabilities from false positives, Arizona State University associate professor Yan Shoshitaishvili said. Speaking at Black Hat USA 2026, he described the agentic loop as automating tasks such as root cause analysis and malware reverse engineering.

Key Facts

  • Artificial intelligence agents can run commands, observe software behavior and test whether a suspected flaw can be triggered.
  • That capability reduces false positives, a long-standing burden in software security, Shoshitaishvili said.
  • Shoshitaishvili is an associate professor at Arizona State University and focuses on automated program analysis and vulnerability detection.
  • He led Shellphish's participation in the DARPA Cyber Grand Challenge, which produced a fully autonomous hacking system.

Agentic Loop Mechanics

AI agents go beyond code review by running commands, observing behavior and testing whether a suspected flaw can be triggered, Shoshitaishvili said. Traditional program analysis tools often rely on abstractions and assumptions that create imprecision. Agents can follow investigative steps, test hypotheses and classify findings at a level closer to human review. Shoshitaishvili described the ability to filter out false positives with an agentic loop as a 'big superpower'.

Automation of Reverse Engineering

Human analysts traditionally spend a long time reverse-engineering malware to understand it, Shoshitaishvili noted. AI agents are increasingly automating that process, he said. The agentic approach also automates root cause analysis and malware reverse engineering. Shoshitaishvili compared LLM vulnerability discovery with static and dynamic analysis techniques at Black Hat USA 2026.

Research and Background

Shoshitaishvili is an associate professor at Arizona State University, where he focuses on cybersecurity research, education and real-world security impact. His research centers on automated program analysis and vulnerability detection. He has published dozens of research papers in the field. He led Shellphish's participation in the DARPA Cyber Grand Challenge, creating a fully autonomous hacking system.