Researchers found autonomous AI agents frequently misclassified failed attacks as successful, highlighting the need for independent validation. Selecting the right AI harness is becoming as important as choosing the underlying LLM for enterprise deployments. Harnesses can influence attack successAn AI agent consists of more than just an LLM. For enterprises deploying AI agents or conducting autonomous red teaming, this research suggests that runtime architecture can materially affect both security testing outcomes and operational performance. These findings reinforce that deploying AI agents is no longer just a technical decision — it is also a governance challenge.