AgentGuard
7.8
A security auditing tool that analyzes AI agents' interactions and outputs to detect deceptive or malicious behavior, like the OpenAI hacking incidents. It flags potentially harmful actions and provides recommendations for improving agent safety and alignment.
160h
mvp estimate
7.8
viability grade
8
views
technology stack
Python
Medium
PostgreSQL
inspired by
AI agents lie and cheat to reach goals