GCSA Agent Achieves 91.3% on CyberGym, Joining Global Leading AI Security Agents
Developed by researchers at UC Berkeley, CyberGym is a large-scale, real-world cybersecurity evaluation framework. It comprises 1,507 historical vulnerability test cases across 188 major software projects, designed to assess AI agents’ practical abilities in genuine vulnerability analysis scenarios. Unlike traditional AI benchmarks that focus on code understanding or static analysis, CyberGym requires agents to interact directly with vulnerable codebases.
In its core Level 1 test, an AI agent receives only a vulnerability description and an unfixed codebase. It must autonomously perform code analysis, locate the flaw, reason about attack paths, construct, and execute a PoC. A task is deemed successful only if the generated PoC triggers the vulnerability in the unpatched version but fails to do so in the patched version. Thus, CyberGym measures whether an AI can complete the full cycle from security analysis to reproducible vulnerability validation.
From Large Models to Security Agents
In this test, GCSA Agent, powered by Grok 4.5 and 4.6 models, achieved its 91.3% success rate. The result highlights a crucial shift: ultimate security capability is no longer determined by the foundation model alone. Real-world vulnerability research is a multi-step process involving code retrieval, attack surface identification, hypothesis formation, input generation, execution, and iterative PoC refinement. GCSA Agent is built around this entire workflow, enabling it to form hypotheses, gather evidence, and validate findings through repeatable execution, moving beyond simple code analysis.
Real-World Vulnerability Research Capabilities
CyberGym’s core value is bridging the gap between AI testing and practical security research. It restores pre-patch code states, requiring agents to navigate large codebases with millions of lines of code. More significantly, further CyberGym research indicates this agentic security capability extends beyond replicating known flaws. In open-ended experiments, AI agents have already discovered previously unknown zero-day vulnerabilities and incomplete historical patches, showing the potential for migrating these skills to genuine discovery.
For GCSA, benchmark scores are not the final goal. The objective is to develop AI Security Agents that serve real cybersecurity needs, participating in the entire lifecycle of discovery, analysis, validation, and remediation.
Building AI-Native Cybersecurity
As AI accelerates software development, it is also transforming vulnerability research and defense. Facing increasingly complex systems, the next generation of cybersecurity will depend on collaboration between human experts and autonomous AI agents. These agents can help security teams by identifying exploitable vulnerabilities earlier, automatically analyzing complex attack paths, generating execution-level PoCs for validation, reducing false positives, and accelerating the entire remediation process.
GCSA Agent’s 91.3% score on CyberGym is a significant milestone in GCSA’s mission to build AI-native security. The alliance will continue advancing research to turn cutting-edge AI into practical security capabilities, fostering a more secure, trustworthy, and resilient digital environment.
Source: GCSA
Website: www.gcsa.org
(資料由客戶提供)



