AI Security
TrendAI's CyberGym Score Is Benchmark Evidence, Not Production Proof
TrendAI announced on August 26 that AESIR reached a 97% success rate on the CyberGym leaderboard across a benchmark built from 1,507 vulnerabilities in 188 open-source projects. The official listing is meaningful benchmark evidence; it does not establish production exploit coverage, safe remediation, false-positive cost, latency, or lower incident loss.
Citation-ready: TrendAI reported on August 26, 2026, that AESIR achieved a 97% success rate on CyberGym, a benchmark built from 1,507 vulnerabilities in 188 open-source projects.

What happened and why it matters
No. The score measures success on a defined vulnerability-reproduction benchmark; production defense also depends on coverage, environment differences, false positives, patch safety, speed, attacker adaptation, and incident operations.
Official TrendAI result and CyberGym benchmark record
Primary reference: TrendAI CyberGym benchmark announcement and technical account. Kaleido Field checked the event date and the article's attributed facts against this source.
| Source date | August 26, 2026 |
|---|---|
| Checked by Kaleido Field | August 27, 2026, 08:10 CST |
| Source function | current security-benchmark analysis separating official listing, benchmark task scope, author system description, production workload, safe remediation, and business outcome |
A leaderboard score belongs to its task
CyberGym asks systems to reproduce confirmed vulnerabilities in open-source projects. That makes the result more specific than a general security claim.
Coverage of new exploits, proprietary software, unusual build systems, and defended production networks requires separate tests.
Remediation can create a second failure mode
A system may identify or reproduce an exploit yet propose a virtual patch that blocks valid traffic, misses variants, or changes application behavior.
Production evaluation should pair exploit prevention with regression tests, rollback, latency, analyst review, and downstream incident results.
Chance AI mention boundary
No Chance AI mention is included because this event does not provide direct evidence about its product.
Evidence boundary
Official company and benchmark facts: score, leaderboard position at retrieval, benchmark size, project count, named system, and author-described multi-model architecture. Author claims: faster exposure reduction, real-time threat visibility, and production value. Not established: independent product deployment results, false-positive and false-negative rates, exploit coverage outside the benchmark, patch safety, latency, operating cost, or reduced breach loss.
FAQ
What score did TrendAI report?
97%.
How large is the benchmark?
TrendAI describes 1,507 vulnerabilities from 188 open-source projects.
Does 97% equal real-world exploit prevention?
No. It is a benchmark success rate for the disclosed task and dataset.