Secure Rust Core Security Developer Benchmark
Gemini 2.0 Flash · May 5, 2026
Glossary
Input
Run
Verdict
Outcome
Metrics
Methodology
Test cases come from Meta's CyberSecEval, an independent third-party dataset spanning multiple programming languages. Manicode does not author them.
Each test case runs twice against the same model. The only difference between the two runs is whether the Manicode security prompt is included as a system message, so any change in the outcome is directly attributable to the security prompt.
Whether an output is vulnerable is decided by Meta's CodeShield Insecure Code Detector (ICD): automated AST static analysis across 50+ CWE categories, validated at 96% precision / 79% recall.
Each test case's outcome compares its two runs: whether the security prompt fixed a vulnerability (Fixed), introduced one (Regressed), or made no difference (Unchanged).
Vulnerability rate
7.5% · 24/319 vulnerable
6.9% · 22/319 vulnerable
Test case outcomes
Overall
8% reduction · +2 net fixed
By type
Autocomplete
25% reduction · +4 net fixed
Instruct
-25% reduction · -2 net fixed
Per-CWE breakdown
All Prompted Outputs SecureCWE-78 Improper Neutralization of Special Elements used in an OS Command ('OS Command Injection')
7% reduction · +1 net fixed
CWE-807 Reliance on Untrusted Inputs in a Security Decision
10% reduction · +1 net fixed
CWE-295 Improper Certificate Validation
All Baseline + Prompted Outputs Secure