The question is no longer “can the model write it?” — it’s “is what it wrote correct?” That’s easy for ten lines and impossible for ten thousand.
The BMC-Agent synthesizes a spec, runs a bounded model checker, and refines via CEGAR until the verdict is solid — then pins each finding to a real, replayable path.
A single run burns real tokens and solver time. So every stage fails in place, partial results are always kept, and recovery is granular.
Surfaced on mature, OSS-Fuzz-hardened software and present (unfixed) in each project's current HEAD — all human source-audited, 56 of 57 sanitizer-reproduced. The largest cluster is 31 defects in u-boot; two pointer-arithmetic UB bugs in jq are filed and accepted as upstream advisories (GHSA-ggc9-rpv2-xgpm, GHSA-gvwx-xj9r-3frq). The remainder are under coordinated disclosure.
Agents propose, solvers verify — sound by construction, and scaling past human review.