authzscan on Real Code: What a 100% Benchmark Score Failed to Predict
The live eval landed at 100% recall and precision. Then I pointed the scanner at a real open-source repo and it found one genuine bug in eleven candidates. The gap was in my benchmark, not the model.
Read the writeup- Reading
- 7 min
- Words
- 1,501
- Difficulty
- hard
- Topics
- #idor #bola #agents