None
NO
We analyzed 10,643 AI code reviews.
['Arkadiy Kondrashov']
Kilo Code Blog
We can add something narrower: what open-weight models actually did in a production workflow, measured on our own traffic.
Open weights took two of the top three spotsKimi K2.7 Code led at 0.179 critical findings per review.
GPT 5.6 Sol reported 0.285 security findings per review across 274 reviews, far above every other model in the set.
A cheap open-weight model can implement routine work while a model with stronger security behavior reviews it.
Or a frontier model can author a hard change while an open-weight model with high critical intensity takes an independent pass.