Research
The State of AI Engineering.
A first-class research program. We publish longitudinal benchmarks of human + AI engineering pairs, with an open scoring rubric and an invitation to independently replicate.
The State of AI Engineering — 2026
Codritium's first longitudinal benchmark of human + AI engineering pairs across 12 task categories. 4,200 real engineering tasks, scored on time-to-fix, regression rate, and reviewer-defensibility.
4,200 tasks · 612 engineers · 14 codebases
Read the reportLatest
Latest from each track.
One link per track. The full archives are linked below.
Report · June 15, 2026
The State of AI Engineering — 2026
Codritium's first longitudinal benchmark of human + AI engineering pairs across 12 task categories. 4,200 real engineering tasks, scored on time-to-fix, regression rate, and reviewer-defensibility.
4,200 tasks · 612 engineers · 14 codebases
Study · June 8, 2026
Security Engineering with AI Pairs
How AI-assisted engineers handle OWASP-class vulnerabilities — by category, by mitigation, by panel review. The disclosures we tracked, the fixes that stuck, and the patterns that fooled both pair and reviewer.
1,341 security tasks · 318 engineers · 22 vulnerability classes
Benchmark · May 20, 2026
The Prompt Pattern Library, 2026
Twenty-four prompt patterns, A/B-tested against the rubric. Which framings move correctness, regression, and defensibility — and which don't. Effect sizes with confidence intervals, paired results, and adoption rates.
24 patterns · 8,640 paired sessions · v2.6 rubric
Surfaces
Three surfaces.
Reports
Quarterly state-of reports on AI engineering, debugging, security, and code review. Public methodology.
Benchmarks
An open scoring rubric and replicable benchmark suite for human + AI engineering pairs across 12 task categories.
AI Engineering Studies
Smaller, targeted studies on specific workflows. Panel-reviewed before publication.