INDEPENDENT AI NEWS / AI-assisted. Human-directed.EDITION 001 / 23 September 2026

Software

We hired AI to check the AI. Somebody check the invoice.

DoorDash measures what an AI code reviewer catches, misses and costs. More comments do not automatically mean better software.

Here’s what happened.

DoorDash describes evaluating AI code review through precision, recall, cost and latency. Its account shows why no single configuration wins on every measure. A review system needs to be judged on the defects it misses as well as the comments it produces.

Source: DoorDash engineering · AI code reviewer ↗

The We Are So Done take

We let AI write the code. Then we asked AI to review it. Now we need a system for deciding whether the review was any good. We have successfully automated the work and created a small department to verify that it happened.

For those of us already clicking

Count the misses

Access
DoorDash’s engineering report describes its own evaluation.
Start here
Use known defects and clean changes. Track missed bugs and false alarms as well as time and cost.
The small print
Results on an internal set of tasks do not establish performance on another repository, or a percentage of a profession that has been automated. This is the production team’s report, not our independent test.

Follow the evidence.

Source-based reporting. No independent hands-on test by this newsroom.
Sources checked 2026-09-23 · Editorial updated 2026-09-23

Still want in?

The tools behind the story.

The whole toolbox ↗