Measured quality
What ocra found and missed on its golden cases and on AACR-Bench, how that was measured, and what the numbers cannot tell you.
This page reports what ocra's evaluation runs have measured so far, how they were scored, and their limits. The sample is small and every run used one model family, so read Limits before relying on a number.
On 8 changes with 16 known issues, the current configuration found 7 (44%) in its one run, and all 7 findings it reported were correct. It found 5 of the 6 issues the cases rate as warnings and 2 of the 10 rated as suggestions. Use ocra as a second reviewer: read what it reports, and do not take an empty report as approval.
How it is measured
ocra-eval runs the real ocra review on changes pinned at fixed commits and scores what it reports (Evaluation shows how to run it). There are two datasets:
- Golden cases (ADR-0011) belong to ocra. Each case pins a public repository at two commits and lists the issues ocra must report, other places where the same issue may rightly be reported, ranges where a finding would be wrong, or that the change has no issue at all. Every run so far used the 8 cases the smoke tier had then, now with 16 expected issues. The smoke tier has since grown to 10 cases and the full tier to 16; neither has been run as it is now.
- AACR-Bench is an external benchmark of real pull requests; its ground truth is the comments human reviewers left. It is the number comparable with published results.
| The 8 cases these runs used | Cases | Expected issues | Languages |
|---|---|---|---|
| Changes from ocra's own history: bugs fixed later, and one change with no issue | 4 | 7 | TypeScript |
| AACR-Bench pull requests, with issues checked against the code | 4 | 9 | C#, C++, Go, TypeScript |
On the golden cases, a finding matches an expected issue when it is in the same file, at most one line from the expected lines or from one of the other places, and a judge model says it describes the same concern.
- Recall is the share of expected issues found at or above the severity the case asks for.
- Precision is the share of reported findings that match an expected issue or are labeled valid. A finding that matches nothing is checked against the code once and labeled valid or invalid, and the label is stored with the case; until then it counts as wrong.
- A critical finding in a range marked wrong, or on a change with no issue, counts as a failure. No run has had one.
AACR-Bench uses the benchmark's own matching: same file and side of the diff, at most one line apart, then the judge model. Its precision is matches over findings, without labels.
Every run used Gemini on Vertex AI: gemini-3.5-flash for the reviewers and verification (gemini-3.5-flash-lite as failback), gemini-3.1-pro-preview for the judge step, gemini-3.5-flash-lite for light tasks, and gemini-flash-lite-latest to decide matches when scoring. Costs are ocra's own accounting: token counts at the models' list prices.
Golden cases
The current configuration
One run on 2026-09-29 (UTC), with ocra's default settings as of main at 664c850 plus the 30-step limit per review agent merged right after it (f966d83):
- Found 7 of the 16 expected issues (44%): 5 of the 6 the cases rate as warnings, 2 of the 10 rated as suggestions. No case expects a critical issue.
- Reported 7 findings, all correct (7 of 7, 100%), and nothing on the change with no issue.
- Cost $12.53 for the 8 changes: $1.57 each on average (median $1.15, from $0.07 to $3.98). A review took 243 seconds at the median and 437 at the 90th percentile.
It found header declarations and callers out of step with their definitions, which breaks the build (FreeCAD, C++); a test's monitor container left on the public image when its server container moved to a mirror (ASP.NET Core, C#); files created before their parent directory exists (Ollama, Go); an import of a name the module does not export, which breaks the build (RAGFlow, TypeScript); and three bugs from ocra's history: documentation promising a .gitignore that no code writes, a helper function registered as a plugin, and reference-style Markdown links through which model-written text could still put live links into pull request comments.
It missed four bugs from ocra's history (a temporary copy of the git index with a new modification time, so a file edited in the same second as the real index can look unchanged and be left out of the review; a blocking verdict returned before a new coverage check; unchanged files counted as not reviewed; a run counted as reviewing nothing when one reviewer failed on files another had finished), four smaller validation, error-handling and test issues in Ollama, and missing translation keys in RAGFlow.
Where the expected issues come from
Recall depends on which issues are expected. Four of the 16 were written from bugs fixed in ocra's history, before any run; the current configuration found 2 of them. The other 12 were added after an ocra run reported them and a check against the code confirmed them: 5 from runs on AACR-Bench pull requests, 7 from runs on these cases, one of them the run above. On those 12, recall shows whether ocra finds again what it has found before; it found 5.
Every run
All runs on those 8 cases, scored against today's cases and labels so that they compare. Three earlier runs, on the first 4 of them, each found 1 of those cases' 7 expected issues and reported nothing wrong. Runs that only checked mechanics, or compared a change on a handful of AACR-Bench pull requests, are left out.
| Date (UTC, 2026) | Configuration | Found | Precision | Cost per change |
|---|---|---|---|---|
| 09-28 | 20 steps per review agent | 6/16 (38%) | 6/7 (86%) | $1.16 |
| 09-28 | The same, run again | 6/16 (38%) | 7/7 (100%) | $0.97 |
| 09-28 | 30 steps per review agent | 7/16 (44%) | 7/8 (88%) | $1.06 |
| 09-28 | The same, run again; 7 of the 8 changes (spend limit) | 8/15 (53%)¹ | 8/9 (89%) | $1.77 |
| 09-28 | 20 steps, with prompt-injection boundaries | 4/16 (25%) | 4/5 (80%) | $1.33 |
| 09-29 | The current configuration | 7/16 (44%) | 7/7 (100%) | $1.57 |
| 09-29 | A later main plus a prompt change that was not merged | 6/16 (38%)² | 8/9 (89%) | $1.26 |
¹ The 16th expected issue was taken from this run.
² One more expected issue was found below the severity its case asks for.
Both configurations that ran twice found about the same number of issues each time (6 and 6, 7 and 8), and their precision differed by one finding. Two runs cannot show the real spread, so treat a difference of one or two issues between configurations as noise.
AACR-Bench
Ten pull requests (--seed 1 --limit 10 --max-change-lines 300: at most 300 changed lines each, 44 annotated issues), with two configurations run twice each on 2026-09-28, before the current configuration:
| Configuration | Precision | Recall | Cost per PR |
|---|---|---|---|
| 20 steps per review agent | 4/6 (67%); 2/5 (40%) | 4/44 (9.1%); 2/44 (4.5%) | $0.77; $0.86 |
| The same, told to report each finding as soon as it is confirmed | 3/14 (21%); 3/7 (43%) | 3/44 (6.8%); 3/44 (6.8%) | $0.97; $1.02 |
The current configuration has not been run on AACR-Bench. Three facts about the benchmark matter when reading these numbers:
- Recall has a ceiling. Of the 530 issues annotated on 94 of its pull requests, 40% are maintainability and readability comments, which ocra does not report by design, so at most 58% are within ocra's reach (
ocra-eval ceiling, which calls no model). - Precision is understated. The ground truth is the comments reviewers happened to leave, so a real issue nobody commented on counts as a false positive. A one-off check of each finding against the code (by an AI agent working on ocra, not recorded as labels) found 6 real defects, 4 minor issues and 1 wrong finding among the 11 findings of the first two runs, where the benchmark matched 6; and 15 real defects, 5 minor issues and 1 wrong finding among the 21 of the last two, where it matched 6.
- Matching allows one line of distance. In one Electron pull request, ocra reported the annotated use-after-free ten lines away from the reference, which counted as a miss and as a false positive.
ocra-evalreports a lenient match (same concern in the same file, at any line) next to the official one for this reason.
Cost
In these runs, at list prices:
- with the current configuration, a review of a golden case cost $1.57 on average (median $1.15), from $0.07 to $3.98;
- with earlier configurations, a review of one of AACR-Bench's smaller pull requests cost $0.77 to $1.02 on average;
- the most expensive single review in any run cost $4.31, for a change to 10 files.
Cost grows with the size of the change and with the reviewers its risk tier runs. --max-cost-usd caps a run's spend (CLI).
Limits
- Small. 8 changes and 16 expected issues, of which only 4 were written independently of ocra's own findings. The current configuration ran once, and the full tier has not run.
- Not independent. The cases and labels are the project's own. The labels so far were made by an AI agent working on ocra, which checked each finding against the code and recorded its reason. A second model (Claude Fable 5.1), not told the labels, re-judged 5 of the 11 labels recorded by then against the code and agreed with all 5; no person outside the project has reviewed them. Recent changes, the 30-step limit among them, were chosen on these cases, which flatters the current configuration.
- A model decides matches. Whether a finding describes an expected issue is decided by
gemini-flash-lite-latest. Its answers are cached, so scoring the same run again gives the same result, but another judge could decide differently. - One model family. Every run used Gemini on Vertex AI. Other models will score differently.
- Runs vary. Two runs of the same configuration differed by one finding in precision.
- Recall is the weak point. ocra missed more than half of the expected issues here, most of them rated as suggestions, and on AACR-Bench it matched under 10% of the annotated issues.
- Changes since the last run are not measured, among them that a review agent which stops before finishing its files is now asked once to continue.
- Not comparable across tools. Other tools publish numbers on other datasets, subsets, matching rules and models. These numbers say how ocra did on its own cases, not how it ranks.
What this means for you
ocra finds real bugs and so far has rarely reported something wrong, but it missed more issues than it found. Use it as a second reviewer next to people, not as a gate: read what it reports, and do not take an empty report as approval. Its verdict is advice, not a security check (How it works).
Reproduce
The cases and their labels are in evals/golden. With models configured as in Evaluation, a run on the 8 cases above cost $8 to $13 at list prices; the smoke tier now has 10:
node packages/eval/dist/main.js run --dataset golden --tier smoke --label mine --max-cost-usd 20 --pr-max-cost-usd 6