ocra

How it works

The review pipeline, from diff to anchored findings.

ocra combines two proven designs: domain-specialized reviewers with a coordinating judge, and deterministic engineering for every step that must not fail.

Ingest → Select → Triage → Bundle → Matrix → Execute → Anchor → Filter → Verify → Judge → Report
StageWhat happens
IngestRead the change set and repository guidelines through a VCS adapter.
SelectDecide per file: review it, or exclude it with a reason (binary, secret, generated, too large…).
TriageAssign a risk tier (trivial, lite, full) from size and sensitive paths such as auth/, PasswordHasher.cs or CI workflows.
BundleGroup related files. Small change sets form one bundle; larger ones are grouped by a light model that answers with file indices. Without that grouping, up to 20 files are reviewed one per task and more by directory.
MatrixDecide which reviewer runs on which bundle from the risk tier and each reviewer's scope (for example, no documentation for a code-only reviewer). Skipped pairs are listed in the report.
ExecuteRun each planned (bundle, reviewer) pair as an isolated agent task with read-only tools.
AnchorResolve every finding's quoted code to exact lines.
FilterDrop findings the repository's memory accepts or a reviewer dismissed, and compare with the previous review of the same pull request, before any model call is spent on them.
VerifyA standard-tier model fact-checks the findings of each file against its diff and the code around them. A finding is dropped only when that code proves it wrong; doubt keeps it, and so does a failed check. Dropped findings and the reason are listed in the report.
JudgeOne top-tier call sees every finding: it merges reports of the same root cause from different reviewers, drops speculation and nitpicks, recalibrates severity, and writes a short summary. Each change carries a reason in the report.

Verdict

The verdict is decided by a fixed rubric over the final findings, not by a model, so the same findings always give the same verdict:

FindingsVerdict
noneapproved
suggestions, or fewer than three warningsapproved_with_comments
three or more warnings, or critical findings the verifier did not confirmminor_issues
any critical finding the verifier confirmedsignificant_concerns

Every finding shows whether Verify confirmed it, was unsure, or did not check it (verification skipped, failed, or out of budget). Unconfirmed critical findings are shown as critical, marked unverified, and cap the verdict at minor_issues, so a single model's unchecked claim cannot block a change. When verification should have run but failed or ran out of budget for a critical finding, the run is incomplete (exit code 3), so it never reads as a pass; with verify: false it is not. The judge can neither drop nor downgrade a confirmed critical finding; an attempt is ignored and listed as a warning.

If the judge is disabled, has no top model or fails, findings are reported unjudged and the rubric still applies.

The verdict is advice, not a security gate. Every model in the pipeline reads the change, and text planted in it can steer a reviewer away from an issue or talk the judge out of one. Use ocra to help reviewers, not to replace required human review or security checks.

Modes

The default mode favors precision. --ultra trades cost for recall:

Default--ultra
ReviewersBy risk tier and scopeEvery reviewer at every tier, including tiers set with reviewers.<id>.minTier (files outside a reviewer's scope are still skipped, and enabled: false still applies)
SamplesOne run per reviewer and bundleTwo runs, merged by fingerprint
PlanOnly for large bundles (5 or more files, or 40,000 characters of diff)Every task. One short model call lists what to check first; if it fails, the task runs without it
CallersReviewers search themselvesUses of the symbols a bundle defines or changes, found outside the bundle, are shown to its reviewers
JudgeDrops speculation and nitpicksKeeps them, marked low confidence; they do not count towards the verdict
CostBaselineAbout twice or more, plus one plan call per task

Verification works the same in both modes.

Reviewers and tools

A reviewer is an agent with a focused prompt, its own rules, and a scope that tells the matrix where it is worth running:

ReviewerLooks forRuns atSkips
correctnessLogic errors, broken contracts, error handlingevery tiernothing
securityExploitable issues reachable from untrusted input: injection, authorization, secrets, crypto, unsafe parsing, CIlite and fulldocumentation, tests
performanceMeasurable regressions on paths that matter: complexity, N+1, blocking work, memorylite and fulldocumentation, tests, config data, CI workflows
docsDocumentation the change makes wrong: a removed or renamed flag, key, endpoint or default that the README, docs or help text still describe (light model)lite and fulltests
agents-mdStatements in AGENTS.md the change makes false: commands, layout, module responsibilities, required tools and variables (light model; only when the repository has AGENTS.md)lite and fulltests

Each one reviews only its own domain and states what not to flag. Every reviewer can only read: read_file, read_diff and code_search answer from the revision under review (a commit in range or commit mode, not your working tree), and findings are submitted through report_finding. Their answers are capped (2,000 characters per line, 50,000 per answer, 2 MB per file read), so a minified bundle cannot flood a model's context. It cannot edit files, run commands or browse the web.

Anchoring

Models are unreliable with line numbers, so they quote code instead. ocra resolves the quote:

  1. normalized match in the file's changed hunks;
  2. match in the whole file;
  3. match in another changed file (the model named the wrong file), only for a whole-line quote found in exactly one place;
  4. a light model is shown the finding and the file's diff and answers with the lines it means, which are matched like any quote (one small call per unmatched finding);
  5. otherwise the finding stays attached to its file, as a file-level finding.

A finding that resolves to a file outside the reviewer's bundle is dropped with a warning: that reviewer was not shown the file.

Whole lines are matched first. Part of a single line counts only when it is at least 12 characters long and only one line contains it: a fragment like err would otherwise land on the first line that happens to contain it. When a quote fits several places equally well (three identical return err; lines), no line is picked; the finding is shown at file level and counted under anchoring.ambiguous. The JSON report (--format json, a session's report.json) gives anchoring for the whole run: findings per method (hunk, file, cross_file, relocated, file_level), the ambiguous ones, and how many relocation calls were made.

Robustness

  • A per-task and a whole-run timeout, enforced even if a model stops responding. A task whose agent shows no progress for five minutes (no new step, text or tool call) is stopped earlier and handed to the next model in the chain.
  • A failed task never fails the run; its files are reported as failed.
  • Model failback chains per tier, with a circuit breaker per model (open after repeated failures, half-open probe after a cooldown).
  • Rate limits are waited out: when a provider says to retry after a short wait (up to 90 seconds), every task pauses that model and retries it. A limit marked as daily, a longer or missing wait, or a fourth limit in a row takes the model out of the chain for the rest of the run, so ocra stops sending requests that would be refused.
  • Every review agent is capped at 30 steps, because each step resends the conversation; most tasks finish in about 15.
  • A review agent that stops before finishing (no done signal, no answer, steps to spare) is told once, in the same session, to finish the files it has not reviewed; the progress line says "resumed after stopping early".
Edit on GitHub

On this page