How it works
The review pipeline, from diff to anchored findings.
ocra combines two proven designs: domain-specialized reviewers with a coordinating judge, and deterministic engineering for every step that must not fail.
Ingest → Select → Triage → Bundle → Matrix → Execute → Anchor → Filter → Verify → Judge → Report| Stage | What happens |
|---|---|
| Ingest | Read the change set and repository guidelines through a VCS adapter. |
| Select | Decide per file: review it, or exclude it with a reason (binary, secret, generated, too large…). |
| Triage | Assign a risk tier (trivial, lite, full) from size and sensitive paths such as auth/, PasswordHasher.cs or CI workflows. |
| Bundle | Group related files. Small change sets form one bundle; larger ones are grouped by a light model that answers with file indices. Without that grouping, up to 20 files are reviewed one per task and more by directory. |
| Matrix | Decide which reviewer runs on which bundle from the risk tier and each reviewer's scope (for example, no documentation for a code-only reviewer). Skipped pairs are listed in the report. |
| Execute | Run each planned (bundle, reviewer) pair as an isolated agent task with read-only tools. |
| Anchor | Resolve every finding's quoted code to exact lines. |
| Filter | Drop findings the repository's memory accepts or a reviewer dismissed, and compare with the previous review of the same pull request, before any model call is spent on them. |
| Verify | A standard-tier model fact-checks the findings of each file against its diff and the code around them. A finding is dropped only when that code proves it wrong; doubt keeps it, and so does a failed check. Dropped findings and the reason are listed in the report. |
| Judge | One top-tier call sees every finding: it merges reports of the same root cause from different reviewers, drops speculation and nitpicks, recalibrates severity, and writes a short summary. Each change carries a reason in the report. |
Verdict
The verdict is decided by a fixed rubric over the final findings, not by a model, so the same findings always give the same verdict:
| Findings | Verdict |
|---|---|
| none | approved |
| suggestions, or fewer than three warnings | approved_with_comments |
| three or more warnings, or critical findings the verifier did not confirm | minor_issues |
| any critical finding the verifier confirmed | significant_concerns |
Every finding shows whether Verify confirmed it, was unsure, or did not check it (verification skipped, failed, or out of budget). Unconfirmed critical findings are shown as critical, marked unverified, and cap the verdict at minor_issues, so a single model's unchecked claim cannot block a change. When verification should have run but failed or ran out of budget for a critical finding, the run is incomplete (exit code 3), so it never reads as a pass; with verify: false it is not. The judge can neither drop nor downgrade a confirmed critical finding; an attempt is ignored and listed as a warning.
If the judge is disabled, has no top model or fails, findings are reported unjudged and the rubric still applies.
The verdict is advice, not a security gate. Every model in the pipeline reads the change, and text planted in it can steer a reviewer away from an issue or talk the judge out of one. Use ocra to help reviewers, not to replace required human review or security checks.
Modes
The default mode favors precision. --ultra trades cost for recall:
| Default | --ultra | |
|---|---|---|
| Reviewers | By risk tier and scope | Every reviewer at every tier, including tiers set with reviewers.<id>.minTier (files outside a reviewer's scope are still skipped, and enabled: false still applies) |
| Samples | One run per reviewer and bundle | Two runs, merged by fingerprint |
| Plan | Only for large bundles (5 or more files, or 40,000 characters of diff) | Every task. One short model call lists what to check first; if it fails, the task runs without it |
| Callers | Reviewers search themselves | Uses of the symbols a bundle defines or changes, found outside the bundle, are shown to its reviewers |
| Judge | Drops speculation and nitpicks | Keeps them, marked low confidence; they do not count towards the verdict |
| Cost | Baseline | About twice or more, plus one plan call per task |
Verification works the same in both modes.
Reviewers and tools
A reviewer is an agent with a focused prompt, its own rules, and a scope that tells the matrix where it is worth running:
| Reviewer | Looks for | Runs at | Skips |
|---|---|---|---|
correctness | Logic errors, broken contracts, error handling | every tier | nothing |
security | Exploitable issues reachable from untrusted input: injection, authorization, secrets, crypto, unsafe parsing, CI | lite and full | documentation, tests |
performance | Measurable regressions on paths that matter: complexity, N+1, blocking work, memory | lite and full | documentation, tests, config data, CI workflows |
docs | Documentation the change makes wrong: a removed or renamed flag, key, endpoint or default that the README, docs or help text still describe (light model) | lite and full | tests |
agents-md | Statements in AGENTS.md the change makes false: commands, layout, module responsibilities, required tools and variables (light model; only when the repository has AGENTS.md) | lite and full | tests |
Each one reviews only its own domain and states what not to flag. Every reviewer can only read: read_file, read_diff and code_search answer from the revision under review (a commit in range or commit mode, not your working tree), and findings are submitted through report_finding. Their answers are capped (2,000 characters per line, 50,000 per answer, 2 MB per file read), so a minified bundle cannot flood a model's context. It cannot edit files, run commands or browse the web.
Anchoring
Models are unreliable with line numbers, so they quote code instead. ocra resolves the quote:
- normalized match in the file's changed hunks;
- match in the whole file;
- match in another changed file (the model named the wrong file), only for a whole-line quote found in exactly one place;
- a light model is shown the finding and the file's diff and answers with the lines it means, which are matched like any quote (one small call per unmatched finding);
- otherwise the finding stays attached to its file, as a file-level finding.
A finding that resolves to a file outside the reviewer's bundle is dropped with a warning: that reviewer was not shown the file.
Whole lines are matched first. Part of a single line counts only when it is at least 12 characters long and only one line contains it: a fragment like err would otherwise land on the first line that happens to contain it. When a quote fits several places equally well (three identical return err; lines), no line is picked; the finding is shown at file level and counted under anchoring.ambiguous. The JSON report (--format json, a session's report.json) gives anchoring for the whole run: findings per method (hunk, file, cross_file, relocated, file_level), the ambiguous ones, and how many relocation calls were made.
Robustness
- A per-task and a whole-run timeout, enforced even if a model stops responding. A task whose agent shows no progress for five minutes (no new step, text or tool call) is stopped earlier and handed to the next model in the chain.
- A failed task never fails the run; its files are reported as
failed. - Model failback chains per tier, with a circuit breaker per model (open after repeated failures, half-open probe after a cooldown).
- Rate limits are waited out: when a provider says to retry after a short wait (up to 90 seconds), every task pauses that model and retries it. A limit marked as daily, a longer or missing wait, or a fourth limit in a row takes the model out of the chain for the rest of the run, so ocra stops sending requests that would be refused.
- Every review agent is capped at 30 steps, because each step resends the conversation; most tasks finish in about 15.
- A review agent that stops before finishing (no done signal, no answer, steps to spare) is told once, in the same session, to finish the files it has not reviewed; the progress line says "resumed after stopping early".