OhMyBug
← Back to your landing/How it works
Technical note

How a code change
becomes a finding.

The review separates finding a possible bug from checking that it is real.

SYSTEM OVERVIEW5 MIN READ11 SEPTEMBER 2026
In brief

OhMyBug reviews a specific diff in an isolated cloud workspace. Review passes propose concrete failure scenarios; a skeptic challenges them. Findings that survive go back to your coding agent for a repository check. You approve the verdicts. Tests and the merge decision stay with you.

01 / INPUT

The change sets the scope.

A review starts with the final diff and the context you select: callers, types, configuration, and nearby code. Your agent shows a manifest before sending local files. For a public or connected repository, the service can retrieve the diff and available context at the reviewed ref.

The goal is not to audit every line. It is to understand what this change can break, including effects outside the edited files.

FIG. 01 — THE REVIEW BOUNDARYDIFF → EVIDENCE → VERDICT
YOUR WORKSPACEDiff + context

Your agent prepares the change.

Approved input →
OHMYBUG CLOUDReview → challenge

Review passes propose bugs. A skeptic tests each claim.

Candidate findings →
YOUR WORKSPACEVerify + decide

Your agent checks the evidence. You decide.

REAL / NOT_REAL / UNCLEAR
The cloud produces candidates. The final verdict comes back from your side.
02 / METHOD

Find a failure. Then try to disprove it.

The protocol looks for different kinds of failure. In the orchestrated path, reviewers work independently; other engine paths use separate passes. Engines and models may change, but every finding still has to show a concrete failure.

  • LogicConditions, boundaries, calculations, and defaults.
  • StateRaces, retries, lifecycle, and cleanup.
  • ContractsCallers, data shapes, nulls, and compatibility.
  • Error pathsFailures, authorization gaps, and unsafe inputs.
  • CalibrationClaims, limits, and tests that disagree with the code.

Duplicate candidates are combined. The skeptic checks reachability, guards, and missing context. A plausible story is not enough; a finding needs a concrete scenario grounded in the code.

Context matters. Reviewers may ask for more files. The service retrieves files it can access or asks your agent for approved local files. Missing context limits what we can establish.
03 / OUTPUT

Every finding is something you can check.

Each result names the location, severity, and failure scenario. Your coding agent checks it against the repository and returns REAL, NOT_REAL, or UNCLEAR, with a reason. Finding a possible bug and confirming it are separate steps.

Illustrative finding / not a customer report

Two requests spend the same credit.

Both requests read the balance before either writes the debit. Each proceeds with the same available credit.

LOCATION
Balance check → debit write
TRIGGER
Two concurrent requests on one account
VERIFY
Check whether a transaction or lock makes that sequence impossible.

The standard charge is $10 when a review has at least one confirmed medium-or-higher bug, no matter how many. Minor-only, rejected, unclear, and clean results cost $0. Your first bug-finding review is free. Billing happens when verdicts are submitted.

04 / EXECUTION

An isolated workspace, not open access.

Each review runs in its own VM on Sprites by Fly.io. Network access is limited to model inference and the OhMyBug API. Review code never runs in your production environment.

Fast review

The diff and selected files anchor the review. The reviewers can ask for more context while it runs.

Optional deep review

After a fast review, we may offer a full-repository review when more context could help—for example, after a clean or incomplete run. It needs your approval and repository access. The reviewed diff still sets the scope.

Each run has time and inference budgets. If it stops before there is enough evidence, that is not the same as a clean review.
05 / DATA HANDLING

Temporary input is not zero retention.

We do not keep a permanent repository copy. We keep some review records so you can inspect results and we can investigate failures.

MaterialWhat happens to it
Submitted diff and filesRemoved from the database after the sandbox receives them. A retry may temporarily store them again.
Repository snapshotTemporarily cached while parallel reviews seed their sandboxes. The retention and seeding limits are in the privacy policy.
Run transcriptKept for 7 days after a completed review; 30 days after a failed review.
Report and findingsKept in your review history until account deletion. They can quote code.

We do not use submissions to train models. Inference providers process them under their API terms. See the privacy policy for processors, regions, retention, and deletion requests.

06 / LIMITATIONS

More evidence, not proof.

Reviewers can miss bugs or be wrong. Missing files, time budgets, and model limits affect the result. Independent passes reduce reliance on one reading; they do not remove model errors.

Keep CI, tests, human judgment, and any specialist security review your system needs. Re-review meaningful fixes on the new diff. A clean review means this run found no supported bug—not that none exists.

Ready for a second look?
Send your next diff.

Try the first review free ↗

This note describes the review workflow, not a warranty or benchmark. FAQ · Terms · Privacy policy