We ran a security audit on ourselves before launch. It found a critical.
Hypervex is a security tool. Before letting anyone else install it, we pointed the same scrutiny at our own codebase that we ask customers to accept from us. The audit found a critical cross-tenant vulnerability in our GitHub install flow — one that would have let any signed-up user read another organisation's private source code.
We fixed it before a single customer existed. Then, over the following week, four more bugs surfaced. This is the write-up, including the parts that are unflattering, because a security vendor that only publishes its wins is asking you to trust a filtered view.
The critical: anyone could claim your installation
When you install a GitHub App, GitHub redirects you back to the app's setup URL with an installation_id in the query string. Our callback took that number and wrote a row binding the installation to whoever was logged in.
It never checked whether that person had anything to do with the installation.
What made this hide in plain sight is that it looked defended. The column had a uniqueness constraint, so you could not steal an installation that already existed — the obvious attack fails, which is exactly why nobody looked harder. But you could claim one that did not exist yet:
- Sign up for the free tier. Install the app on a throwaway organisation to learn roughly where GitHub's installation-id counter currently sits.
- Call the callback repeatedly for a block of ids that have not been allocated yet. Each one writes a row claiming that future installation for you. They are cheap authenticated GET requests.
- Wait. Someone installs Hypervex and is assigned one of them.
- Our webhook handler finds the pre-existing row and attaches the new customer's repositories to your installation.
- Every dashboard query scopes by installation owner — correctly — and therefore serves you their private repository names, their security findings, and the code snippets inside those findings.
Cross-tenant source code disclosure, in a product whose entire purpose is to be trusted with source code. The victim would have seen a confusing error on their own install and no reason to suspect anything worse.
The fix, and the fix we rejected
The callback now proves ownership with GitHub before writing anything: it asks GET /user/installations with the installing user's own token and refuses unless the installation is in the list. Every uncertain path fails closed — a stale token or a GitHub outage is not evidence of ownership.
The tempting fix was a signed state nonce, and it is worth explaining why we did not ship that alone: a nonce binds the user, not the installation. An attacker obtains a perfectly valid nonce for themselves, then replays it with a substituted installation id. It is a reasonable CSRF defence and it would have closed nothing here. Mitigations that address the adjacent problem are how vulnerabilities survive review.
Then we started actually running things
The audit covered our web surface: routes, authorisation, dependencies, secrets. It came back clean apart from that one finding. Then we began exercising the product end to end in production — installing on a real repository, opening real pull requests, deliberately breaking things — and four more bugs appeared within a day.
Every one of them was in code that already had tests.
A whole language quietly unscanned
Our lockfile list mixed capitalisation — cargo.lock lowercase, Gemfile.lock capitalised — while the three places that used it disagreed about whether to normalise case first.
The effect: Rust lockfiles were never scanned for known vulnerabilities. No error, no warning. The dependency scanner simply reported nothing, which is indistinguishable from a project that has no vulnerable dependencies. A Rust team could have used Hypervex for months and concluded their dependencies were clean.
Silent wrong answers are the worst failure mode a security tool has. A crash gets investigated; a clean report gets believed.
A merge gate that could never be cleared
Hypervex can block a merge by posting a GitHub check run. Two separate defects combined into something considerably worse than either.
First, a review that failed after exhausting its retries left its check run marked in progress forever. The code to conclude it correctly existed and was tested — it was simply never called from anywhere in production. Second, we only handled the opened webhook event, not synchronize, so pushing new commits to a pull request triggered no new review at all.
Put together: Hypervex blocks your merge over a genuine vulnerability. You fix it and push. No new review runs, so no check appears on the new commit, so the required check never passes. Your pull request is now permanently unmergeable, and the failure comment helpfully advised pushing a new commit to retrigger the review — describing something that could not happen.
We found this because our own API credit ran out mid-review. An accident exercised the failure path, which no test had.
The free tier delivered half of what we advertised
Our pricing page says ten reviews a month on the free tier. It was five.
Two separate places incremented the same usage counter for a single review: a leftover from before we built the quota service, and the quota service itself. Nobody noticed because nobody had run ten reviews in a calendar month yet. Customers would have hit their limit at half the advertised allowance and had every reason to assume we were being sharp about it.
A transient error became a permanent downgrade
Our embedding provider enforces a per-minute quota. We had no retry, so a momentary rate limit meant the review continued with no codebase context — the exact capability that distinguishes Hypervex from a diff-only reviewer. To its credit the product already disclosed this in the review comment rather than pretending otherwise. It still should have waited two seconds and tried again.
What we take from this
A mitigation that solves the adjacent problem is worse than none, because it stops the search. The uniqueness constraint genuinely prevented one attack and concealed a better one.
Tested is not the same as run. Four of these five bugs were in code with passing tests. The tests encoded what we believed the code did. Production disagreed within a day.
Silence is the dangerous failure. Unscanned Rust, halved quota, dropped context — none of them errored. Each produced a plausible-looking result that was wrong. We now treat "found nothing" as a claim needing evidence rather than a happy path.
Why publish this
Every security vendor publishes the benchmark they win. We are also sitting on a favourable benchmark result we have refused to quote, because its methodology has an unresolved confound we have not eliminated yet. Same instinct, opposite direction: if we only tell you the good things, our good things are worth less.
These bugs were found before we had customers because we went looking before we had customers. That is the standard we would want from a tool we let read our own source code.
Found something in Hypervex we haven't? admin@hypervex.ai. We will credit you and write it up the same way.