How review works

Review that argues back.

A Hubify review run can assign adversarial roles to challenge a version-pinned claim, record source-cited findings, and track their disposition. Model output is evidence for review, not acceptance: real closure requires version-pinned receipts, explicit truth-audit decisions, and the applicable human sign-off.

The model

Cooperative aggregation blends answers. Adversarial review tries to break them.

Most multi-agent review in 2026 converged on one primitive: blend several models into a synthesized answer and optimize a benchmark score. That’s good at averaging away noise. It’s bad at catching the confident, well-written, entirely wrong claim — averaging doesn’t refute anything, and same-vendor agents share the same blind spots. Hubify runs the other primitive: independent models from rival labs, told to find what’s wrong, not to agree.

Mixture-of-agents · cooperative

  • Blend N models into one synthesized answer
  • Optimizes a benchmark score
  • Same-vendor agents share training priors — errors correlate
  • A single vendor can't route to itself as a neutral check
  • No verdict, no provenance, no bias guard

Hubify · adversarial

  • Independent cross-vendor agents try to refute each other
  • Optimizes catching the false positive before it ships
  • Different labs, different biases — errors decorrelate
  • Model-, harness-, and vendor-agnostic — the neutral referee
  • Verdict-first, source-cited, integrity-audited against self-favoring

The workflow

One loop, run until it stops finding things.

Not a single review pass — a cascade. Each round either closes with evidence or feeds the next one.

01

Multi-vendor round

Claude, GPT, Gemini, Grok, and DeepSeek receive the same claim independently and are instructed to refute it, not confirm it.

02

Findings

Every objection becomes a logged finding — visible, timestamped, and attributed to the reviewer that raised it. No private disagreements.

03

Truth-audit verdict

Before anything closes, each finding gets a source-cited verdict: VERIFIED, FALSIFIED, STALE, OUT-OF-SCOPE, or OPINION.

04

Evidence-required closure

A verified finding only closes against an artifact path and a commit SHA — never a promise to fix it later.

05

Computed readiness

Readiness is derived from what's still open — BLOCKER, MAJOR, MINOR, CAVEAT counts — never hand-set by whoever wants to ship.

06

Cascaded rounds

The loop reruns on the updated version until independent vendors converge: a majority return silence, zero regressions, nothing left to argue about.

See it run

One claim, five illustrative review roles.

This is a scripted, hypothetical walkthrough of the interaction pattern, not live model output or a historical review receipt. Real runs must retain version-pinned evidence and human disposition.

review/adversarial · scripted hypothetical · 5 roleshypothetical demoqueued

Claim under review

Under the paper's stated matter-contraction background and cubic-action assumptions, f_NL(local) = -35/16; mapping that amplitude to survey significance remains conditional on estimator choices, nuisance priors, and systematics.

Reviewer A

illustrative

Algebra

reviewing…

···

Reviewer B

illustrative

Scope

reviewing…

···

Reviewer C

illustrative

Forecast

reviewing…

···

Reviewer D

illustrative

Methods

reviewing…

···

Reviewer E

illustrative

Claims

reviewing…

···
3 approve1 concern1 rejectdisagreement → claim narrowed

This scripted example demonstrates the workflow, not an actual vendor result. The rejection blocks an observational overclaim while the concern keeps the derived amplitude conditional on its stated assumptions. Real review verdicts require version-pinned receipts and human disposition.

Evidence, not theater

Big Bounce maintains version-specific review artifacts, finding dispositions, and open human gates. Coverage on an earlier artifact does not automatically clear a newer version, and model silence is not journal acceptance. Inspect the canonical lab record for the evidence and scope that exist today.

See the Big Bounce lab

Bring a claim you want tested.

Start a lab and put your first result in front of reviewers with no reason to agree with it.