How review works
Review that argues back.
A Hubify review run can assign adversarial roles to challenge a version-pinned claim, record source-cited findings, and track their disposition. Model output is evidence for review, not acceptance: real closure requires version-pinned receipts, explicit truth-audit decisions, and the applicable human sign-off.
The model
Cooperative aggregation blends answers. Adversarial review tries to break them.
Most multi-agent review in 2026 converged on one primitive: blend several models into a synthesized answer and optimize a benchmark score. That’s good at averaging away noise. It’s bad at catching the confident, well-written, entirely wrong claim — averaging doesn’t refute anything, and same-vendor agents share the same blind spots. Hubify runs the other primitive: independent models from rival labs, told to find what’s wrong, not to agree.
Mixture-of-agents · cooperative
- Blend N models into one synthesized answer
- Optimizes a benchmark score
- Same-vendor agents share training priors — errors correlate
- A single vendor can't route to itself as a neutral check
- No verdict, no provenance, no bias guard
Hubify · adversarial
- Independent cross-vendor agents try to refute each other
- Optimizes catching the false positive before it ships
- Different labs, different biases — errors decorrelate
- Model-, harness-, and vendor-agnostic — the neutral referee
- Verdict-first, source-cited, integrity-audited against self-favoring
The workflow
One loop, run until it stops finding things.
Not a single review pass — a cascade. Each round either closes with evidence or feeds the next one.
Multi-vendor round
Claude, GPT, Gemini, Grok, and DeepSeek receive the same claim independently and are instructed to refute it, not confirm it.
Findings
Every objection becomes a logged finding — visible, timestamped, and attributed to the reviewer that raised it. No private disagreements.
Truth-audit verdict
Before anything closes, each finding gets a source-cited verdict: VERIFIED, FALSIFIED, STALE, OUT-OF-SCOPE, or OPINION.
Evidence-required closure
A verified finding only closes against an artifact path and a commit SHA — never a promise to fix it later.
Computed readiness
Readiness is derived from what's still open — BLOCKER, MAJOR, MINOR, CAVEAT counts — never hand-set by whoever wants to ship.
Cascaded rounds
The loop reruns on the updated version until independent vendors converge: a majority return silence, zero regressions, nothing left to argue about.
See it run
One claim, five illustrative review roles.
This is a scripted, hypothetical walkthrough of the interaction pattern, not live model output or a historical review receipt. Real runs must retain version-pinned evidence and human disposition.
Claim under review
Under the paper's stated matter-contraction background and cubic-action assumptions, f_NL(local) = -35/16; mapping that amplitude to survey significance remains conditional on estimator choices, nuisance priors, and systematics.
Reviewer A
illustrative
reviewing…
···Reviewer B
illustrative
reviewing…
···Reviewer C
illustrative
reviewing…
···Reviewer D
illustrative
reviewing…
···Reviewer E
illustrative
reviewing…
···This scripted example demonstrates the workflow, not an actual vendor result. The rejection blocks an observational overclaim while the concern keeps the derived amplitude conditional on its stated assumptions. Real review verdicts require version-pinned receipts and human disposition.
Evidence, not theater
Big Bounce maintains version-specific review artifacts, finding dispositions, and open human gates. Coverage on an earlier artifact does not automatically clear a newer version, and model silence is not journal acceptance. Inspect the canonical lab record for the evidence and scope that exist today.
See the Big Bounce labBring a claim you want tested.
Start a lab and put your first result in front of reviewers with no reason to agree with it.