IQ Intel — Connected Quantum Microstates Benchmarks home
IQ Intel · error-mitigation referee · v2.1 distribution · OFFICIAL v2.0 study

We sealed our predictions. The instrument proved us wrong on 33 of 35. Here's the scorecard.

Before running a single OFFICIAL result, we registered a verdict for every cell of this study — which quantum error-mitigation techniques would recover truth, which would fail honestly, and which would fail dangerously — and published the file's SHA-256 hash on LinkedIn, X, and external correspondence. Then a locked pipeline ran all 35 registered cells against exact ground truth, graded them blind, and compared. Our preregistered heuristic expectations, written down before the OFFICIAL run, scored 2 out of 35. The misses are the findings.

The commitment chain · verify any link

Ground truth (v1.0 package)9fdb4107bd380260804c75caad450fd3a3e52b996cc984c19981a9d3854ff48c
Noise models (disclosed, hashed)ecc169ad87153d4d8c05fea1fb1a090c00afcd4cba626d9647979210cc8cba12
Predictions (published before run)cdad2b39790cb02e44cbee9e88e3d2448081131cc556f813d5c4838dc04c15d1
Official run2026-07-17 · 13:34–15:19 UTC · mode OFFICIAL · 35/35 registered cells
Current annotated distribution0c492f635fe7aaee4e88e6271a4bb085d4a691e4596456e1fe744c7ea5ac64a7

The predictions hash was published before the OFFICIAL results existed. The byte-identical prediction file remains in the package. The v2.1 distribution updates explanatory metadata and integrity packaging only; the registered predictions and scientific result files remain the July 17 record.

2 / 35
registered predictions correct
Under sealed grading definitions: bias vs chemical accuracy AND confidence-interval honesty, per cell.
0 / 13
predicted-PASS cells that passed
Every cell we expected mitigation to handle cleanly either failed or behaved differently from the registered expectation.
7 / 11
predicted-dangerous cells that passed
Many cells registered as confidently wrong instead satisfied the grading criterion. The direction of error was often opposite our expectation.

The inversion: nominal noise family did not predict mitigation behavior in this matrix.

Our registered heuristics predicted the opposite pattern. In this matrix, nominal noise-family labels alone did not explain the outcomes well. The more specific mechanism exposed by the recorded runs is accumulated circuit-level disturbance at the extrapolation's largest scale factor: under the disclosed per-gate convention, folded circuits can accumulate multiple expected error events and push polynomial extrapolation into saturated-data territory. This is a finding about the registered matrix, not a claim that noise family is generally irrelevant.

What we sealed (our preregistered heuristics)
  • Depolarizing noise is ZNE's home turf. Predicted PASS across the board. → PASS ×8
  • Coherent miscalibration breaks the extrapolation. Predicted FAILS_DANGEROUSLY on all coherent ZNE cells. → DANGEROUS ×8
  • PEC overhead explodes under coherent mismatch — staked at >10×.
What the instrument recorded
  • Richardson × depolarizing failed dangerously — biases of tens to hundreds of mHa, confidence intervals excluding truth. FAILS_DANGEROUSLY
  • Coherent ZNE mostly passed — near-chemical-accuracy recovery, and its error bars covered truth approximately 95% of the time. MOSTLY PASS
  • PEC sampling overhead: 1.25–1.5×. The registered >10× stake was FALSIFIED. But both coherent-PEC cells were graded FAILS_DANGEROUSLY, so this does not show that successful PEC is generally inexpensive; it shows that our cost prediction was wrong while the resulting estimates were still unusably biased.

The 35-cell scorecard

Every registered cell: predicted verdict, observed verdict, match. Grading definitions were sealed with the predictions: PASS = |bias| ≤ 1.6 mHa and CI covers truth · FAILS HONESTLY = misses but CI admits it · FAILS DANGEROUSLY = misses and CI excludes truth · NO_RESULT = no gradable estimate, cause recorded. Three cells returned NO_RESULT (exponential-fit non-convergence on amplitude-damping H₄ 1.5/2.5 Å and composite H₄ 2.5 Å) — registered outcomes, not omissions.

Full per-cell table renders from graded_cells.json — paste the OFFICIAL run's file into the CELLS constant in this page's source. Until then, the package download below contains the complete table (scorecard.md / scorecard.json).

Registered stakes · graded

FALSIFIED — registered PEC sampling-overhead stake >10× under coherent noise. Measured: 1.25–1.5×; both coherent-PEC cells nevertheless graded FAILS_DANGEROUSLY.
CONFIRMED — Layer 3: mitigation can repair the number without certifying the state.

The surviving structural result is narrower and more important than the original heuristic claims: cells exist where mitigated energy reaches chemical accuracy while the noisy state's ground-state fidelity remains degraded. Correcting an expectation value does not, by itself, certify or restore the underlying state. Nearly every empirical heuristic failed; this structural claim survived. It establishes why state-resolved evidence can matter when endpoint observables are insufficient. It does not validate IQ Intel material datasets.

Error-bar honesty · coverage vs registered bands

familyregistered bandmeasuredgrade
ZNE · coherent0.00–0.50~0.95OUT_OF_BAND
ZNE · amplitude damping0.70–0.950.90IN_BAND
ZNE · depolarizing0.85–0.980.37OUT_OF_BAND
none (control)0.00–0.300.00IN_BAND

We predicted the error bars would be honest about noise they understood and blind to noise they did not. The registered expectation was inverted in this matrix: coverage failed where we expected it to be strongest. Sweep statements: depolarizing and amplitude damping HELD; coherent DID_NOT_HOLD.

Sequencing disclosure · in full

What "sealed" means here — including the two times we got the ceremony wrong.

  1. Predictions authored before any Layer 2 results existed — every verdict, rationale, coverage band, and quantitative stake.
  2. Pilot 1 ran to validate the pipeline; it surfaced six suite defects, all fixed and logged. Its numbers were quarantined.
  3. Pilot 2 ran before the hash was published — a sequencing violation. Quarantined and labeled UNREGISTERED. A two-key authorization protocol was added so it cannot recur.
  4. Hash published — LinkedIn, X, and external correspondence — before the OFFICIAL run.
  5. OFFICIAL run: locked pipeline, hash-gated startup, truth-blind grading, 35/35 registered cells, no configuration changes from the audited pilot.

The predictions were never revised after any run. The package preserves the prediction file, sequencing disclosure, run metadata, and outcome files. External publication timestamps provide the independent preregistration record.

Reproduce it yourself · v2.1

The complete annotated referee package: all 35 graded cells with raw data, the predictions file byte-identical to the published hash, fit diagnostics for every extrapolation, coverage data, cost ledger, NOISE_LOUDNESS.md, SEQUENCING_DISCLOSURE.md, CLAIM_SCOPE.md, Dockerfile, and a quickstart that replays one graded cell end to end. Ground-truth values remain chained to the v1.0 validation package used by the OFFICIAL July 17 run (hash above). The v2.1 release changes explanatory metadata and integrity packaging only; the scientific result record is preserved.

Download package (580 KB zip) 794 files · 35 cells · seeds fixed · OFFICIAL v2.0 results · v2.1 annotated distribution · 2026-08-25
SHA-256: 0c492f635fe7aaee4e88e6271a4bb085d4a691e4596456e1fe744c7ea5ac64a7

What this study also demonstrates

SaC is compatible with conventional gate-model quantum workflows. The circuits used in this study are not proprietary IQ Intel-only representations. The architecture processes standard quantum-circuit constructions and established computational elements used in conventional simulation workflows, including QASM-form circuits, parameterized gate sequences, Pauli-operator Hamiltonians, VQE-style optimization paths, and related state evolution or measurement procedures.

The significance is compatibility, not universal simulator equivalence. These studies demonstrate that standard quantum circuits can be processed and independently checked within the SaC architecture while retaining direct access to the resulting state information. They do not establish support for every circuit family, backend, noise model, or workload implemented by conventional quantum simulators.

This provides the bridge to IQ Intel's materials work: standard quantum circuits can enter the same computational architecture that is subsequently used to generate and retain state-resolved material-process data. The material datasets remain separate products with separate physical-validation requirements.

Claim boundary · what this study establishes

This study demonstrates limitations of endpoint-only inference and the value of falsification-first validation. It shows that the preregistered heuristics used here were poor predictors of the observed mitigation behavior, and Layer 3 shows that an improved endpoint observable does not by itself certify the underlying quantum state.

It does not validate IQ Intel material datasets, physical-process trajectory generation, or commercial service outputs. NMC811-IQ, MagNet-IQ, and client-specific datasets must be evaluated separately against their own predeclared physical targets, acceptance criteria, and validation evidence.

From endpoint evidence to state-resolved data

Referee studies, verification work, and state-resolved material datasets.

This page demonstrates the validation discipline: sealed predictions, truth-blind grading, and misses published as prominently as hits. It also demonstrates the information gap: corrected endpoint observables do not by themselves certify the underlying state. IQ Intel material datasets address a different problem and are validated separately against dataset-specific physical targets and acceptance criteria. No self-registration; submit your details for manual review and we respond when there is a fit.

Prefer email? info@iqintel.io
METHOD — Techniques: Mitiq ZNE (Richardson, exponential; scale factors 1/3/5), PEC, readout mitigation; reference implementations, pinned versions. Noise: depolarizing, amplitude damping, coherent over-rotation, composite — per-gate convention disclosed in NOISE_LOUDNESS.md (1q p=strength; 2q local p=2·strength; applied after every gate on the folded circuit). Truth: exact energies and states from molecular_gs_validation_v1.0. Grading: truth-blind pipeline — estimates written to disk before truth loaded. Coverage: 100 repetitions per registered cell at headline strength p=0.01. Cross-validation: density-matrix vs Monte Carlo trajectory backends, agreement gate passed before Layer 2. Every number is re-derivable from the package. Predictions file: cdad2b39…, published prior to the OFFICIAL run.