Every practitioner knows VQE runs get trapped. What is usually unavailable is why — the failure occurs in the prepared state, while hardware returns measurement outcomes rather than a directly inspectable state vector. So we ran the standard benchmark — four molecular systems, hardware-efficient ansatz, 16 restarts each — with an instrument that retains every optimizer-evaluated state exactly, then compared the final states against the exact eigenstates. One system passed. Three failed in three different ways you can inspect, evaluation by evaluation — including a state whose energy lands 0.56% from exact with essentially zero energy variance while containing none of the ground state. This page is the record.
On a quantum computer, energy is estimated by repeated measurement: more shots, tighter error bars. This simulation replays that process against each recorded final state. The estimate always converges — to the state's exact energy (amber). Whether that energy is the truth (cyan) is something the shots can never say, because on a real problem the cyan line does not exist.
This result falsified our own working hypothesis. Before state inspection, the plausible expert diagnosis was leakage onto the near-degenerate triplet band 2.9 mHa above ground. The overlap data ruled it out: the state carries no weight on the band either. It is a superposition living entirely above the four-state window whose energy expectation happens to mimic the ground state. No energy measurement, at any shot count, distinguishes this state from success — it even passes the energy-variance check, because it is an exact eigenstate of the Hamiltonian, merely the wrong one (see below). No plausible guess replaced the measurement. Only state-level inspection settled it.
An exact eigenstate has zero energy variance — ⟨H²⟩ − ⟨H⟩² = 0 — so hardware can, at extra cost, estimate the variance and flag states that merely average to a plausible energy. It is the strongest verification available from measurement statistics, and the correct first objection to everything above. So we computed it, exactly, for every final state:
| system | σ² (Ha²) | σ (Ha) | variance verdict | actual state |
|---|---|---|---|---|
| H₂ · 0.735 Å | 1.4×10⁻¹¹ | ≈0 | certified eigenstate | ground state — correct |
| H₄ · 0.9 Å | 7.77×10⁻² | 0.279 | flagged | HF reference — caught |
| H₄ · 1.5 Å | 6.65×10⁻² | 0.258 | flagged | excited-state mixture — caught |
| H₄ · 2.5 Å | 3.8×10⁻¹⁰ | 2×10⁻⁵ | certified eigenstate | exact excited level — zero ground-state content |
The variance check catches the two sloppy failures — and certifies the mimic. The 2.5 Å state is not merely energy-plausible: it is an exact eigenstate of the Hamiltonian — 95.8% + 4.2% across a degenerate excited level at −1.8617403629 Ha, matching its energy to seven decimals — just not the ground state. It passes the energy test, passes the variance test, and contains 8.6×10⁻¹³ of the answer. The strongest verification hardware can perform blesses the worst failure in the series.
This is our second falsified hypothesis in this study. Before computing, we predicted the opposite on both counts: that variance would flag the superposition-like mimic and miss the near-reference trap. The data reversed both — the HF-shaped trap carries large variance (σ = 0.28 Ha, since the reference state is not an eigenstate of the correlated Hamiltonian), and the mimic carries none. Twice now, the plausible expert guess about these states was wrong until the state itself was inspected. That is the argument for instruments over inference.
Cost note: estimating ⟨H²⟩ for the H₄ Hamiltonian requires sampling 1,774 non-identity Pauli terms versus 185 for the energy itself — roughly a 10× larger measurement campaign, per optimizer evaluation, to run a check that fails on precisely the case that needs it. All variances above are exact inner products, re-derivable from the validation package (pauli_ops.json + stored states, one page of NumPy).
Everything on this page is re-derivable from the public validation package: Hamiltonians and Pauli operators for all four systems, exact FCI energies and state vectors, the winning VQE optimizer path as one QASM 2.0 file per recorded optimizer evaluation, Trotter circuit templates, optimizer logs, the full seed table, and the eigenstate-overlap data behind every fidelity shown above — plus the generator scripts, pinned dependencies, and a Dockerfile for exact environment reproduction. A two-minute quickstart re-verifies the headline H₂ energy and fidelity with nothing but Python, NumPy, and SciPy. CC BY 4.0; per-file checksums included.
SaC is compatible with conventional gate-model quantum workflows. The circuits used in this study are not proprietary IQ Intel-only representations. The architecture processes standard quantum-circuit constructions and established computational elements used in conventional simulation workflows, including QASM-form circuits, parameterized gate sequences, Pauli-operator Hamiltonians, VQE-style optimization paths, and related state evolution or measurement procedures.
The significance is compatibility, not universal simulator equivalence. These studies demonstrate that standard quantum circuits can be processed and independently checked within the SaC architecture while retaining direct access to the resulting state information. They do not establish support for every circuit family, backend, noise model, or workload implemented by conventional quantum simulators.
This provides the bridge to IQ Intel's materials work: standard quantum circuits can enter the same computational architecture that is subsequently used to generate and retain state-resolved material-process data. The material datasets remain separate products with separate physical-validation requirements.
This study establishes the need for state-resolved evidence when endpoint metrics are insufficient to identify the state obtained. The retained H₄ cases show that a plausible energy — and, in the 2.5 Å case, essentially zero energy variance — does not necessarily establish ground-state identity.
It does not validate IQ Intel material datasets, physical-process trajectory generation, or commercial service outputs. The correctness, physical fidelity, predictive value, and validation status of NMC811-IQ, MagNet-IQ, and client-specific datasets must be established separately by their own predeclared validation protocols and evidence.
A quantum computer returns samples, not states. Verifying ground-state content requires quantum state tomography — a cost that grows exponentially with qubit count and destroys the state being measured. The optimizer is structurally blind to every failure on this page: a state that never moved, a 75% head start destroyed, an excited state reported as the answer, a near-perfect energy from an empty state. It navigates by energy, and energy lies.
For this benchmark, the SaC architecture deterministically retains and replays the states evaluated along the conventional VQE optimizer path, with the complete state available at every recorded evaluation — cloneable, inspectable, replayable. Every fidelity on this page is an exact inner product between stored state vectors, computed after the fact, from recorded runs. The verification that costs hardware an exponential experiment is, here, a file read.
This study establishes the information gap: endpoint observables can be insufficient to identify the state actually obtained. It does not validate NMC811-IQ, MagNet-IQ, any client dataset, or IQ Intel physical-process trajectory generation. Those are separate claims evaluated against their own predeclared, dataset-specific physical targets and acceptance criteria. IQ Intel engagements address the demonstrated gap with independently addressable, lineage-connected microstates and multidimensional observable records. No self-registration; submit your details for manual review and we respond when there is a fit.