Research note

The Correction That Hurt, and the Theorem That Stopped It

A pre-registered, two-round trial of raw vs. Lupine-corrected uMLIPs on two of the five portfolio targets, with a LAMMPS classical leg and kernel-checked correction licenses


If a correction layer is worth trusting, it must survive a trial designed to let it fail. We pre-registered one: two candidate groups from the five-target portfolio, four foundation potentials, a classical LAMMPS baseline, disjoint calibration sets, and success criteria written down before any GPU cycle ran. Round 1 falsified our own first correction. Round 2, licensed by the theorem that failure produced, is the one we are keeping.

What we asked

Can a correction layer improve single-model predictions on materials it has never seen — and can a reference-free gate tell, in advance, which predictions deserve refusal? We picked two groups where the local instrument is honest: fcc Cantor-subset alloys (CoCrNi, CoCrFeNi, with CoNi and FeNi anchors — the one class with a true classical LAMMPS leg, via the shipped 2NN-MEAM potential) and lead-free halide perovskites (CsSnCl₃, CsSnBr₃, CsSnI₃, CsGeI₃, with CsPbI₃ as control). MOF sorbents, LMR cathodes, and ammonia catalysts stayed out: their failure modes are barrier and framework problems the local cubic-statics lane cannot measure without bias, and an unbiased trial does not force off-instrument classes into it.

Every reference value was locked before the run — experimental where it exists (single-crystal constants for CoCrNi from resonant ultrasound; tetrataenite lattice constants), DFT where it does not, unfindable values recorded as null rather than invented. The instrument: 108-atom random-solid-solution supercells (seed 20260713), Birch–Murnaghan E–V relaxation, relaxed-ion stress–strain elastic probes at ±0.5 % strain, identical code paths for every arm. The MEAM leg runs the same cells through the same probe functions with LAMMPS supplying forces — the integration we validated earlier the same day on real Ni/EAM physics, exe and python-module legs agreeing to 10⁻⁹ GPa.

What did not survive

The cross-class correction. Round 1’s de-bias arm divided each prediction by the model’s median prediction-to-reference ratio, learned on the elemental fcc metals of our 21-material baseline — a set disjoint from every candidate. On the perovskites it helped marginally. On the alloys it hurt: median bulk-modulus error went 9.1 % → 16.9 %. The mechanism is plain once seen: the potentials underpredict elemental fcc B₀ (CHGNet’s median ratio is 0.856), so the correction inflates — but on the alloys the raw predictions already sat above their references. A correction with the wrong sign is not a weaker correction; it is a multiplier on the error.

A defect-rich ceramic surface beneath a scanning probe, with one smooth polished region ending sharply at a cracked rough region — The probe maps where a smooth correction field ceases to be supported at the defect cliff

The synthetic demo. Our LAMMPS-bridge demo theorems had been generated from a fabricated log whose C11 sat comfortably inside tolerance. Real Ni/EAM physics gives C11 = 233.3 GPa against the 246.5 GPa experimental reference — outside the 5 % gate — and the regenerated module now carries an honest exceeds_tol theorem where the synthetic one flattered.

The theorem in the loop

The Round-1 failure is now a law, not an anecdote. Two lemmas in Shapes/Certificates.leanwrong_direction_inflation_worsens and its deflation twin — prove that a correction applied against the true error direction strictly increases error, with the Round-1 alloy numbers instantiated as a kernel-checked witness. A decidable directionVerified predicate turns that into a license: a correction may act only where in-class evidence shows the calibration and target err the same way. No license, no correction — abstention is provably risk-free.

What survived

Round 2, direction-gated — and honestly exploratory. With corrections gated per (model, property, class) by leave-one-out evidence inside each candidate group — the held-out candidate never sees its own reference — the rule applied 69 corrections and abstained 71 times. Where it acted: alloy lattice error 0.84 % → 0.17 %, perovskite lattice error 1.41 % → 0.48 %. Where evidence was absent or inconsistent it abstained. Two corrections from our own review temper this. First, Round 2’s rule was chosen after Round 1 failed and evaluated on the same nine candidates: it is exploratory rule selection, not a completed trial, and it is now frozen verbatim in the Round-3 preregistration for evaluation on candidates it has never seen. Second, the gate is weaker than the theorem it invokes: the law requires the target’s error direction, the rule checks only its classmates’ — and one alloy cell (FeNi B₀) walked through that gap and received the exact wrong-direction correction the theorem forbids. The gate provably excludes one failure mode on calibration members; the magnitude cap that closes the remaining gap is registered for Round 3. Part of the lattice gains also absorb a known reference-convention offset (high-temperature lattice references against athermal predictions) — the decomposition is registered work.

The gates as selective prediction — split by class. Across all nine candidates, the reference-free verdicts refused five. Pooled, the refused candidates carry higher raw error than the issued ones — but that pooled ratio is mostly class composition (alloys are easier than perovskites), and within class the picture is humbler: the perovskite gate barely separates, and among the alloys the gate refused the most accurate candidate. One refusal was a false positive on our own control: CsPbI₃ was refused because CHGNet’s predicted tensor is Born-unstable — defensible physics for a 0 K cubic phase, but a refusal of a subject we designated known-good, and we count it as such. The refusals that carry real content are the physics-gated ones: LiS rocksalt fails Born under every model, and no calibration choice can change that.

The reproducibility rerun. The full Round-1 measurement repeated end-to-end returns all nine verdicts identically, with worst-case raw-value drift of 8×10⁻⁴ relative (GPU float nondeterminism amplified through finite differences). The instrument is verdict-stable.

The defect cliff

The five-materials brief names the perovskite failure mode precisely: tin vacancy formation, at under-coordinated sites the potentials were not trained to see. We measured it. Neutral B-site vacancies, metal-rich limit, 2×2×2 supercells, four models: cross-model spreads on vacancy formation energy run 0.5–1.8 eV — at the brief’s own conversion (100 meV ≈ 50× in rate), ranking candidates on any single raw potential here is meaningless. Relative to an energy observable of the same kind, defect disagreement runs 2–34× the bulk-modulus disagreement on the same compounds (comparisons against lattice-constant dispersion inflate the ratio and we do not headline them). CHGNet cannot form the α-Sn reference at all — so the tin rows are a three-model, single-family spread — and predicts negative vacancy formation in CsGeI₃, a cell we flag as invalid rather than average: spontaneous self-destruction of the lattice, the softening signature in its purest form. The literature consensus agrees with the mechanism (the tin vacancy is the accepted oxidation gateway) but publishes its numbers in the halide-rich, ionized-defect convention; our metal-rich neutral panel is therefore graded on disagreement, not on references, and says so.

A cantilever coupon in a displacement-test jig with three adjacent physical pointers showing raw prediction, corrected prediction, and measured deflection without scales — The pointers compare whether correction moves the prediction toward the measured deflection

Boundaries

Nine candidates demonstrate a method, not a survey. Random-solid-solution cells are only statistically cubic. Alloy B₀ references derive from room-temperature polycrystal moduli set against athermal predictions — part of the residual there is convention, not physics. The vacancy panel is neutral-only, single supercell size, no charge corrections. Two of the four potentials share architecture and training data, so dispersion understates independent error. And no per-property thresholds exist yet for defect energies — the gate that would refuse a vacancy prediction is the next thing to calibrate.

Registered next — and what happened when we ran it

Round 3 froze the direction-gated rule verbatim, added a magnitude cap, and evaluated it on eight candidates it had never seen — four alkali halides with primary-source elastic references, four out-of-sample perovskites. Kill condition, registered in advance: if the frozen rule fails both groups, the correction layer’s scope claim narrows to same-class lattice constants, in all public material.

It failed both groups, and the kill fired. One property survived, decisively: lattice constants, corrected 1.60 % → 0.33 % (ionics) and 1.75 % → 0.74 % (perovskites) on candidates the rule never saw, with the gate confident enough to fire on every cell. On every other property the gate abstained on most cells — correctly — and the few firings that slipped past the frozen cap made predictions worse, exactly the behavior the stronger cap theorem says the frozen cap permits. So the claim this program now makes is the one that survived: for lattice constants, within a chemical class, a direction-gated correction transfers out of sample; for everything else, the honest verdict is abstention — and we can prove why.

Evidence trail: preregistration and both rounds in the rhizo repository (docs/plans/2026-07-13-unbiased-accuracy-campaign.md, data/candidates/round1, round2, round1_lammps, round2-verify, perovskite_vacancy_panel); thresholds and re-verdicts under data/discovery_gates and data/climate_targets/halide_panel; correction-direction laws, threshold-migration laws, and ensemble-hull refusals in lean-spec/LupineEvidence/Shapes/Certificates.lean and Discovery/Certificates_LiS_V2.lean, all kernel-checked by lake build LupineEvidence, zero sorry; ledger claims landed under discovery_gates_* and correction_direction_validation_*, agent local-gpu-discovery-lane.

A safer pilot program: an unsupported candidate remains at the laboratory boundary while verified paths continue toward battery and structural test rigs