proof-pack

A smooth, environment-resolved error field underlies the systematic property errors of universal machine-learned interatomic potentials

Universal machine-learned interatomic potentials (uMLIPs) err systematically away from equilibrium; the test is whether those errors share a single, measurable shape.

Narrated summary: A smooth, environment-resolved error field underlies the systematic property errors of universal machine-learned interatomic potentials

Video page · Download MP4 · Download captions

Claim. The systematic errors of universal machine-learned interatomic potentials on fcc metals are projections of a single smooth error field over local atomic environments, with coordination deficit as the leading coordinate.

Foundation machine-learned interatomic potentials (uMLIPs) match near-DFT accuracy on bulk properties but fail on the surfaces, vacancies, and planar faults that dominate real materials practice. The evidence shows that those failures are not independent: they follow a smooth field over local atomic environments. The field is measured from three standard observables, predicts a fourth never-fitted observable with zero adjustable parameters, and converts into a run-time correction that runs beside a live calculator.

Published


Evidence summary

The error landscape

Figure 1: Bulk observables are accurate, but defect-family observables err 15–60× worse per model.

Two parallel rows of four unmarked engineering coupons on stepped mechanical pedestals, both rows preserve the same order while the pedestal heights differ — The pedestal ordering shows rank agreement despite mismatched measured magnitude

Bulk observables (lattice constants, formation enthalpies) are accurate: median relative errors < 0.5 % and ≈ 3 % respectively. Defect-family observables — surface energies, vacancy energies, stacking-fault energies — err 15–60× worse. This defect/bulk asymmetry is the signature the field explains.

Rankings survive where magnitudes fail

Across materials, predicted rankings track reference rankings closely for surfaces (Spearman ρ = 0.88–1.00), vacancies (0.84–0.93), and bulk moduli (0.82–0.85), even while magnitude errors reach tens of percent. Ordinal faithfulness is the invertibility condition for any monotone error model; its selective failure marks where no such model can apply.

The field and its blind test

Figure 4: The environment error field predicts the never-fitted γ₁₁₀ observable with r = 0.906.

A sealed prediction cartridge locked in a cradle outside an independent compression-test chamber, with the specimen entering through a separate hatch — The sealed prediction remains untouched until the independent measurement is complete

The core hypothesis is simple: the model’s energy error is a smooth function of local coordination, accumulated per atom.

E_model(config) − E_ref(config) ≈ Σᵢ Δε(cᵢ), with Δε(12) ≡ 0 for fcc bulk.

Each property samples the field at its characteristic coordinations: γ₁₀₀ at c = 8, γ₁₁₁ at c = 9, vacancy first neighbors at c = 11. The never-fitted γ₁₁₀ error probes coordination 7. Across 36 cells the prediction is r = 0.906 — not because the model is fitted to γ₁₁₀, but because the field shape is measured elsewhere and extrapolated.

From field to run time

Figure 5: The correction recovers fitted observables and improves the blind facet.

Because the field is a function of environments, its inverse is an additive energy with analytic forces. Deployed beside a live CHGNet calculator:

Provable boundaries

Correction has jurisdiction only where order survives. Where rankings invert — for example, MACE-MP-small ordering SFE(Ni) ≤ SFE(Al) while references order the reverse — a machine-checked proof shows that no monotone correction can recover both. The proof kernel certifies data-analysis arithmetic and stated inequalities over SHA-256-provenance data; 899 build-locked Lean 4 theorems, zero sorry.

A compact correction plate inserted between an interatomic-model cartridge and a finite-element fatigue-test fixture for one turbine-blade coupon — The plate applies a bounded local correction before the coupon simulation reaches the fatigue fixture

Conclusion

Conclusion. For fcc metals, uMLIP errors away from equilibrium are not a collection of independent failures. They are the projection of a smooth, environment-resolved error field that can be measured from three observables, predicts a fourth, and runs as a correction beside a live calculator — with machine-checked boundaries on where it does and does not apply.

The field is the first concrete instance of the larger program: measure the structured wrongness of a predictor, prove it, correct it, and make the evidence inspectable.

Audit links