proof-pack
Sharing DFT anchors reduces projected evaluations by 72.4%
A frozen, artifact-backed analysis of 29 analyzable diffusion paths compares four separately evaluated sparse-anchor sets with their shared union.
Reference calculations are the scarce part of many atomistic workflows. This proof pack asks a narrow accounting question: if four models nominate sparse images on the same paths, how many distinct reference evaluations remain after identical image indices are evaluated once and shared?
Abstract
We analyze recorded prediction profiles from four universal machine-learned interatomic potentials on a locked panel of diffusion paths. A frozen rule selects endpoints, a predicted minimum, and a bounded neighborhood around the predicted maximum for each available model–path profile. The separate-model accounting contains 558 projected reference evaluations. Deduplicating equal image indices within each path leaves a union of 154 projected evaluations, or 72.4% fewer DFT evaluations. One panel path is excluded because all four recorded artifacts lack a converged prediction profile; three additional paths retain only their available model records, without imputation. This is prediction-side accounting, not a claim that the shared set has already been evaluated by DFT, not an accuracy result, and not a universal savings rate.
Result
| Accounting basis | Analyzable paths | Projected DFT evaluations |
|---|---|---|
| Four model-specific anchor sets, counted separately | 29 | 558 |
| Within-path union of identical image indices | 29 | 154 |
The percentage is reported exactly as the reviewed public claim. No dollar amount, wall-time conversion, annualization, or additional economic multiplier is inferred in this paper.
Method
For model (m) and path (p), let (A(m,p)) be the anchor indices selected from the recorded predicted image-energy profile. The frozen rule includes both endpoints, the predicted minimum, and a clamped neighborhood around the predicted maximum. The separate-model total counts every available (A(m,p)). The shared total counts each index in the within-path union only once.
The analysis reads the locked panel and four recorded campaign artifacts. It does not rerun DFT. Input SHA-256 digests are stored in the machine-readable record. The accompanying script reproduces the JSON from repository-relative local artifacts or the recorded object locations; a pinned recording timestamp supports byte-for-byte JSON verification.
Exclusions and fail-closed handling
The panel contains 30 paths, but the primary denominator is 29 analyzable paths. One path lacks a converged prediction profile in every model artifact and is excluded rather than imputed. Three paths have partial model coverage; their available profiles are included and their missing profiles remain missing. The record preserves per-model failure entries.
The result is deliberately separate from any realized execution record. This paper does not combine projected and realized denominators, does not substitute one anchor count for another, and does not infer a new percentage from a different campaign.
What the result establishes
- It establishes the projected reference-evaluation count under one frozen selection rule on one locked panel.
- It shows that identical image indices nominated by different model guides can be deduplicated before reference evaluation.
- It provides a machine-readable record, input digests, a deterministic analysis script, and a locked figure.
What the result does not establish
- It does not show that the 154-member union has been evaluated by DFT.
- It does not establish accuracy parity between shared and separately evaluated anchors.
- It does not estimate dollars, wall time, energy use, or emissions.
- It does not claim that 72.4% transfers to other panels, chemistries, engines, model sets, or anchor rules.
- It does not upgrade literature items that remain marked unverified or abstract-only in the underlying research digests.
Historical context
The nudged elastic band method made a discrete chain of images a standard way to optimize minimum-energy paths. Later surrogate-assisted NEB work reduced the number of expensive evaluations required inside a single search. The present result addresses a different reuse boundary: multiple model guides nominate images on the same path, and identical reference-evaluation sites are shared across those guides. The context is established by independent literature; the 558-to-154 measurement itself is supported by the internal auditable record linked below, not by those papers.
Reproduction and audit
From the released data pack, verify all file digests with:
sha256sum -c MANIFEST.sha256
Recompute the union-anchor JSON with the released union_anchor_economics.py, the recorded local inputs, and the pinned --recorded-at timestamp documented in the pack README. Compare the resulting JSON against its SHA-256 sidecar. The proof-pack manifest separately locks the figure input and generated PDF.
Bibliography
- Henkelman, G.; Uberuaga, B. P.; Jónsson, H. A climbing image nudged elastic band method for finding saddle points and minimum energy paths. Journal of Chemical Physics 113, 9901–9904 (2000). DOI: 10.1063/1.1329672.
- Garrido Torres, J. A.; Jennings, P. C.; Hansen, M. H.; Boes, J. R.; Bligaard, T. Low-Scaling Algorithm for Nudged Elastic Band Calculations Using a Surrogate Machine Learning Model. Physical Review Letters 122, 156001 (2019). DOI: 10.1103/PhysRevLett.122.156001.
- Deng, B. et al. CHGNet as a pretrained universal neural network potential for charge-informed atomistic modelling. Nature Machine Intelligence 5, 1031–1041 (2023). DOI: 10.1038/s42256-023-00716-3.
Audit links
- Released data and reproduction pack
- Machine-readable union-anchor record
- SHA-256 sidecar
- Analysis note
- Reproduction script
- Citation identifier audit
- Complete data-pack manifest
Author: Alex Welcing · Institution: Lupine Science · Evidence cutoff: 2026-07-21 · Editorial status: Independently reviewed and approved for publication 2026-08-03.