Pushing a Single GPU to Its Limits and Scaling to Tens of Thousands: RL-Guided, Physically Consistent KMC for Nuclear Materials Simulation

Wednesday, June 24, 2026 2:55 PM to 3:15 PM · 20 min. (Europe/Berlin)
Hall E - 2nd Floor
Research Paper
Chemistry and Materials ScienceExtreme-scale SystemsNovel Algorithms

Information

Kinetic Monte Carlo is a cornerstone for rare-event dynamics in materials, but its strict sequentiality and super-basin trapping create a \emph{scalability deadlock}: massive computation advances physical time only marginally, throttling both spatial and temporal scalability.

We present EscapeKMC, the first reinforcement learning–guided KMC framework that resolves the scalability deadlock under strict physical alignment. EscapeKMC introduces three techniques:
(1) \emph{Poisson Clock Alignment} anchors event timing to the Poisson law and enables $O(1)$ escape-time estimation; (2) \emph{Adaptive Swarm Reasoning} distributes decision-making across lightweight atomic agents coordinated by a centralized critic, yielding system-size-invariant policies transferable across scales; and (3) \emph{Sparse Dynamics Regularization} restructures irregular dynamics into batched, accelerator-friendly execution efficiently.

On reactor-pressure-vessel benchmarks spanning concentrations from $10^{-5}$ to $10^{-1}$ and temperatures from 663–773 K, EscapeKMC achieves $17.5\times$ over OpenKMC while maintaining physical fidelity. At extreme scale, it simulates $5.8\times 10^{10}$ atoms on a single NVIDIA A100 GPU with a $500\times$ improvement in memory efficiency. Beyond single-GPU execution, EscapeKMC scales efficiently to large GPU clusters. It sustains $91\%$ parallel efficiency on up to 16{,}384~GPUs and enables KMC simulations of up to $1.7\times10^{14}$ atoms, pushing atomistic irradiation modeling to previously unattainable scales.
Contributors:
Format
on-site