1. Probability, integration, and random estimators
Position in the course: Lesson 1 of 10. Complete the preceding derivation and use the explained exercises to check understanding.
Prerequisites: Calculus; sums and integrals
Learning goal: explain the mathematical steps, reproduce the analytical examples, and state the conditions under which the conclusions hold.
1. Distinguish a state from its probability density
A Monte Carlo calculation replaces a difficult integral by an average over random configurations. The randomness belongs to the numerical procedure; it does not change the physical model. For a discrete state \(x\), \(p(x)\) is a probability. For a continuous coordinate, \(p(x)\) is a density and \(p(x)dx\) is a probability. A density may exceed one because it has units inverse to its coordinate. Always specify the integration measure before defining a target.
Normalization and expectation mean
For independent identically distributed samples, linearity gives \(\mathbb E[\bar A]=\mu\). Expand the centered square: \(\operatorname{Var}(\bar A)=M^{-2}\sum_{ij}\mathbb E[(A_i-\mu)(A_j-\mu)]\). Independence makes the \(i\ne j\) terms vanish, leaving \(\sigma_A^2/M\). Finite variance is essential; a normal-looking histogram does not establish it. The law of large numbers concerns convergence; the central limit theorem supplies the usual approximate Gaussian uncertainty under additional conditions.
2. A worked integral and its exact variance
Take \(x\) uniform on \([0,1]\) and \(A=x^2\). Then \(I=\int_0^1x^2dx=1/3\). The second moment is \(1/5\), so
At \(M=1000\), the expected standard error is approximately \(0.00943\). This is an analytical planning value, not an executed random experiment. Four times as many independent samples halve the error. In higher dimensions the familiar \(M^{-1/2}\) scaling survives when the variance remains finite, but variance can grow drastically with dimension. Monte Carlo avoids a regular grid's exponential point count; it does not make every high-dimensional integral cheap.
3. Transform coordinates without losing the measure
For \(x=h(u)\), the change of variables is \(dx=|h'(u)|du\). A radial integral in three dimensions contains \(4\pi r^2dr\); sampling radius uniformly is not sampling a sphere uniformly. If \(u\) is uniform on \([0,1]\), \(r=Ru^{1/3}\) yields uniform volume because \(P(r<a)=(a/R)^3\). Angular variables require their own correct measure. This distinction returns in volume moves, molecular orientations, and electronic radial densities.
4. A numerical contract
Before implementing a sampler, write down the target, support, observable, units, and one known integral. Keep random seeds for reproducibility, but do not infer independence merely from distinct seeds. Pseudorandom generators are deterministic algorithms whose statistical quality matters. Separate random-number quality from physical-model and discretization errors. A deterministic quadrature check on a one-dimensional model is often more informative than a large unexplained stochastic output.
A confidence interval describes repeated-sampling behavior under stated assumptions. It is not automatically a posterior probability for the unknown value. Unknown variance is usually estimated with the sample variance using the M minus one denominator, and that estimate itself is uncertain. Repeatability across a few seeds can coexist with missed rare events. Check a generator against support and known moments as well as the mean.
5. Exercises and explained solutions
Exercise: For uniform \(x\in[-1,1]\), estimate \(\int_{-1}^1x^2dx\). Is averaging \(x^2\) enough?
Solution
The density is \(1/2\), hence \(\mathbb E[x^2]=1/3\). The integral is the interval length times the expectation, \(2/3\). Forgetting the factor two confuses a normalized expectation with an unnormalized integral.
Exercise: Why is a sample mean of a heavy-tailed variable potentially unsafe?
Solution
If the expectation or variance diverges, the preceding standard-error derivation fails. Increasing the sample count cannot justify a finite-variance formula; examine tails and choose a better estimator or target.
Original teaching schematic of the mathematics or algorithm; it is not simulation or experimental data.
6. Sources and connections
Related theory: Molecular dynamics · Quantum Monte Carlo
Software connection: CP2K · Quantum ESPRESSO
These software courses provide related background on energies, orbitals, or convergence management; they do not imply that the Monte Carlo or QMC examples on this page were executed there.
7. Related theory and practice
Molecular dynamics · Quantum Monte Carlo · RASPA3