Skip to content
LESSON NOTES · 04

4. Optimizing a trial state without optimizing noise

Position in the course: Lesson 4 of 10. Complete the preceding derivation and use the explained exercises to check understanding.

Prerequisites: VMC; derivatives; covariance

Learning goal: explain the mathematical steps, reproduce the analytical examples, and state the conditions under which the conclusions hold.

1. Parameters affect both the estimator and its distribution

Let a real normalized-in-expectation trial state depend on parameters \(\theta\). Define \(O_k=\partial_{\theta_k}\ln|\Psi_T|\). Changing a parameter changes the probability density and local energy together, so differentiating a fixed sample average while ignoring the density is generally wrong. Starting from the Rayleigh quotient with normalization \(N=\int\Psi_T^2\), self-adjointness gives \(\partial_k\int\Psi_TH\Psi_T=2\int(\partial_k\Psi_T)H\Psi_T\). The quotient rule then yields

\[ \frac{\partial E_V}{\partial\theta_k}=2\left[\langle O_kE_L\rangle-\langle O_k\rangle\langle E_L\rangle\right]. \]

The expression assumes a suitable differentiable real family and boundary conditions that justify moving derivatives and the Hamiltonian. Parameter-dependent nodes or singular terms require care. The covariance form connects optimization directly to sampling statistics.

2. Worked Gaussian energy derivative

For \(\Psi_a=e^{-ax^2/2}\), \(O_a=-x^2/2\) and \(E_L=a/2+(1-a^2)x^2/2\). Thus \(\operatorname{Cov}(O_a,E_L)=-(1-a^2)\operatorname{Var}(x^2)/4\). Since \(\operatorname{Var}(x^2)=1/(2a^2)\), twice the covariance is \((1-a^{-2})/4\), matching direct differentiation of \(E_V=(a+a^{-1})/4\). This identity is a stringent analytical test of signs and factors.

3. Energy and variance are different objectives

Variance minimization uses \(\sigma_L^2=\langle(E_L-E_V)^2\rangle\) to favor eigenstate-like trial functions and stable sampling. Energy minimization targets the lowest expectation in the chosen family. An exact excited state has zero variance, so the variance objective alone does not identify the ground state. Parameters that move nodes can behave differently from positive Jastrow parameters. Check the physical symmetry sector and compare energies with independent validation samples.

4. Linearized wavefunctions and noisy matrices

Around current parameters, expand \(\Psi(\theta+\delta\theta)\approx\Psi+\sum_k\delta\theta_k\partial_k\Psi\). Minimizing the Rayleigh quotient in this local derivative space leads to a generalized eigenvalue problem \(Hc=ESc\), with overlap and Hamiltonian matrix elements estimated stochastically. Nearly dependent derivative directions make \(S\) ill-conditioned. Regularization, parameter scaling, controlled steps, and fresh samples help prevent noise-driven instabilities. This is a conceptual derivation, not a substitute for the conventions of a specific linear-method implementation.

Correlated sampling can compare nearby parameter choices using weights \(|\Psi_{\theta'}|^2/|\Psi_\theta|^2\), but far changes may lose overlap or produce extreme weights. Optimization and final reporting should not reuse the same data without identifying the selection bias. A smooth-looking minimization trace is not a convergence proof.

Selecting the lowest noisy energy among many candidates can bias the selected sample mean downward, even when every exact trial expectation is variational. Evaluate the selected state again with fresh data. Track energy and variance, and compare independent optimizations. Optimization failure is a limitation of the numerical procedure rather than a physical effect. Diagnose sample size, redundant directions, and steps before interpreting differences as wavefunction improvement.

5. Exercises and answers

Exercise: What happens to the covariance gradient for an exact eigenstate?

Solution

\(E_L\) is constant, so every covariance with \(O_k\) vanishes, subject to admissibility. This is consistent with a stationary Rayleigh quotient, but may describe an excited state as well.

Exercise: Why can adding hundreds of poorly scaled parameters hurt optimization?

Solution

Derivative directions can be redundant or noise dominated. Estimated matrices become unstable; more expressiveness needs enough samples, sound regularization, and independent validation rather than just more parameters.

Optimizing a trial state without optimizing noise

Original teaching schematic of the mathematics or algorithm; it is not simulation or experimental data.

6. Sources and connections

Related theory: Classical Monte Carlo · Molecular methods · Density functional theory

Software connection: CP2K · Quantum ESPRESSO · Gaussian

These software courses provide related background on energies, orbitals, or convergence management; they do not imply that the Monte Carlo or QMC examples on this page were executed there.

Quantum mechanics · Molecular methods · Monte Carlo · Molecular dynamics · QMCPACK


Previous lesson · Course overview · Next lesson