Signature
z = (L_best - mu_n) / s_n; EI = (L_best - mu_n) * Phi_z + s_n * exp(-z^2 / 2) / 2.5066283
| Inputs | Definition | Unit |
|---|---|---|
L_best | Best (lowest) loss among the runs made; with noisy runs, the lowest posterior mean at the evaluated values | loss units |
mu_n | Predicted loss at the candidate from the Gaussian process | loss units |
s_n | Square root of the posterior variance at the candidate, above 0 | loss units |
Phi_z | Probability that a standard normal variable is at most z, entered from NORM.S.DIST(z, TRUE) or tables | probability |
z | Best loss so far minus the posterior mean, divided by the posterior standard deviation | none |
|---|---|---|
EI | Expected reduction of the best loss from running the model at the candidate | loss units |
Function
Bayesian optimisation of a model calibration loss with a Gaussian process surrogate
Searches a bounded parameter region for the values that minimise a goodness-of-fit loss between model outputs and calibration targets when each model run is expensive. A Gaussian process fitted to the losses already observed gives a normal posterior for the loss at any untried value, and an acquisition function, here expected improvement, chooses the next run. The notation follows the Bayesian Optimisation article and its one-parameter incidence calibration.
Try this function
Implementations
Excel
Expected improvement from named posterior summaries
With BestLoss, PostMean and PostSD named, the first formula returns the standardised gap, held in StdGap, and the second the expected improvement, held in ExpImp.
=(BestLoss-PostMean)/PostSD; =(BestLoss-PostMean)*NORM.S.DIST(StdGap,TRUE)+PostSD*NORM.S.DIST(StdGap,FALSE)
Assumptions
Normal posterior for the loss at the candidate
The loss at the candidate is normally distributed with mean mu_n and standard deviation s_n, as a Gaussian process gives; at an evaluated value without noise s_n is 0 and the expected improvement is 0.
Phi_z evaluated at the z of the same candidate
Phi_z must be the normal distribution function at the z computed from L_best, mu_n and s_n; a value for a rounded or different z changes the result.
Worked examples
Expected improvement at the candidate 0.04 in the article example
The gap 16.1385 minus 17.8394 over 7.3698 gives z of about minus 0.2308, where Phi is about 0.40874; the expected improvement is about minus 0.695 plus 2.863, or 2.1676, as in the article, which rounds z to minus 0.23 and gives 2.17.
L_best = 16.1385; mu_n = 17.8394; s_n = 7.3698; Phi_z = 0.40874; z = -0.2308; EI = 2.1676
Expected improvement at the candidate 0.05 in the article example
With posterior mean 16.651 and standard deviation 4.6028, z is about minus 0.1113 and Phi about 0.45567, so the expected improvement is about 1.5914 (1.59 in the article's table).
L_best = 16.1385; mu_n = 16.651; s_n = 4.6028; Phi_z = 0.45567; z = -0.1113; EI = 1.5914
Expected improvement at 0.035 in the second iteration
After the run at 0.04 the best loss is 3.5167; at 0.035 the posterior mean is 3.7091 and the standard deviation 1.3888, so z is about minus 0.1385, Phi about 0.44491 and the expected improvement about 0.4632, as in the article (0.46).
L_best = 3.5167; mu_n = 3.7091; s_n = 1.3888; Phi_z = 0.44491; z = -0.1385; EI = 0.4632
Common errors
Applying the maximisation form to a loss
Frazier states expected improvement for maximising an objective; used unchanged for a loss, with the gap taken as posterior mean minus best value, it rewards candidates predicted to fit worse and gives about 3.87 instead of 2.17 at 0.04.
Using the best observed loss with a noisy microsimulation
With noisy runs the lowest observed loss is itself uncertain; implementations typically use the lowest posterior mean at the evaluated values in its place.
Reading the candidate with the highest expected improvement as the answer
Expected improvement picks the next run, not the optimum: the run it chose at 0.04 had a loss of 3.52, and the search ends only when a stopping rule or the run budget is reached.
Sources
Expected improvement in closed form, traced to Močkus and Jones and colleagues
Frazier PI. A tutorial on Bayesian optimization. arXiv preprint arXiv:1807.02811 [stat.ML], version 1, 8 July 2018. Section 4.1, equations 7 and 8: expected improvement is the expected value of the positive part of the improvement over the best value observed so far under the normal posterior, evaluated in closed form by integration by parts as in Jones et al. (1998); it was first proposed by Močkus (1975) and is increasing in both the expected gain and the posterior standard deviation. The record applies it to the maximisation of minus the loss.
Noisy evaluations: posterior mean in place of the best observed value
Frazier PI. A tutorial on Bayesian optimization. arXiv preprint arXiv:1807.02811 [stat.ML], version 1, 8 July 2018. Section 5, noisy evaluations: direct use of expected improvement presents conceptual challenges, and authors typically use the maximum of the posterior mean at the previously evaluated points in place of the best observed value.
Expected improvement used to calibrate a lung cancer microsimulation
Gómez-Guillén D, Díaz M, Arcos JL, Cerquides J. Bayesian optimization with additive kernels for a stepwise calibration of simulation models for cost-effectiveness analysis. International Journal of Computational Intelligence Systems. 2024;17:249. doi:10.1007/s44196-024-00646-x. Section 3.2: the acquisition function employed was Expected Improvement, in Bayesian optimisation compared with Nelder-Mead, simulated annealing and particle swarm optimisation for calibrating a lung cancer Markov microsimulation.
Canonical Identity
Stable URI · Machine-readable · Resolvable · CC BY 4.0