Signature
c_ab = exp(-(theta_a - theta_b)^2 / (2 * ell^2)); k_ab = sf2 * c_ab
| Inputs | Definition | Unit |
|---|---|---|
theta_a | Value of the calibrated parameter, such as an annual transition probability | parameter units |
theta_b | Another value of the same parameter | parameter units |
ell | Distance in the parameter over which the loss is expected to change appreciably | parameter units |
sf2 | Variance of the Gaussian process prior, the kernel value at zero distance | squared loss units |
c_ab | Correlation of the Gaussian process prior between the losses at theta_a and theta_b | none |
|---|---|---|
k_ab | Kernel value, the prior covariance between the two losses | squared loss units |
Function
Bayesian optimisation of a model calibration loss with a Gaussian process surrogate
Searches a bounded parameter region for the values that minimise a goodness-of-fit loss between model outputs and calibration targets when each model run is expensive. A Gaussian process fitted to the losses already observed gives a normal posterior for the loss at any untried value, and an acquisition function, here expected improvement, chooses the next run. The notation follows the Bayesian Optimisation article and its one-parameter incidence calibration.
Try this function
Implementations
Excel
Squared exponential kernel from named cells
With ThetaA, ThetaB, LengthScale and SFVar named, the formulas return the prior correlation and covariance, held in KernelCorr and KernelCov.
=EXP(-(ThetaA-ThetaB)^2/(2*LengthScale^2)); =SFVar*KernelCorr
Assumptions
Smoothness of the loss set by the kernel
The squared exponential kernel assumes a smooth loss surface; a loss with a sharp minimum, such as an absolute difference, is approximated less well near the minimum.
Kernel parameters fixed or estimated from the runs
The article fixes the prior mean, variance and length-scale for clarity; in practice they are usually estimated from the runs made so far by maximum likelihood or a maximum a posteriori estimate.
Worked examples
Correlation between the two initial runs at 0.01 and 0.06
With a length-scale of 0.02, 2 times 0.02 squared is 0.0008, and the squared distance 0.0025 gives exp(minus 3.125), about 0.0439, a covariance of about 4.3937 with prior variance 100, as in the article.
theta_a = 0.01; theta_b = 0.06; ell = 0.02; sf2 = 100; c_ab = 0.0439; k_ab = 4.3937
Correlation of the candidate 0.04 with the run at 0.01
The squared distance 0.0009 gives exp(minus 1.125), about 0.3247, as in the article.
theta_a = 0.04; theta_b = 0.01; ell = 0.02; sf2 = 100; c_ab = 0.3247; k_ab = 32.4652
Correlation of the candidate 0.04 with the run at 0.06
The squared distance 0.0004 gives exp(minus 0.5), about 0.6065, as in the article.
theta_a = 0.04; theta_b = 0.06; ell = 0.02; sf2 = 100; c_ab = 0.6065; k_ab = 60.6531
Common errors
Length-scale on the wrong scale
A length-scale of 2 for a parameter that runs from 0.01 to 0.06 makes every correlation about 0.9997, so the surrogate is almost flat and the search learns little from each run (computed here for illustration).
Too smooth a kernel for a loss with a sharp minimum
In the article's example one more iteration pushed the posterior mean below zero (about minus 1.09 at 0.030), a value an absolute-difference loss cannot take (computed here for illustration).
Sources
Squared exponential covariance with a length-scale and a variance factor
Rasmussen CE, Williams CKI. Gaussian Processes for Machine Learning. Cambridge, MA: MIT Press; 2006. Chapter 2, Regression. Section 2.2, equation 2.16: cov(f(x_p), f(x_q)) = k(x_p, x_q) = exp(minus half |x_p minus x_q| squared); replacing |x_p minus x_q| by |x_p minus x_q| / l changes the characteristic length-scale, and a positive pre-factor controls the overall variance.
Squared exponential kernel in a lung cancer microsimulation calibration
Gómez-Guillén D, Díaz M, Arcos JL, Cerquides J. Bayesian optimization with additive kernels for a stepwise calibration of simulation models for cost-effectiveness analysis. International Journal of Computational Intelligence Systems. 2024;17:249. doi:10.1007/s44196-024-00646-x. Section 2.3 and section 3: the squared exponential (SE) kernel, k(x, x') = sigma squared times exp(minus ||x minus x'|| squared / 2 l squared), used in Bayesian optimisation (BO-SE) to calibrate a lung cancer Markov microsimulation from a published cost-effectiveness analysis.
Kernel correlation larger for closer points; hyperparameters by maximum likelihood or MAP
Frazier PI. A tutorial on Bayesian optimization. arXiv preprint arXiv:1807.02811 [stat.ML], version 1, 8 July 2018. Section 3: the kernel is chosen so that points closer in the input space have a large positive correlation; section 3.2: kernel and mean parameters are commonly set by maximum likelihood or maximum a posteriori estimates.
Canonical Identity
Stable URI · Machine-readable · Resolvable · CC BY 4.0