Squared exponential kernel covariance between losses at two parameter values

Gives the prior covariance between the calibration losses at two parameter values: the prior variance times a correlation that falls with the squared distance between them, scaled by the length-scale. Nearby values are expected to have similar losses. For several parameters the squared difference becomes the squared Euclidean distance.

Signature

c_ab = exp(-(theta_a - theta_b)^2 / (2 * ell^2)); k_ab = sf2 * c_ab
Inputs
InputsDefinitionUnit
theta_aValue of the calibrated parameter, such as an annual transition probabilityparameter units
theta_bAnother value of the same parameterparameter units
ellDistance in the parameter over which the loss is expected to change appreciablyparameter units
sf2Variance of the Gaussian process prior, the kernel value at zero distancesquared loss units
Output
c_abCorrelation of the Gaussian process prior between the losses at theta_a and theta_bnone
k_abKernel value, the prior covariance between the two lossessquared loss units

Function

Bayesian optimisation of a model calibration loss with a Gaussian process surrogate

Searches a bounded parameter region for the values that minimise a goodness-of-fit loss between model outputs and calibration targets when each model run is expensive. A Gaussian process fitted to the losses already observed gives a normal posterior for the loss at any untried value, and an acquisition function, here expected improvement, chooses the next run. The notation follows the Bayesian Optimisation article and its one-parameter incidence calibration.

Try this function

Implementations

  • Excel

    Squared exponential kernel from named cells

    With ThetaA, ThetaB, LengthScale and SFVar named, the formulas return the prior correlation and covariance, held in KernelCorr and KernelCov.

    =EXP(-(ThetaA-ThetaB)^2/(2*LengthScale^2)); =SFVar*KernelCorr

Assumptions

  • Smoothness of the loss set by the kernel

    The squared exponential kernel assumes a smooth loss surface; a loss with a sharp minimum, such as an absolute difference, is approximated less well near the minimum.

  • Kernel parameters fixed or estimated from the runs

    The article fixes the prior mean, variance and length-scale for clarity; in practice they are usually estimated from the runs made so far by maximum likelihood or a maximum a posteriori estimate.

Worked examples

  • Correlation between the two initial runs at 0.01 and 0.06

    With a length-scale of 0.02, 2 times 0.02 squared is 0.0008, and the squared distance 0.0025 gives exp(minus 3.125), about 0.0439, a covariance of about 4.3937 with prior variance 100, as in the article.

    theta_a = 0.01; theta_b = 0.06; ell = 0.02; sf2 = 100; c_ab = 0.0439; k_ab = 4.3937
  • Correlation of the candidate 0.04 with the run at 0.01

    The squared distance 0.0009 gives exp(minus 1.125), about 0.3247, as in the article.

    theta_a = 0.04; theta_b = 0.01; ell = 0.02; sf2 = 100; c_ab = 0.3247; k_ab = 32.4652
  • Correlation of the candidate 0.04 with the run at 0.06

    The squared distance 0.0004 gives exp(minus 0.5), about 0.6065, as in the article.

    theta_a = 0.04; theta_b = 0.06; ell = 0.02; sf2 = 100; c_ab = 0.6065; k_ab = 60.6531

Common errors

  • Length-scale on the wrong scale

    A length-scale of 2 for a parameter that runs from 0.01 to 0.06 makes every correlation about 0.9997, so the surrogate is almost flat and the search learns little from each run (computed here for illustration).

  • Too smooth a kernel for a loss with a sharp minimum

    In the article's example one more iteration pushed the posterior mean below zero (about minus 1.09 at 0.030), a value an absolute-difference loss cannot take (computed here for illustration).

Sources

  • Squared exponential covariance with a length-scale and a variance factor

    Rasmussen CE, Williams CKI. Gaussian Processes for Machine Learning. Cambridge, MA: MIT Press; 2006. Chapter 2, Regression. Section 2.2, equation 2.16: cov(f(x_p), f(x_q)) = k(x_p, x_q) = exp(minus half |x_p minus x_q| squared); replacing |x_p minus x_q| by |x_p minus x_q| / l changes the characteristic length-scale, and a positive pre-factor controls the overall variance.

    View source →

  • Squared exponential kernel in a lung cancer microsimulation calibration

    Gómez-Guillén D, Díaz M, Arcos JL, Cerquides J. Bayesian optimization with additive kernels for a stepwise calibration of simulation models for cost-effectiveness analysis. International Journal of Computational Intelligence Systems. 2024;17:249. doi:10.1007/s44196-024-00646-x. Section 2.3 and section 3: the squared exponential (SE) kernel, k(x, x') = sigma squared times exp(minus ||x minus x'|| squared / 2 l squared), used in Bayesian optimisation (BO-SE) to calibrate a lung cancer Markov microsimulation from a published cost-effectiveness analysis.

    View source →

  • Kernel correlation larger for closer points; hyperparameters by maximum likelihood or MAP

    Frazier PI. A tutorial on Bayesian optimization. arXiv preprint arXiv:1807.02811 [stat.ML], version 1, 8 July 2018. Section 3: the kernel is chosen so that points closer in the input space have a large positive correlation; section 3.2: kernel and mean parameters are commonly set by maximum likelihood or maximum a posteriori estimates.

    View source →

Canonical Identity

Squared exponential kernel covariance between losses at two parameter values | HealthEconomics.wiki