mathematics//optimization//Bayesian optimization
Bayesian optimization is a method for minimizing a function that is expensive to evaluate and offers no usable gradient, by fitting a probabilistic model of the function to the points tried so far and choosing the next point where that model expects the most useful result; it is used to tune a handful of parameters when every trial costs a bench run, a test flight, a batch of product or hours of training. The model is called the **surrogate**, and it is usually a Gaussian process, which gives at every untried point both a predicted value and an uncertainty.
Bayesian optimization is a method for minimizing a function that is expensive to evaluate and offers no usable gradient, by fitting a probabilistic model of the function to the points tried so far and choosing the next point where that model expects the most useful result; it is used to tune a handful of parameters when every trial costs a bench run, a test flight, a batch of product or hours of training. The model is called the surrogate, and it is usually a Gaussian process, which gives at every untried point both a predicted value and an uncertainty.
Picture three gains of an attitude controller tuned on a drone test stand, where each trial is a step test that drains a battery and risks the airframe (PID tuning). A grid of ten values per gain is a thousand trials; a person tuning by hand might take fifty and never know how close they came. Bayesian optimization keeps a belief about the whole response surface and asks, at each step, which single trial is worth running next.
1Fit the surrogate to every trial so far2Find the point with the best acquisition value3Run that trial on the real system4Add the result to the data
The choice of the next point is made by an acquisition function, cheap to evaluate on the surrogate. Expected improvement, the most common, scores each candidate by how much it is expected to beat the best result so far, which rewards both points the surrogate predicts to be good and points where it knows little. That balance is the exploration-exploitation trade-off solved one trial at a time.
It pays where trials are few and dear. Fitting a Gaussian process costs a time cubic in the number of trials, trivial for a few hundred, so the method spends computation to save experiments. It works best with up to roughly a dozen parameters; beyond that the surrogate needs too many trials to say anything (curse of dimensionality). When gradients are cheap, gradient descent wins; when trials are cheap, random search is simpler and nearly as good.
On hardware, safety comes first. A surrogate that expects improvement in an unexplored corner will send the drone there, so the search box is bounded by gains known to be stable, and safe variants only try points the model predicts to be safe with high probability. Measured costs are noisy, and the Gaussian process carries that noise as a term of its own instead of chasing it.
The word Bayesian refers to the belief about the objective, updated after every trial; the parameters being tuned get no posterior of their own. The book counts the method as a niche in control, while machine learning uses it routinely for hyperparameters.