Bayesian methods for performing inference using neural networks

Tom Charnock

Institut d'Astrophysique de Paris


Sorbonne Université
ANR
IAP
CNRS
Aquila

Neural networks for parameter regression

What is a neural network?


An arbitrary, non-linear function $(\mathscr{f}:\mathbb{R}^{\bf d}\rightarrow\mathbb{R}^\boldsymbol{t})$ with fittable parameters $(\boldsymbol{w})$

Let's talk about parameter estimation by regression


Take in data and predict model parameters


Only provide intrinsically biased estimates!

Some stats

A model

The generator $({\bf d}\in\mathcal{M}(\boldsymbol{\theta}))$ of data $({\bf d})$ with model parameters $(\boldsymbol{\theta})$

The likelihood and the posterior


Likelihood $\mathcal{L}({\bf d}|\boldsymbol{\theta}^*)$
What is the probability that the model generates data ${\bf d}$ given a set of model parameters $\boldsymbol{\theta}^*$?

Posterior $\mathcal{P}(\boldsymbol{\theta}|{\bf d}^*)$
What is the probability that model parameters $\boldsymbol{\theta}$ generate the data ${\bf d}^*$?

Fisher information


How much information does the data ${\bf d}$ contain about the model parameters $\boldsymbol{\theta}$?

$${\bf F}_{\alpha\beta} = \left.\left\langle\frac{\partial^2\ln\mathcal{L}({\bf d}|\boldsymbol{\theta})}{\partial\theta_\alpha\partial\theta_\beta}\right\rangle\right|_{\boldsymbol{\theta}=\boldsymbol{\theta}^*}$$
For a Gaussian
$${\bf F}_{\alpha\beta} = \frac{\partial\mu}{\partial\theta_\alpha}^T{\bf C}^{-1}\frac{\partial\mu}{\partial\theta_\beta}.$$

Maximum likelihood estimates


Which value of the model parameters $\boldsymbol{\theta}$ is most likely to generate the observed population of data $\{{\bf d}\}$?
$$\mathcal{L}(\{{\bf d}\}|\boldsymbol{\theta}_\textsf{MLE})=\underset{\boldsymbol{\theta}\in\boldsymbol{\Theta}}{\textsf{Sup}}~\mathcal{L}(\{{\bf d}\}|\boldsymbol{\theta})$$

Intrinsically biased estimators


Intrinsically biased estimators

A summary of the parameters which is conditioned on the wrong likelihood.

Training a network


Neural network is a model $\mathcal{M}(\boldsymbol{w}, {\bf d})$ with parameters $\boldsymbol{w}$ and initial conditions ${\bf d}$.

Training amounts to finding the maximum likelihood estimates of the weights for a chosen cost function $\boldsymbol{l}$.




$$\mathcal{L}(\boldsymbol{l}\,|\boldsymbol{w}_\textsf{MLE}, {\bf d})=\underset{\boldsymbol{w}\in\boldsymbol{W}}{\textsf{Sup}}~\mathcal{L}(\boldsymbol{l}\,|\boldsymbol{w}, {\bf d})$$

Why do neural networks give intrinsically biased estimates?

The likelihood surface is not knowably convex (or concave)




The likelihood surface is conditional on the architecture and the data

There is no known correct minimum which is complete for all data

Even well converged networks are unknowably biased

In general outputs are highly informative



Outputs can look like model parameter estimates

They cannot be trusted as true predictions of the parameters

The can be used as informative summaries!

How can we use a neural network safely?

Build it into the physical model

Likelihood-free inference using machine learning

Approximate Bayesian computation


Why even introduce a neural network?

The curse of dimensionality


Inadequate sampling

Impossibly large numbers of simulations become necessary to correctly sample the posterior

We need to do some compression

Any pretrained neural network could be used

Is there a well motivated choice?

Information maximising neural networks


Which function $\mathscr{f}: \mathbb{R}^{\bf d}\to \mathbb{R}^\boldsymbol{\theta}$ maximises the Fisher information of the summaries ${\bf x}$ from that function?


$$\mathcal{L}({\bf d}|\boldsymbol{\theta})\to\mathcal{L}({\bf x}|\boldsymbol{\theta}, {\bf d})$$

Fisher information of the network summaries

Simulations $\{{\bf d}_i|i\in[1, n_{\bf d}]\}$ at a single parameter value $\boldsymbol{\theta}^*$

Seed matched simulations $\{{\bf d}_i^{\pm}|i\in[1, n_{\bf p}]\}$ at perturbed fiducial paremeters $\Delta\boldsymbol{\theta}^\pm=\boldsymbol{\theta}\pm\boldsymbol{\delta}$

Calculate the covariance

Calculate the derivative of the mean of the summaries with respect to the parameters

Calculate the Fisher information

$${\bf F}_{\alpha\beta}=\frac{\partial\mu_\mathscr{f}}{\partial\theta_\alpha}^T{\bf C}^{-1}_\mathscr{f}\frac{\partial\mu_\mathscr{f}}{\partial\theta_\beta}$$

This Gaussian form forces the summaries to be Gaussianised


Makes a non-linear mapping of (non-Gaussian) data to the compressed set of Gaussian summaries

Optimise the tunable parameters of the network such that $\textsf{ln}|{\bf F}_{\alpha\beta}|$ is maximised!

Once converged the network compresses and Gaussianises the data without losing information*





*in the optimal case...

And we get free maximum likelihood estimates...



$$\boldsymbol{\theta}^\textsf{MLE}_\alpha=\boldsymbol{\theta}_\alpha^*+{\bf F}_{\alpha\beta}^{-1}{\bf C}_\mathscr{f}^{-1}\frac{\partial\mu_\mathscr{f}}{\partial\theta_\beta}({\bf x}-\mu_\mathscr{f})$$

And we can get free approximations of the posterior



$$\textsf{Cov}[\boldsymbol{\theta}_\alpha^\textsf{MLE},\boldsymbol{\theta}_\beta^\textsf{MLE}] \ge {\bf F}_{\alpha\beta}^{-1}$$

Or do likelihood-free inference

We pass our observed data through the network
(put it to the side)

Draw simulations (cleverly?) from the prior model parameters, $\mathcal{P}(\boldsymbol{\theta})$ and pass them through the network

Measure the difference between the summarised simulations and the observed summary


Compare distance between observed summaries and simulation summaries and select results within $\epsilon$

Conclusions

Neural networks are not to be trusted

They can make trusty companions - when the correct framework is introduced

Using statistics we can build them into the forward model to give us lossless summaries

Great, but is there something better?

We could infer the neural network...

without any training data

Our observation

What does our observation actually look like

What does our observation actually actually look like

What do we need to infer the parameters of a neural network?

Neural physical engines

Build networks using physical principles.

  • Reduces number of parameters

  • Increases computational efficiency

  • Decreases overfitting

  • Improves interpretability

Convolutions for translational invariance

Use only causally relevant features

Use only informative features



github:multipole_kernels

Neural density estimators

Halo mass distribution function is a smooth function of mass given a density environment

Use a mixture density network

Our neural density estimator

$$\begin{align*} {\tiny n(M|\delta) =}&{\tiny \sum_i^N\alpha(\boldsymbol{\psi}, \boldsymbol{\theta}_i^{\boldsymbol{\alpha}})\mathcal{N}\left(\mu(\boldsymbol{\psi}, \boldsymbol{\theta}_i^{\boldsymbol{\mu}}), \sigma(\boldsymbol{\psi}, \boldsymbol{\theta}_i^{\boldsymbol{\sigma}})|M\right),}\\ {\tiny =}&{\tiny \sum_i^N\frac{\alpha(\boldsymbol{\psi}, \boldsymbol{\theta}_i^{\boldsymbol{\alpha}})}{\sqrt{2\pi\left(\sigma(\boldsymbol{\psi}, \boldsymbol{\theta}_i^{\boldsymbol{\sigma}})\right)^2}}\exp\left[-\frac{\left(\log(M) - \mu(\boldsymbol{\psi}, \boldsymbol{\theta}_i^{\boldsymbol{\mu}})\right)^2}{2\left(\sigma(\boldsymbol{\psi}, \boldsymbol{\theta}_i^{\boldsymbol{\sigma}})\right)^2}\right],}\end{align*} $$
$$ \begin{align*} {\tiny\alpha_i} & {\tiny=\text{softplus}(w_{i}^\alpha\boldsymbol{\psi}+b^\alpha_i),}\\ {\tiny\mu_i }&{\tiny = \left\{\begin{array}{ll} w_{i}^\mu\boldsymbol{\psi}+b^\mu_i&i=0\\ \textrm{Max}\left[0,~ w_{i}^\mu\boldsymbol{\psi}+b^\mu_i\right]+\mu_{i-1}&i>0\\ \end{array}\right.,}\\ {\tiny\sigma_i} &{\tiny = \text{softplus}(w_{i}^\sigma\boldsymbol{\psi}+b^\sigma_i)\;,} \end{align*} $$
2 Gaussians with
${\Tiny\boldsymbol{b}^\alpha \to \boldsymbol{b}^\alpha +\log(10^{-3})}$, ${\Tiny b_0^\mu \to b_0^\mu + \log\left(2\times10^{12}\right)}$ and ${\Tiny\boldsymbol{b}^\sigma\to \boldsymbol{b}^\alpha + \log(10^3)}$.

Likelihood

Poisson likelihood of the observed catalogue given the forward model



$$\begin{align*} {\Tiny\mathcal{L} =}&{\Tiny \sum_{j\in\textsf{catalogue}}\log\left[\sum_i^N\frac{\alpha_{i,j}}{\sqrt{2\pi\sigma_{i,j}^2}}\textsf{exp}\left[-\frac{\left(\textsf{ln}(M_j) - \mu_{i,j}\right)^2}{2\sigma_{i,j}^2}\right]\right]}\\ &{\Tiny - V\sum_{j\in\textsf{voxels},i=1}^N\frac{\alpha_{i,j}}{2}\textsf{exp}\left[\frac{\sigma_{i,j}^2}{2}\right]\textsf{erfc}\left[\frac{\textsf{ln}\left(M_\textsf{th}\right) - \mu_{i,j} - \sigma_{i,j}^2}{\sqrt{2\sigma_{i,j}^2}}\right].} \end{align*}$$

HMCLET

We want to infer the weights of the neural bias model

Hamiltonian Monte Carlo

Introduce momentum ${\small {\bf p}\leftarrow\mathcal{N}({\bf 0},{\bf M})}$ and solve Hamilton's equations

$$\begin{align*} {\Tiny\mathcal{H}(\boldsymbol{\theta}, {\bf p}) }&{\Tiny~= \mathcal{V}(\boldsymbol{\theta}) + \mathcal{K}({\bf p})}\\ &{\Tiny~=\mathcal{L}(\boldsymbol{\theta}|\boldsymbol{\delta})-\textsf{ln}\left[\pi(\boldsymbol{\theta})\right]+\frac{1}{2}{\bf p}^T{\bf M}^{-1}{\bf p}} \end{align*}$$ Accept samples according to probability $$\begin{equation*} {\Tiny\alpha = \textsf{Min}\left[\textsf{exp}\left(\Delta\mathcal{H}\right), 1\right]} \end{equation*}$$



Evolve using ${\Tiny\dot{\boldsymbol{\theta}}= {\bf M}^{-1}{\bf p}}$ and ${\Tiny\dot{\bf p}= -\nabla\mathcal{V}(\boldsymbol{\theta})}$

Leapfrog algorithm

Acceptance criterion

How does this work for neural networks?

Likelihood surface is extremely flat and highly degenerate


Nearly impossible to know the mass matrix, ${\bf M}$.

Use second order geometric information

How well does it work?

Sample the weights of the neural bias model


✓ Weight values are properly sampled after burn in
✓ NPE acts as a contrast enhancer

See the effects of the neural physical engine


✓ Non-local information is used to improve fit

And what does the halo mass distribution function look like?

Halo mass distribution function from neural bias model


✓ Fits data!

Second conclusion

We can infer neural networks built for physics

Highly efficient and bias free

Take home message

Stop doing machine learning, think, then start doing machine learning again!