Self complete guide on Variational Diffusion Models
Posted on Jul 30, 2026
Why This Post
Diffusion models are the state of the art in terms of generative models and I wanted to contribute in my own way with a self-complete article that summarizes the diffusion model training and sampling.
The Model
We define first some common terminology, then we describe the training and sampling process. Finally we define the most common ODE samplers used for generating data and we focus on the Classifier-Free guidance.
VDM Formulation
In this article we assume intuitive understand of what a generative model and what a diffusion models does (inject noise, denoise). We adopt the Variational Diffusion Model (VDM) formulation. To describe the noising process we adopt the convention
So runs forward from data to noise during the diffusion process, and the model’s job will be to run it backward.
Training
Training a diffusion model means teaching a network parametrized by to predict the noise that was added at an arbitrary point along this schedule. Concretely, each training step does three things.
1. Sample a timestep.
2. Inject noise. Given the clean sample , we form the noisy sample
where the signal and noise coefficients and come straight from the schedule:
We can see how both terms are determined by , which is the log Signal-to-Noise ratio. Down below we will describe how the noise schedule can be defined.
3. Predict the noise. The network outputs , an estimate of the we just added, optionally conditioned on a context vector (see Classifier-Free Guidance ).
Noise Schedule. The schedule itself can take several forms. Once the endpoints and are fixed, the common choices are:
where is a positive monotonic network, i.e.
We can go one step further with a multivariate learned contextual schedule: part of is fixed and learned globally, while another part is estimated from the context , letting the noise schedule adapt to the input.
Loss Function
With predictions in hand, the loss combines three terms:
The first term is the usual denoising objective. The last two terms anchor the endpoints of the process and, in our experience, are particularly useful when the noise schedule itself is being learned.
Once the network is trained, how do we turn pure noise back into data? That is the sampling problem. But before we get there, it’s worth pausing on the first term above — where does it actually come from, and why is it valid to leave it unweighted?
From the ELBO to the Diffusion Loss
The “standard diffusion loss” is not an arbitrary design choice. It is exactly what remains of the variational lower bound (ELBO) on once the number of diffusion steps is taken to infinity — but only under a specific parameterization and a specific way of sampling . Getting either of those wrong silently turns the training objective into something that no longer bounds the likelihood. It’s worth deriving this once end-to-end.
From Discrete Steps to a Continuous Integral
For a -step discretization with adjacent timesteps , the diffusion term of the discrete-time ELBO takes the form
where . Intuitively, each step penalizes the -reconstruction error, weighted by how much the signal-to-noise ratio drops over that step — steps where the signal degrades sharply matter more.
As , the finite difference converges to , and the sum over steps becomes an integral over continuous time:
Since is monotonically decreasing in , we have , so is a genuine, non-negative weight — larger wherever the SNR is falling fastest.
Changing Variables from to
Let . Since , the chain rule gives , so
which turns the integral over into an integral over log-SNR:
Switching to Noise Prediction
The network in our formulation predicts noise, not the clean signal directly, so we relate the two through the same forward equation: . Subtracting this from the identical expression for the true and — noting that both are evaluated at the same realized , so that term cancels exactly —
Substituting into the integral above, the factor introduced by the change of variables cancels exactly against the introduced by this reparameterization — the two are reciprocal by construction. What remains is
This is the elegant result underlying the standard diffusion loss: written in terms of noise prediction and integrated uniformly over log-SNR, the correct ELBO term is literally an unweighted mean-squared error. The catch is that “integrated uniformly over log-SNR” is doing real work here — it is not automatically satisfied just because we sample uniformly.
Sampling the Timestep Correctly
The boxed identity holds only when is itself distributed uniformly on . If , the standard 1-D change-of-variables rule tells us the induced density of is
For a schedule with constant slope — the linear schedule — this density is itself constant, so sampling uniformly happens to give uniformly, and no correction is required. For any schedule whose slope varies with (cosine, or a learned/contextual schedule that may saturate near the endpoints), over-samples whichever noise levels correspond to the flattest regions of — wherever is small, is large, so those log-SNR values get visited disproportionately often. Averaging the unweighted MSE under this mismatched sampling distribution converges to a biased quantity, not to .
There are two equivalent fixes, differing only in where the correction is paid.
Option A — importance-weight by . Keep , but reweight each sample by the density ratio between the target (uniform-in- ) and the induced (mismatched) distribution:
This requires the derivative of the schedule with respect to . When is vector-valued — for instance, a separate log-SNR curve per output dimension — this is naturally suited to forward-mode automatic differentiation (a Jacobian-vector product): since is a single scalar input mapping to many outputs, one JVP call recovers the entire per-element derivative in a single extra forward-like pass, exactly, without the truncation/round-off tuning that finite differences would require.
Option B — sample directly in log-SNR space. Draw first, then invert the schedule to recover the corresponding timestep,
Because is monotonic by construction, this inversion is well-posed and can always be solved reliably by bisection on , regardless of how nonlinear is. Since is now drawn directly from the target distribution, there is no density mismatch left to correct — only the constant prefactor from the interval length survives, and the loss is again plain, unweighted MSE. This is the approach adopted in the original VDM paper, and it trades a derivative for a root-find.
Both estimators are unbiased for ; they simply relocate the computational cost. Option A is preferable when is cheap to differentiate but awkward to invert (e.g. a small MLP with no closed form); Option B is preferable when has a closed-form or cheaply-invertible structure (e.g. the cosine schedule), since it avoids computing a derivative altogether.
Sampling
We develop two types of sampling: ancestral sampling and ODE sampling. While the first relies partially on stochastic variables, the second is purely deterministic.
Ancestral (Stochastic) Sampling
To sample with the VDM we walk the schedule backward, drawing from the conditional distribution with :
with
where . Equivalently, a single ancestral step can be written explicitly as
with . The first bracket pulls the sample toward the denoised estimate; the second injects fresh noise.
The stochastic sampler is simple but tends to need many steps. If we are willing to drop the noise injection, we can instead describe the reverse process as a deterministic ODE, which opens the door to fast, high-order numerical solvers.
ODE Sampling and the VDM Coefficients
Every diffusion process has an associated forward SDE of the form
and a corresponding probability-flow ODE that shares the same marginals but is fully deterministic:
In VDM we adapt the terms and as follows:
In other words, the VDM drift is just a time-dependent scaling of the current state, and the diffusion coefficient measures how fast variance is being pumped in net of that scaling. The only learned object is the score , which we read off directly from the noise prediction:
Substituting these into the probability-flow ODE gives a closed-form vector field that depends only on the schedule and the network output, which is exactly the function that the solvers below will integrate.
Likelihood Estimation for Diffusion Models using ODE
Continuous-time diffusion models formulate the generative process over a continuous time domain as the path of a Stochastic Differential Equation (SDE). By mapping this SDE to an equivalent deterministic Probability-Flow Ordinary Differential Equation (ODE), we cast the generative model as a Continuous Normalizing Flow (CNF). This permits exact likelihood estimation via the instantaneous change-of-variables theorem.
Problem Formulation
Let be a sample from the underlying data manifold. The forward diffusion process corrupts progressively over to yield noised latents , governed by the Gaussian transition kernel:
Equivalently, this forward dynamic maps exactly to the solution of the linear It^o SDE:
where is the standard forward Wiener process. In the VDM parameter space, the drift vector field and the scalar diffusion coefficient expand analytically as:
Following Anderson’s theorem, the reverse-time generation—flowing from back to —is characterized by a corresponding reverse SDE:
where denotes a reverse-time Wiener process, and is the Stein score of the marginal density . Crucially, Song et al. (2021) demonstrated that there exists a deterministic probability-flow ODE sharing the identical marginal probability densities as the SDE:
In practice, the exact score is intractable. We approximate it using our neural network via the standard score-to-noise parameterization, creating our parameterized ODE drift field .
Because this ODE describes a continuous normalizing flow mapping the known prior to the complex targeted distribution , we can invoke the instantaneous change-of-variables formula from Neural ODEs (Chen et al., 2018) to compute the exact log-likelihood:
where the divergence operator acts as the infinitesimal volume expansion term, and the parameterized field defines as:
Consequently, evaluating the exact likelihood condenses down to integrating the divergence of the neural ODE drift field continuously tracked along the generated trajectory.
Hutchinson Trace Estimator
Directly computing requires a full Jacobian and scales as . To avoid this, we use Hutchinson’s estimator. For any matrix ,
if and .
To see unbiasedness, expand
Taking expectation,
Because , we obtain
In diffusion likelihood estimation this is crucial: we never materialize the full Jacobian, and instead compute Jacobian-vector products, reducing practical cost from quadratic to roughly linear in dimension per probe. In practice, averaging a small number of probes (often around - ) is already effective.
Choice of the Projection Vector
The projection vector must satisfy
Two common choices are:
- Gaussian probes: .
- Rademacher probes: each component is independently sampled as
Both satisfy the required moment conditions and therefore produce unbiased trace estimates.
ODE Solvers
Terminology. Given a step size , we want to advance from timestep to starting from our sample , where denotes the ODE vector field derived above. The different solvers below trade off accuracy against the number of function evaluations (NFE).
Euler Method
While computationally inexpensive, this first-order method accumulates local errors on the order of , at a cost of NFE per step.
Midpoint Method
A second-order Runge–Kutta method that estimates the vector field at the midpoint of the integration step, giving a more accurate gradient estimate:
Error: , at NFE.
Fourth-Order Runge–Kutta Method
Pushing the same idea further, RK4 blends four slope evaluations:
Error: , at NFE.
So far every method uses a fixed step size. The smarter move is to let the solver pick on the fly based on how hard the trajectory is to integrate — this is what adaptive solvers do.
Dormand–Prince (dopri5)
This method is adaptive: the NFE depends on the local approximation error. The trick is to compute two Runge–Kutta solutions of different order (5th and 4th) from the same stages and use their difference as an error estimate. The stages are
with fractional time locations
and internal coefficient matrix
The two approximations are
Their difference gives the local error estimate
We then normalize the error,
where and are hyperparameters. The step is accepted if ; when it is rejected we keep and shrink the step:
Adaptive Heun
Adaptive Heun is a lighter-weight cousin of dopri5 that combines Euler and Heun. We compute
and form the Euler and Heun solutions:
The estimation error and its normalization mirror dopri5:
with . We accept if ; otherwise we reduce the step as
where the exponent follows the rule for an order- solution (here , with the safety factor usually around ).
Good solvers get us samples efficiently — but so far the generation is unconditional. The last ingredient is a way to steer the model toward a desired class or behavior.
Classifier-Free Guidance
Classifier-free guidance lets us guide the generation process of a diffusion model without training a separate classifier. We denote by the conditioning class.
During training. Instead of always estimating the conditional score , with probability (typically – ) we replace the condition with a null token and predict . This single network thus learns both the conditional and unconditional scores.
During sampling. We evaluate both scores and combine them:
For this can be rewritten as a convex-style blend
where is the guidance strength: it dictates how strongly the generation follows the class-conditioned signal. Larger means tighter adherence to the condition at some cost to diversity.