Refactor from derivative
Link to score function

Cleanup the below

Rearranging this gives us the fundamental form:

Terminology Note: The term is called the score function in statistics, hence “score function estimator.”

Why This Matters: The Core Problem

Consider this ubiquitous problem in ML: we want to compute gradients of expectations:

The challenge: we can’t interchange the gradient and integral operators directly when the distribution depends on the parameters we’re differentiating with respect to.

The Trick in Action

Step-by-Step Derivation

Starting with our expectation gradient:

Key Insight: We’ve transformed a gradient of an expectation (hard to compute) into an expectation of a gradient-weighted function (easy to estimate via sampling).

Prerequisites & Mathematical Foundation

Required Concepts

  1. Gradient operator:

  2. Chain rule:

  3. Leibniz integral rule: Under suitable conditions,

Regularity Conditions

  • wherever (to avoid division by zero)
  • Sufficient smoothness for interchange of differentiation and integration
  • The expectation must exist

Core Applications in ML/AI

1. Policy Gradient Methods (Reinforcement Learning)

In RL, we maximize expected return:

where is a trajectory and is the trajectory distribution under policy .

The gradient:

For a trajectory :

Since transition dynamics don’t depend on :

2. Variational Inference (ELBO Maximization)

In variational inference, we maximize the Evidence Lower BOund (ELBO):

The gradient:

Alternative names: This is also called the REINFORCE gradient estimator in the VAE literature.

3. Black-Box Optimization

For non-differentiable objective :

This enables gradient-based optimization even when itself isn’t differentiable!

Concrete Example: Gaussian Distribution

Let where .

The log probability:

Score functions:

Verification of a key property:

This always holds (it’s the derivative of ).

Variance Considerations

The Variance Problem

The log derivative trick produces unbiased but potentially high-variance gradient estimates:

Variance Reduction Techniques

  1. Baseline subtraction: Replace with where is a baseline

    • Maintains unbiasedness since
    • Optimal baseline:
  2. Control variates: Use correlated variables with known expectations

  3. Reparameterization trick: When possible, use pathwise gradients instead (e.g., in VAEs)

Fisher Information

The Fisher Information Matrix:

This measures the information content about parameters in the distribution.

Natural Gradients

Natural gradient descent uses:

This is intimately connected to the score function.

Importance Sampling

When we can’t sample from directly:

The log derivative trick can be combined with importance sampling for off-policy learning.

Implementation Considerations

Numerical Stability

  • Log-sum-exp trick: When computing , use stable implementations
  • Gradient clipping: Score functions can have large magnitudes
  • Standardization: Often helpful to standardize advantages/returns

Sample Efficiency

  • More samples → lower variance in gradient estimates
  • Trade-off between computation and variance reduction
  • Batch size selection is critical

Common Pitfalls & Clarifications

  1. Not zero-mean: While , the term is generally not zero-mean

  2. Biased with finite samples: Individual gradient estimates are unbiased, but nonlinear functions of these estimates (e.g., Adam optimizer updates) can introduce bias

  3. Distribution support: The trick assumes we can sample from - this isn’t always tractable

Advanced Extensions

Rao-Blackwellization

If and we can compute analytically:

This reduces variance by computing part of the expectation exactly.

Multiple Sampling Distributions

The multiple importance sampling variant:

Summary: When to Use the Log Derivative Trick

Use when:

  • The expectation’s distribution depends on parameters you’re optimizing
  • The function is non-differentiable or unknown
  • You can sample from but can’t compute expectations analytically
  • Working with discrete distributions (where reparameterization is impossible)

Consider alternatives when:

  • Reparameterization is possible (continuous distributions with known inverse CDFs)
  • Analytical gradients are tractable
  • Variance is prohibitively high even with reduction techniques

This technique bridges the gap between probabilistic modeling and gradient-based optimization, making it fundamental to modern ML approaches that involve stochastic computation graphs.