Refactor from derivative
Link to score function
Cleanup the below
Rearranging this gives us the fundamental form:
Terminology Note: The term is called the score function in statistics, hence “score function estimator.”
Why This Matters: The Core Problem
Consider this ubiquitous problem in ML: we want to compute gradients of expectations:
The challenge: we can’t interchange the gradient and integral operators directly when the distribution depends on the parameters we’re differentiating with respect to.
The Trick in Action
Step-by-Step Derivation
Starting with our expectation gradient:
Key Insight: We’ve transformed a gradient of an expectation (hard to compute) into an expectation of a gradient-weighted function (easy to estimate via sampling).
Prerequisites & Mathematical Foundation
Required Concepts
-
Gradient operator:
-
Chain rule:
-
Leibniz integral rule: Under suitable conditions,
Regularity Conditions
- wherever (to avoid division by zero)
- Sufficient smoothness for interchange of differentiation and integration
- The expectation must exist
Core Applications in ML/AI
1. Policy Gradient Methods (Reinforcement Learning)
In RL, we maximize expected return:
where is a trajectory and is the trajectory distribution under policy .
The gradient:
For a trajectory :
Since transition dynamics don’t depend on :
2. Variational Inference (ELBO Maximization)
In variational inference, we maximize the Evidence Lower BOund (ELBO):
The gradient:
Alternative names: This is also called the REINFORCE gradient estimator in the VAE literature.
3. Black-Box Optimization
For non-differentiable objective :
This enables gradient-based optimization even when itself isn’t differentiable!
Concrete Example: Gaussian Distribution
Let where .
The log probability:
Score functions:
Verification of a key property:
This always holds (it’s the derivative of ).
Variance Considerations
The Variance Problem
The log derivative trick produces unbiased but potentially high-variance gradient estimates:
Variance Reduction Techniques
-
Baseline subtraction: Replace with where is a baseline
- Maintains unbiasedness since
- Optimal baseline:
-
Control variates: Use correlated variables with known expectations
-
Reparameterization trick: When possible, use pathwise gradients instead (e.g., in VAEs)
Connections to Related Concepts
Fisher Information
The Fisher Information Matrix:
This measures the information content about parameters in the distribution.
Natural Gradients
Natural gradient descent uses:
This is intimately connected to the score function.
Importance Sampling
When we can’t sample from directly:
The log derivative trick can be combined with importance sampling for off-policy learning.
Implementation Considerations
Numerical Stability
- Log-sum-exp trick: When computing , use stable implementations
- Gradient clipping: Score functions can have large magnitudes
- Standardization: Often helpful to standardize advantages/returns
Sample Efficiency
- More samples → lower variance in gradient estimates
- Trade-off between computation and variance reduction
- Batch size selection is critical
Common Pitfalls & Clarifications
-
Not zero-mean: While , the term is generally not zero-mean
-
Biased with finite samples: Individual gradient estimates are unbiased, but nonlinear functions of these estimates (e.g., Adam optimizer updates) can introduce bias
-
Distribution support: The trick assumes we can sample from - this isn’t always tractable
Advanced Extensions
Rao-Blackwellization
If and we can compute analytically:
This reduces variance by computing part of the expectation exactly.
Multiple Sampling Distributions
The multiple importance sampling variant:
Summary: When to Use the Log Derivative Trick
Use when:
- The expectation’s distribution depends on parameters you’re optimizing
- The function is non-differentiable or unknown
- You can sample from but can’t compute expectations analytically
- Working with discrete distributions (where reparameterization is impossible)
Consider alternatives when:
- Reparameterization is possible (continuous distributions with known inverse CDFs)
- Analytical gradients are tractable
- Variance is prohibitively high even with reduction techniques
This technique bridges the gap between probabilistic modeling and gradient-based optimization, making it fundamental to modern ML approaches that involve stochastic computation graphs.