Interpreting Parton Distributions with Shapley Values

Raphaël Bonnet-Guerrini1,3, Stefano Carrazza2,3, Stefano Forte2,3,
Eva Groenendijk2,3, and Ramon Winterhalder2,3


1Computer Science department, University of Milan; 2TIF lab, Physics Department, University of Milan;
3Istituto Nazionale di Fisica Nucleare, Sezione di Milano

Parton Distribution Functions (PDFs) and NNPDF

\[ f_i(x, Q_0^2) = x^{(1-a)}(\text{NN}_i(x)-\text{NN}_i(1)) \]
  • Parton carrying a fraction of momentum \(x\) at scale \(Q^2\)

  • NNPDF: neural net, no functional form assumed


PDF fitting: an inverse problem

PDFs groundtruth do not exist they fit from experimental data

Theory (QCD, QED) allows us to compute the observable values \(\mathcal{O}_n\) from the PDF space.

PDF space

Why does a given flavor look like this in a given \(x\) region? Which datasets are responsible for this behavior?

What if we perturb the PDF directly in the PDF space ?

• Multidimensionality of the PDF space • Rotation and convolutions • PDF-Observable correlations • DGLAP Evolution

Difficult to perform systematic tests and obtain an understanding of independent flavor impact on the fit.

NNPDF a black box model ?

Black-box systems are a recurrent interpretability issue in modern DL. We can take inspiration from the methods developed there to solve the interpretability problem of PDFs.

XAI answer : What's the impact of one feature on the output of the model ?

Shapley Values

Game theory: average marginal contribution of player \(i\) across all coalitions \(S\).

\[ \phi_i = \sum_{S \subseteq N \setminus \{i\}} \frac{|S|! \, (|N| - |S| - 1)!}{|N|!} \left[ v(S \cup \{i\}) - v(S) \right] \]

Cost \(\sim 2^N\): intractable for high dimensional attribution.

SHAP buys tractability by assuming feature independence, known to distort HEP results when features are correlated.

PDF flavors as players

Take a determined PDF set as given, each flavor is a player. Blind to how it was fitted.

  • \(\phi_i\): Shapley value of flavor \(i\)
  • \(S\): subset of flavors, \(i\notin S\)
  • \(v(S)=\chi^2\) with the flavors in \(S\) perturbed, the rest untouched

The Shapley analysis in NNPDF

Load trained PDFs from n3fit.

  1. Build all coalitions of perturbed PDFs
  2. Compute \(\chi^2\) for each
  3. Per flavor: compare \(\chi^2\) perturbed vs not

How to perturb the PDF space ?

How do we perturb a PDF ? A smooth calibrated Gaussian bump:

How to perturb the PDF space ?

How do we perturb a PDF ? A smooth calibrated Gaussian bump:
\[ xf_j^{\text{pert}\pm}(x,Q_0^2)=xf_j+\delta_j^\pm,\qquad \delta_j^\pm(x)=\pm\Delta_j^\pm(x_\mu)\,e^{-\log_{10}^{2}(x/x_\mu)/(2w^{2})} \]

\(\Delta^\pm_j(x_\mu)\) = 68% CI bounds of the replicas → calibrated to the PDF uncertainty, comparable across flavors.

How to avoid directional bias ? average the Shapley value of both directions \[ v(S)=\frac{v_+(S)+v_-(S)}{2}. \]

Coalitions: what does \(\phi_g\) measure?

Flavors join the perturbation one at a time, in a random order.
For each possible \(S\), we ask what \(g\) adds to the \(\chi^2\) at the moment it walks in.

order 1ΣT3gVV3T8T15V8S = {Σ, T3}
order 2gVT8ΣV3T15V8T3S = { }
order 3VT3ΣV8gT8T15V3S = {V, T3, Σ, V8}
order 4V3T15T8VΣV8T3gS = {V3, T15, T8, V, Σ, V8, T3}
… all \(8!\) orders

blue perturbed flavor  ·  grey non-perturbed flavor  ·  g perturbed flavor \(g\)

\[ \phi_g \;=\; \Big\langle\, \chi^2(S\cup\{g\}) - \chi^2(S) \,\Big\rangle_{\text{all orders}} \]

\(8! = 40320\) orders, but only 128 distinct \(S\).

Correlated perturbation

Correlations between flavors come from the sum rules:

\[ \text{Momentum sum rule:} \quad\int_0^1dxx(g(x,Q)+\Sigma(x,Q))=1, \] \[ \text{Valence sum rule:} \int_0^1dxV(x,Q)=\int_0^1dxV_8(x,Q)=3, \qquad \int_0^1dxV_3(x,Q)=1, \]

Flavors in \(S\) get the bump; the others have their normalization \(A_i\) readjusted so the rules still hold.

How to interpret the Shapley values?

  • \(\phi_j> 0 \): \(\chi^2\) worsens → flavor constrained there
  • \(\phi_j\approx 0 \): dataset insensitive to \(j\)
  • \(\phi_j< 0 \): \(\chi^2\) improves → hint of data tension

  • Depends on the \(x\) region (local perturbation) and on the dataset (\(\chi^2\)).

    Completeness: \[\sum_{j \in N} \phi_j = v(N)-v(\emptyset)\] \(v=\chi^2\), so \(\phi_j\) is in units of \(\chi^2\).

    Uncertainty on the Shapley value

    Replay the whole game on each replica \(f^k\):

    \[ \phi_i^{(k)}=\phi_i(f^k), \qquad \bar\phi_i = \frac{1}{N_{\text{rep}}}\sum_k \phi_i^{(k)}, \qquad \sigma^2_{\phi_i}=\frac{1}{N_{\text{rep}}-1}\sum_k\left(\phi_i^{(k)}-\bar\phi_i\right)^2. \]
    • 100 replicas × 128 coalitions
    • Plots: median + 68% band \([\phi^{16},\phi^{84}]\)
    • Quick scans: central PDF only, no uncertainty

    Scanning the whole \(x\) range

    50 log-spaced \(x_\mu\), central NNPDF4.0, flavor basis.

    One row misbehaves: the gluon dips to zero twice.

    A blind spot in the gluon

    We infer the gluon from how data change with \(Q^2\)
    (scaling violations)

    At \(N=2\), momentum conservation creates a non-evolving mode

    No evolution = little signal to constrain the gluon

    Mellin mapping: \(N_0=2 \;\mapsto\; x\approx0.07\)

    The non-evolving mode lands at the dip

    Why this matters

    • Expect an inflated uncertainty there. It does not happen
    • The band interpolates straight across: new physics in the gluon here could be absorbed with no visible distortion
    • Invisible to standard diagnostics: the Shapley value makes it explicit

    Is it universal? NNPDF4.0 vs MSHT20 vs CT18

    • The argument is pure QCD evolution \(\Rightarrow\) the dip must be universal
    • MSHT20: same dip
    • CT18: same dip, but sitting higher
    • Rescale CT18 to \(0.5\,\delta_{\text{pert}}\) and it lands on the others: CT18 uncertainties are more conservative
    • Holds for fixed functional forms too, not just the NN

    Application: optimizing the hyperopt folds

    • 4 folds, each meant to represent the whole dataset, built by intuition
    • Profiles are similar across folds \(\Rightarrow\) the design was adequate
    • At large \(x\), \(T_3,T_8\) go negative for the full set, for no fold → tension with post-hyperopt data
    • Next: choose folds minimizing the distance between profiles

    Outlook

    • Quantitative methodology-agnostic metric of which flavor is constrained where
    • Recovers the standard lore, and exposes what it misses


    Future steps

    • Data not in the fit → impact of a new measurement before including it
    • PDF K-folding, combined sets
    • Larger feature space (\(\sim10^9\) coalitions) → controlled approximation
    • Not specific to PDFs: any black box with a likelihood x

    Thank you !

    Contact : raphael.bonnet-guerrini@unimi.it

    Other interpretability methods I work on:

    Sparse autoencoders on a neutrino foundation model

    Interpretable transient classification

    This work was supported by the European Union's Horizon Europe research and innovation programme under the Marie Sklodowska-Curie grant agreement No 101168829, Challenging AI with Challenges from Physics: How to solve fundamental problems in Physics by AI and vice versa (AIPHY).

    Backup Slides

    Replica scan, six values of \(x_\mu\)

    DIS: HERA vs fixed target

    Jets and top pair production

    Dataset regularization

    • Perturbation can drive a cross section to \(\sim0\) → ratio observables blow up
    • Iteratively drop any \(\mathcal{D}_i\) with \(v_{\mathcal{D}_i}(S)>10^5\)
    • In practice only DYE906 / DYE866, for a few \(x_\mu\)

    Shapley Values Pseudocode

    # INPUT: observables, mu,sigma,amplitude, n_samples, n_flavors = n
    # OUTPUT: shapley_vals[], baseline_chi2, cache
    
    baseline_chi2 = evaluate_chi2(observables, flavor_subset=[])   # v({})
    
    cache = {}   # map subset -> v(S)
    all_subsets = power_set(0..n-1)     # exclude full-set if desired
    
    for i in 0..n-1:
    SV = 0
    for S in all_subsets:
        if i in S: continue
    
        # v(S)
        if S not in cache:
        cache[S] = evaluate_chi2(observables, flavor_subset=list(S), mu,sigma,amplitude, n_samples)
        vS = cache[S]
    
        # v(S ∪ {i})
        S_with = S ∪ {i}
        if S_with not in cache:
        cache[S_with] = evaluate_chi2(observables, flavor_subset=list(S_with), mu,sigma,amplitude, n_samples)
        vSw = cache[S_with]
    
        Δ = vSw - vS                                   # marginal contribution
        s = |S|
        w = factorial(s) * factorial(n - s - 1) / factorial(n)
        SV += w * Δ
    
    shapley_vals[i] = SV
    
    # RETURN: shapley_vals, baseline_chi2, evaluated_coalitions=|cache|
    # COMPLEXITY: time ~ O(n · 2^n · cost_eval), space ~ O(2^n) (memoized)
                        

    Shapley Values weight

    In game theory, Shapley Values represent the contribution \(\phi_i\) of a player \(i\) within a coalition \(S\) by comparing the outcomes of scenarios where the player is present \(v(S \cup \{i\})\) versus absent \(v(S)\). \[ \phi_i = \sum_{S \subseteq N \setminus \{i\}} \text{w}(S) \left[ v(S \cup \{i\}) - v(S) \right] \label{eq:comb_SV} \] With: \[ w(S) = \frac{|S|! \, (|N| - |S| - 1)!}{|N|!} \] \(w(S)\) weights for the importance of the coalition \(S\) being tested. It is the probability that the set of players who come after \(i\) is exactly \(S\).
    • \(|S|!\) is the number of ways to order the predecessors of \(i\)
    • \((|N| - |S| - 1)!\) is the number of ways to order the players before \(i\)
    • \(|N|!\) is the number of ways to order all players

    Shapley Values weight

    This result is derived in the following way: Given \( |N| = n \), for a fixed player \(i\), let \( S \subseteq N \setminus \{i\} \) with \( |S| = s \). The probability that players before \(i\) are exactly \(S\) can be decomposed into two conditions: First, \(i\) can be in any position in \(N\), so: \[ p(\text{pos}_i = s+1) = \frac{1}{n} \] Second, given that \( i \) is in position \( s+1 \), the \( s \) players before \( i \) form a uniformly random subset of the remaining \( n-1 \) players of size \(\binom{n-1}{s}\): \[ p(\text{predecessors} = S \mid \text{pos}_i = s+1) = \frac{1}{\binom{n-1}{s}} \] To fulfill these conditions, we multiply those probabilities: \[ \frac{1}{n}\cdot \frac{1}{\binom{n-1}{s}} = \frac{1}{n}\cdot \frac{s!(n-1-s)!}{(n-1)!} = w(S) \]