Skip to main content
Regularization in iws.DataFit takes two forms. A prior pairs one fit parameter with a distribution from iws.stats. A penalty or constraint is a pybamm expression over any of the fit parameters, attached to an objective, for statements a distribution on one parameter cannot make. For why regularization is needed and how to choose prior strengths, see the Regularization Guide.

Distributions

A regularized fit

The priors dict adds a regularization term to the cost function so deviations from the prior mean are penalised in proportion to the prior’s inverse variance.

Penalties and constraints on expressions

A prior acts on one parameter. When the physics is a relation between several, write it as a pybamm expression over pybamm.InputParameter symbols named after the fit parameters and attach it to an objective. The example keeps a cathode’s remaining capacity at cut-off, its headroom, at or above 0.15 Ah with a one-sided quadratic penalty. It is zero wherever the headroom is comfortable and grows only where the fit would otherwise let the cathode fill up exactly as the anode empties.
Pick regularizer_weight so the penalty is comparable to the cost where it should bite: here a 0.1 Ah shortfall costs 10 mV against a voltage RMSE of about 9 mV. A symbol fun serialises to pybamm’s JSON form in to_config() and is rebuilt on the server, so the same object works locally and through client.pipeline.create.

Attaching priors via Parameter

Priors can also be attached directly to a Parameter rather than passed as a separate dict — useful when the prior is intrinsic to that parameter:

Dict-form priors

Schema configs can also be written as plain dicts — useful when a config is loaded from JSON or YAML. Two equivalent forms are accepted in the priors mapping:
The mapping key is the authoritative parameter name. Any name embedded inside the prior dict is ignored — including when the parameter itself is literally named "distribution".

Strict validation

Priors, distributions, and samplers are validated against discriminated unions at submission time. This catches typos and stale configs before a run starts, rather than letting them surface as opaque runtime crashes. Mistakes that are now rejected with a ValidationError:
  • An unknown distribution name (e.g. {"distribution": "Guassian", ...}).
  • A stray or misspelled key in a distribution, prior, or sampler config.
  • A type discriminator that doesn’t match the field (e.g. {"type": "Penalty", ...} placed under priors).
  • A bare scalar or list passed where a prior, distribution, or sampler is expected.
The legacy type alias is still accepted in place of the distribution discriminator inside a distribution config — e.g. the nested form’s inner dict {"type": "Normal", "mean": 3.0, "std": 0.2} resolves to a Normal — so existing serialized configs round-trip unchanged. Note that type only aliases distribution at the distribution level: at the top level of a prior mapping, type is the prior discriminator (it must be "Prior"), so a flat-form prior must name its distribution with distribution, not type.

Why LogNormal for diffusivities

Solid-phase diffusivities span many orders of magnitude (often 10−1610^{-16} to 10−1010^{-10} m²/s). A Normal prior on the raw value is hard to specify — mean ± std doesn’t reflect order-of-magnitude uncertainty. A LogNormal prior treats the parameter on a log scale, so “mean ± 1 std” corresponds to a factor of ee — much more natural. The mean=-32.2 in the example above is the natural log of ∼10−14\sim 10^{-14}, so the prior is centred on a typical particle diffusivity.

Regularization (theory)

Ridge regression, MAP estimation, bias–variance tradeoff.

Data Fitting overview

Putting priors together with objectives and optimisers.