DataFit has two coupled pieces:
- Objectives (
iws.objectives.*) — what experiments to compare model output against. - Cost (
iws.costs.*) — how the per-point disagreements are aggregated into a single number.
Available cost functions
For MLE, see
iws.costs.GaussianLogLikelihood — it accepts per-variable noise standard deviations or can estimate them alongside the fitting parameters. It produces a Gaussian negative log-likelihood suitable for Bayesian and MAP estimation.
Wiring a cost into a fit
cost is omitted, the optimizer’s default cost function is used (typically a least-squares form).
cost accepts a cost schema instance (e.g. iws.costs.RMSE()) or a config dict with an explicit type key (e.g. {"type": "RMSE"}). A bare name string like cost="RMSE" is rejected with a validation error — wrap it as {"type": "RMSE"} instead.Wasserstein weighted point-cloud mode
By defaultiws.costs.Wasserstein() compares the model and data samples for each objective variable with uniform weights (sorted point-wise comparison). Set both position_variable and weight_variable to switch to weighted point-cloud mode: one variable supplies the positions, the other supplies the (sign-stripped, renormalised) weights, and a single Wasserstein-1 distance is computed per objective.
Use this when you want to match a density by position rather than sample-by-sample values — for example, lining up dQ/dV peaks in voltage rather than penalising every dQ/dV residual.
Both iws.objectives.MSMRFullCell and iws.objectives.ElectrodeBalancing expose the matching Differential capacity [Ah/V] values alongside their Voltage [V] (dQdU) masked-axis sibling, so either can drive a weighted point-cloud fit:
position_variable and weight_variable must be set together — providing only one raises a validation error. Weights are taken as absolute values and renormalised internally, so sign conventions on dQ/dV don’t matter. Residual-array output is not available in this mode.ElectrodeBalancing options for OCV fitting
ElectrodeBalancing accepts the following keys in its options dict to control how the full-cell OCV is processed before the objective is evaluated. These apply regardless of which cost function the fit uses (not only weighted Wasserstein):
Scoping a cost with calculation_structure
By default every cost on a DataFit consumes every objective and every objective variable in the outputs. Set calculation_structure on a cost to scope it explicitly: a mapping from objective name to the list of variable names that cost should compute, or None to compute all of that objective’s variables (an empty list computes none).
Objectives you leave out of the mapping are not dropped. Inside a DataFit each unscoped objective is bound to all of its variables — the same as mapping it to None — so scoping one objective (e.g. {"ocp": ["Voltage [V]"]} while a "cc" objective also exists) still computes "cc" in full.
Use this when one cost should only see a subset of variables — most commonly when you pair a per-variable cost (e.g. SSE) with a weighted Wasserstein. The Wasserstein owns the dQ/dV variables (whose model and data sides may have different lengths by construction), and the SSE is scoped to skip them so the lengths never collide.
calculation_structure replaces the deprecated objective_names field (a flat list of objective names with no per-variable control). Specifying both on the same cost raises a validation error.Length-mismatch warning
Element-wise costs (SSE, MSE, RMSE, MAE, Max) combine the model and data arrays point-by-point, so a variable whose model and data sides have different lengths almost never gives a meaningful score. At fit setup, DataFit checks the shapes of every variable each cost is configured to score and emits a UserWarning for each mismatch — for example:
calculation_structure so it only sees variables whose model and data lengths match — and route any model-axis variables to a Wasserstein cost (or another distribution metric) instead. Distribution costs like Wasserstein are skipped by the check, since unequal-length sample sets are expected there.
Aligning dQ/dV peaks on the model voltage axis
iws.objectives.ElectrodeBalancing can emit dQ/dV on the model’s own full-window voltage axis in addition to (or instead of) the data voltage grid. Set dQdU model axis: True in options and add the two model-axis variables — "Differential capacity [Ah/V] (model axis)" and "Voltage [V] (model axis)" — to objective variables.
Use this when you want a weighted cost (typically Wasserstein in point-cloud mode) to position-shift — i.e. align dQ/dV peaks in voltage rather than residual-by-residual on the data grid. The model and data sides have different lengths by construction, so only a weighted cost should consume them; pair them with a sibling per-variable cost scoped via calculation_structure (see above) to keep the rest of the fit honest.
The existing data-axis variables ("Differential capacity [Ah/V]" plus the masked siblings "Voltage [V] (dQdU)" / "Capacity [A.h] (dQdU)") remain available — both axes can be requested side by side.
Available objectives
Combine several by passing a
dict[str, objective] to DataFit.objectives.
The objectives that run a simulation — CurrentDriven, Pulse, EIS, CalendarAgeing, and CycleAgeing — need a model to simulate against, and construction fails without one. Give it as options={"model": pybamm.lithium_ion.SPMe()}, or point at a model stored on the platform with options={"parameterized_model_id": "<id>"}. The remaining objectives fit data directly and take no model.
MSMRHalfCell and MSMRFullCell project fitted host-site fractions Xj onto the bounded simplex so they sum to exactly 1 and stay within bounds. MSMRFullCell always does this; on MSMRHalfCell it is the "project" default of the constrain Xj method option, and the recommended setting.
constrain Xj method also accepts "reformulate" and "explicit", which are kept for reproducing older fits. They enforce the constraint more weakly: "reformulate" replaces the final Xj with the complement of the others, and "explicit" applies a soft constraint the optimizer can trade away. Neither can now pass unnoticed — both objectives check the fitted fractions after the fit and fail when |sum(Xj) - 1| exceeds 0.05 or any Xj is negative, warning above 1e-6. Set the DataFit "validate" option to False to skip the check and get the unphysical values back. Set penalize Xj complement bounds on MSMRHalfCellOptions to additionally penalize a reformulated complement that falls outside the final Xj bounds.
GITTModel: diffusion-only model for GITT and pulse fits
GITTModel is a fitting-only model intended for extracting solid-phase diffusivities (and a single lumped ohmic resistance) from GITT or pulse-relaxation measurements. It solves x-averaged spherical particle diffusion in each modelled electrode, with the surface flux set by the applied current, and computes the cell voltage from the electrode open-circuit potentials evaluated at the particle-surface stoichiometries, minus an ohmic drop through a lumped "Ohmic resistance [Ohm]" parameter.
There are no reaction kinetics (Butler-Volmer), no electrolyte dynamics, and no thermal effects — all parameters are constant except the OCPs. Use it when you want fast, well-conditioned fits to diffusion-dominated portions of GITT or pulse data, and reach for SPM / SPMe / DFN when you need a full physics simulation.
Select the cell configuration via the "working electrode" option:
Each modelled electrode is parameterised with the standard full-cell parameter names (thickness, active material volume fraction, particle radius, diffusivity, OCP, maximum and initial concentrations), plus the current function, electrode cross-sectional area, initial temperature, and
"Ohmic resistance [Ohm]".
Fitting a full-cell GITT measurement
Fitting a half-cell pulse measurement
Pass"working electrode": "positive" to model a single electrode against a lithium-metal counter. Only the working-electrode parameters are needed.
"working electrode" only accepts "both" or "positive" — anything else fails schema validation. Any other keys in options are forwarded to the underlying battery-model options for parameter bookkeeping; they do not change the diffusion-only physics.EIS: fitting impedance spectra
iws.objectives.EIS compares model impedance against measured electrochemical impedance spectroscopy (EIS) data in the frequency domain, using PyBaMM’s EISSimulation. Use it to identify kinetic parameters — the two electrodes’ reference exchange-current density (the charge-transfer arc diameters) and double-layer capacity (the arc frequencies) — plus the ohmic offset, which time-domain discharges alone leave degenerate.
The data must contain Frequency [Hz], Z_Re [Ohm], and Z_Im [Ohm] columns (capacitive band with Z_Im < 0, matching the SDK’s EIS upload validator).
The model must be built with
"surface form": "differential" — EISSimulation rejects the default form, and "algebraic" drops the double-layer capacity being fitted.Refining the mesh via simulation_kwargs
options["simulation_kwargs"] forwards mesh and discretisation kwargs (var_pts, submesh_types, geometry, spatial_methods) straight through to EISSimulation — the same shape you already use for CurrentDriven / Pulse objectives. Use it to refine the particle mesh when the default resolution smears the charge-transfer arc.
Time-domain-only keys (solver, solver_kwargs, solve_kwargs, output_variables, experiment, …) are silently dropped with an info log, so a simulation_kwargs dict shared with a time-domain objective is accepted without raising.
Fitting across multiple SOC set-points
Fit oneEIS objective per SOC set-point and combine them in the same DataFit. Each objective pins its operating point with an objective-level Initial SOC [%] (or Initial voltage [V]) parameter, and the shared kinetic parameters are identified jointly across all spectra:
Specifying data_input
Every objective’s data_input (and any other data field on a calculation or interpolant) accepts the same set of forms:
- A reference string:
"db:<id>"to reference an uploaded measurement."file:..."and"folder:..."are read from your local machine and inlined into the config by the API client on submit, so they work both locally and when you submit a fit to Ionworks — subject to the same 1,000-row inline limit as a bareDataFrame. For larger datasets, upload a measurement and reference it with"db:<id>". - An
ionworksdata.DataLoader(local or fetched withDataLoader.from_db(...)). - A bare pandas or polars
DataFrameof pre-loaded columns. - An
ionworksdata.AnalysisLoader, which reads one of a measurement’s stored analyses as the data instead of its time series — see Taking the targets from a stored analysis.
DataFrame is passed, it is auto-wrapped on serialization to match the parser’s expected {"data": <columns>} shape — so data_input=df and data_input={"data": df} behave the same. String paths and already-wrapped dicts are left untouched.
A string
data_input must start with db:, file:, or folder: to say where the data is read from — a bare path such as "data/1C.csv" is rejected when the objective is constructed. Write it as "file:data/1C.csv".data_input is matched against the payload shapes above by its keys — {"time_series": ..., "steps": ...}, {"data": ..., "options": ...}, {"data": "db:<id>", "analysis": ...}, or {"data": ..., "metadata": ...} — and validated strictly, so a misspelt key or loading option is reported when the objective is constructed rather than being ignored. A dict matching none of those shapes is treated as a plain mapping of column names to values. time_series, steps, data, options, time_range, analysis, and metadata are therefore reserved: a column literally named one of them is read as a payload key instead of as a column.
Inline DataFrames are capped at 1,000 rows per call. For larger datasets, upload as a measurement and reference it by ID instead. See inline time series size limit.
Generating a CycleAgeing experiment from data
iws.objectives.CycleAgeing normally requires an explicit pybamm.Experiment describing the cycling protocol. When the protocol is already encoded in the cycler step information attached to your data, set experiment="from data" to skip rebuilding it by hand. The experiment is generated lazily, when the fit starts, by calling DataLoader.generate_experiment() on the loaded step table.
Use this when:
- The fitted data carries its own step information (a local
ionworksdata.DataLoader, or one fetched withDataLoader.from_db(...)). - You want the simulated protocol to track the measurement protocol exactly — the per-step current, power, or voltage setpoint the cycler recorded, and either the duration of each step or the endpoint it reached (see Ending each step on its endpoint).
DataLoader as experiment instead — the steps come from that loader, while the residuals are still computed against data_input:
experiment="from data" requires data_input to resolve to a DataLoader (or a dict whose "data" entry is a DataLoader) that carries step information. When you pass a separate DataLoader as experiment, that loader must carry the step information instead. Either way, configurations missing steps fail fast at objective construction with a clear error, before any simulation runs.Taking the targets from a stored analysis
A per-cycle summary is often already stored on the platform as an analysis of the measurement it was extracted from — for example degradation modes per RPT. Pointdata_input at it with ionworksdata.AnalysisLoader, and the fit reads that analysis as its targets while the same measurement’s steps give the experiment:
AnalysisLoader selects the analysis by analysis_id, analysis_type, and/or name; every field given must match, except that analysis_id names one analysis exactly and any other fields are then ignored. Give at least one — an empty selector is rejected — and separate two analyses of the same type by name. A selector that matches more than one analysis raises rather than picking one. The equivalent dict form is {"data": "db:<measurement-id>", "analysis": {"analysis_type": "lam_lli_from_rpt"}} (iws.AnalysisSpec builds the analysis part). An analysis is a feature table, not a trace, so it cannot be combined with options or time_range.
To check which analysis a selector will pick before submitting, use client.analysis.find_one.
Ending each step on its endpoint
By default each generated step runs for the duration the cycler recorded. Settermination="events" to end it where the measured step ended instead, in the variable it was controlling — a constant-current or constant-power step at its final voltage, a voltage hold at its final current. The step then carries no duration, so its length is the model’s own:
"events"for endpoint-driven protocols — rate capability, CC-CV cycling — where the question is how much capacity to this endpoint? A model whose capacity differs from the cell’s then disagrees in capacity rather than ending a discharge at a dangling voltage."duration"(the default) for time-driven protocols, where the question is what voltage at this time?
Simulating a voltage hold instead of replaying its current
A generated experiment emits each constant-voltage step as a current interpolant: the model is driven with the current the cycler recorded. During a hold that current is what the cell’s impedance produced, so prescribing it asks the model to reproduce the cell’s impedance rather than to simulate the hold. Setuse_cv=True to emit holds as voltage holds instead — and with termination="events", each one ends at its measured final current, the way the cycler’s taper cut-off did:
A taper endpoint drifts as the cell ages, so on a long cycling run
use_cv=True with "events" gives almost every hold its own distinct step. CycleAgeing defaults experiment_model_mode="unified" precisely to build one model for a repeated step, and distinct holds cost that reuse — cut-off voltages, which repeat exactly, do not.pybamm.Experiment raises at objective construction — build what you want into that experiment instead.
Which cycle each data row describes
Your data’sCycle number column says which cycle of the simulation each row is compared against. For a generated experiment it is matched against the steps table’s Cycle count rather than used as a position: the experiment holds one cycle per Cycle count, but numbers its own cycles from zero whatever the table says.
Cycle count is not the cycler’s own label — that is Cycle from cycler. It is a counter ionworksdata derives, starting at 0 and running contiguously, so on a cycler labelling its cycles 1, 2, 3 you get:
first_step / last_step keeps the original counts instead of renumbering, so a loader sliced to start at the second cycle has Cycle count [1, 2] behind experiment cycles 0 and 1. Number those rows 1, 2 — what the steps table says — not 0, 1.
A per-cycle summary built with ionworksdata.cycle_metrics.get_cycle_metrics already carries the right values in its own Cycle count column, so renaming that column to Cycle number is all the fit needs.
An experiment you pass in directly carries no Cycle count to match against, so there a Cycle number stays a 0-based position among its cycles.
A by_cycle metric’s values are compared with the data row for row, over the rows where that variable has a value. Leave cycles= unset and the cycles are taken from the data. Set it and it must name the cycles the data has a value of that variable for — a cycle with no such row raises at build, instead of silently comparing simulation cycle 0 against a row describing a different cycle of the cell.
Targets measured at different cycles
Not every quantity an ageing test produces is measured at every cycle — a resistance sweep might run only on alternate RPTs. Leave the missing values asNaN: a NaN in a target column means there is no measured value at that cycle, and that row is left out of the fit for that variable alone. The other variables still use it, so only measured points are scored and nothing needs interpolating onto a shared cycle grid.
OCPHalfCell options
Interpolating the driving current
CurrentDriven and Pulse drive the model with the measured current, held as
an interpolant over your data. By default that interpolant is compressed —
samples are dropped wherever doing so stays within tolerance — which keeps the
solve fast on long or densely sampled traces.
Reach for these when the driving current is being distorted — a short pulse
flattened, or a sharp transition rounded off — because the compression
tolerance is loose relative to the feature you care about. Tighten
interpolant_atol / interpolant_rtol, or set interpolant_lossless=True to
reproduce the trace exactly:
A lossless interpolant assumes the trace’s time column is monotonically
non-decreasing. It also costs solve time on a dense trace, which is why
compression is the default — prefer tightening the tolerances first, and
reserve
interpolant_lossless for when you need the input reproduced exactly.Tuning the auto-built solver
Simulation-backed objectives (CurrentDriven, Pulse, CalendarAgeing, CycleAgeing, MSMRFullCell, …) build an IonworksSolver for you when no explicit solver is provided. Pass solver_kwargs inside simulation_kwargs to override individual pieces of that default without restating the rest:
- Nested
optionsare merged over the default IDAKLU options. For example,{"options": {"compile": True}}flips on model compilation but keeps every other tuned option. - Other top-level keys (
atol,rtol,on_extrapolation, …) override the corresponding default solver kwargs.
solver_kwargs is ignored (with a warning) when an explicit solver is supplied — configure those on the solver instance directly. It is also ignored when the model’s default solver isn’t IDAKLU-based.
Forwarding kwargs to the runtime solve
simulation_kwargs also accepts solve_kwargs, a dict forwarded to the runtime sim.solve(...) call on every objective evaluation. Use it for arguments that belong on the solve itself rather than the solver — for example starting_solution to warm-start from a previous solution, or any other pybamm.Simulation.solve argument.
solve_kwargsis applied regardless of whether the objective auto-built the solver or you supplied an explicitsolver. It is the recommended way to pass solve-time arguments that work with any solver.solver_kwargs(above) tunes the auto-built solver at construction time;solve_kwargsconfigures each solve call. The two are independent and can be combined.- Keys the objective controls directly —
inputs,initial_soc,t_eval,t_interp,frequencies— are reserved and raise aValueErrorif passed viasolve_kwargs. - For
CycleAgeing,save_at_cyclesis derived automatically from the metrics; any value passed viasolve_kwargsis ignored with a warning so that the cycles required by the metrics are preserved. - For
CurrentDrivenfits against models with open-circuit potential hysteresis enabled, pass"direction": "charge"or"direction": "discharge"viasolve_kwargsto seed the initial hysteresis state on the corresponding OCP branch. Without a direction, the model starts on its default branch, which can bias the first few seconds of the predicted voltage — and, for short experiments, the fitted parameters. Matchdirectionto whichever half-cycle the dataset represents (typically the sign of the measured current).
CycleAgeing: automatic store_first_last for first/last-only metrics
CycleAgeing lets you supply metrics — a mapping from each objective variable to a .by_cycle() metric that pulls the value of interest out of the simulation. Defaults are provided for "LLI [%]", "LAM_ne [%]", and "LAM_pe [%]", all of which read a single per-step sample.
When every metric in that mapping reads only the first or last sample of a step — i.e. the defaults, or any First/Last .by_cycle() metric — CycleAgeing now defaults solver_kwargs["store_first_last"] to True. The solver then stores only the endpoints of each step, which is far more memory-light for long cycling solves and produces identical results for these metrics.
The flag is only auto-set when it is safe to do so:
- Metrics that read interior points (e.g.
Mean(...).by_cycle()) leave the default off so no samples are dropped. - Composed metrics (arithmetic of
First/Last) are conservatively left alone. - An explicit
store_first_lastinsolver_kwargsis always respected. - Supplying your own
solverskips solver-kwargs injection entirely (as elsewhere).
CycleAgeing: unified experiment model for cheaper cycling
Long cycling protocols repeat the same handful of steps thousands of times. By default pybamm builds a separate switching model per step, which is wasteful when every cycle is the same shape. CycleAgeing now defaults simulation_kwargs["experiment_model_mode"] to "unified", so a single switching model covers the whole experiment — much cheaper to build and solve for repeated cycling, with identical results.
The default is applied whenever an experiment is available (passed as the experiment option, or generated from data via experiment="from data"). It is only a default: any explicit experiment_model_mode you pass in simulation_kwargs is respected.
Only
CycleAgeing sets this default — other simulation-backed objectives keep pybamm’s usual experiment_model_mode. If you need the same behaviour on a different objective, pass experiment_model_mode="unified" in simulation_kwargs explicitly.Objective Functions (theory)
Residual vs. canonical form, MLE interpretation.
Data Fitting overview
Putting objectives, parameters, and optimisers together.