> ## Documentation Index
> Fetch the complete documentation index at: https://docs.ionworks.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Objective Functions

> Configure objectives and cost functions for data fitting with iws.objectives and iws.costs.

A `DataFit` has two coupled pieces:

* **Objectives** (`iws.objectives.*`) — what experiments to compare model output against.
* **Cost** (`iws.costs.*`) — how the per-point disagreements are aggregated into a single number.

For the math behind each cost, see the [Objective Functions Guide](/guide/data-fitting/objective-functions).

## Available cost functions

| Schema | Formula | When to use |
| - | - | - |
| `iws.costs.SSE()` | $\sum_i r_i^2$ | Default; works with every optimiser |
| `iws.costs.MSE()` | $\frac{1}{N}\sum_i r_i^2$ | Scale-aware mean of squared residuals |
| `iws.costs.RMSE()` | $\sqrt{\frac{1}{N}\sum_i r_i^2}$ | Interpretable units; scalar-only (won't work with residual-array optimisers) |
| `iws.costs.MAE()` | $\frac{1}{N}\sum_i \lvert r_i \rvert$ | Robust to outliers |
| `iws.costs.Max()` | $\max_i \lvert r_i \rvert$ | Minimise the worst-case (largest absolute) residual |
| `iws.costs.Wasserstein()` | $\frac{1}{N}\sum_i \lvert \tilde y_{\text{model},i} - \tilde y_{\text{data},i} \rvert$ | Match distributions (sorted samples) rather than point-wise time series. Set `position_variable` and `weight_variable` for [weighted point-cloud mode](#wasserstein-weighted-point-cloud-mode) |

For MLE, see `iws.costs.GaussianLogLikelihood` — it accepts per-variable noise standard deviations or can estimate them alongside the fitting parameters. It produces a Gaussian negative log-likelihood suitable for Bayesian and MAP estimation.

## Wiring a cost into a fit

```python theme={null}
import pybamm
import ionworks_schema as iws

fit = iws.DataFit(
    objectives={
        "1C": iws.objectives.CurrentDriven(
            data_input="file:.../1C.csv",
            options={"model": pybamm.lithium_ion.SPMe()},
        ),
    },
    parameters={
        "Negative particle diffusivity [m2.s-1]": iws.Parameter(
            "Negative particle diffusivity [m2.s-1]",
            initial_value=2e-14,
            bounds=(1e-14, 1e-13),
        ),
    },
    cost=iws.costs.RMSE(),
)
```

If `cost` is omitted, the optimizer's default cost function is used (typically a least-squares form).

<Note>
  `cost` accepts a cost schema instance (e.g. `iws.costs.RMSE()`) or a config dict with an explicit `type` key (e.g. `{"type": "RMSE"}`). A bare name string like `cost="RMSE"` is rejected with a validation error — wrap it as `{"type": "RMSE"}` instead.
</Note>

## Wasserstein weighted point-cloud mode

By default `iws.costs.Wasserstein()` compares the model and data samples for each objective variable with uniform weights (sorted point-wise comparison). Set both `position_variable` and `weight_variable` to switch to **weighted point-cloud mode**: one variable supplies the positions, the other supplies the (sign-stripped, renormalised) weights, and a single Wasserstein-1 distance is computed per objective.

Use this when you want to match a *density by position* rather than sample-by-sample values — for example, lining up dQ/dV peaks in voltage rather than penalising every dQ/dV residual.

Both `iws.objectives.MSMRFullCell` and `iws.objectives.ElectrodeBalancing` expose the matching `Differential capacity [Ah/V]` values alongside their `Voltage [V] (dQdU)` masked-axis sibling, so either can drive a weighted point-cloud fit:

```python theme={null}
import ionworks_schema as iws

fit = iws.DataFit(
    objectives={
        "ocp": iws.objectives.ElectrodeBalancing(
            data_input="file:.../ocv.csv",
            options={
                "objective variables": [
                    "Differential capacity [Ah/V]",
                    "Voltage [V] (dQdU)",
                ],
            },
        ),
    },
    parameters={...},
    cost=iws.costs.Wasserstein(
        position_variable="Voltage [V] (dQdU)",
        weight_variable="Differential capacity [Ah/V]",
    ),
)
```

<Note>
  `position_variable` and `weight_variable` must be set together — providing only one raises a validation error. Weights are taken as absolute values and renormalised internally, so sign conventions on dQ/dV don't matter. Residual-array output is not available in this mode.
</Note>

## `ElectrodeBalancing` options for OCV fitting

`ElectrodeBalancing` accepts the following keys in its `options` dict to control how the full-cell OCV is processed before the objective is evaluated. These apply regardless of which cost function the fit uses (not only weighted `Wasserstein`):

| Option | Type | Default | Purpose |
| - | - | - | - |
| `objective variables` | list of str | `["Voltage [V]", "Differential voltage [V/Ah]"]` | Variables compared between model and data. Add `"Differential capacity [Ah/V]"` to also emit model dQ/dU on the data voltage grid (plus the masked-axis siblings `"Voltage [V] (dQdU)"` and `"Capacity [A.h] (dQdU)"`) for weighted Wasserstein costs. |
| `dUdQ cutoff` | float \| None | `None` | Drop data points whose `dU/dQ` exceeds this value — useful for masking the near-vertical regions at the voltage limits. |
| `dQdU cutoff` | float \| None | `None` | Drop data points whose `dQ/dU` exceeds this value — useful for masking flat OCV regions where `dQ/dU` diverges. Negative or zero values are always dropped so the resulting weights stay non-negative for Wasserstein. |
| `direction` | `"charge"` \| `"discharge"` \| None | `None` | Direction of the OCV scan. `None` makes no directional assumption. |
| `GITT` | bool | `False` | Treat the data as sparse GITT samples and upsample by interpolation before computing the derivatives. |
| `dQdU model axis` | bool | `False` | When `True`, additionally emit dQ/dV on the model's own full-window voltage axis as `"Differential capacity [Ah/V] (model axis)"` / `"Voltage [V] (model axis)"` — see [Aligning dQ/dV peaks on the model voltage axis](#aligning-dqdv-peaks-on-the-model-voltage-axis). |

```python theme={null}
import ionworks_schema as iws

fit = iws.DataFit(
    objectives={
        "ocp": iws.objectives.ElectrodeBalancing(
            data_input="file:.../ocv.csv",
            options={
                "direction": "discharge",
                "GITT": True,
                "dUdQ cutoff": 1.0,
                "dQdU cutoff": 50.0,
                "objective variables": [
                    "Differential capacity [Ah/V]",
                    "Voltage [V] (dQdU)",
                ],
            },
        ),
    },
    parameters={...},
    cost=iws.costs.Wasserstein(
        position_variable="Voltage [V] (dQdU)",
        weight_variable="Differential capacity [Ah/V]",
    ),
)
```

## Scoping a cost with `calculation_structure`

By default every cost on a `DataFit` consumes every objective and every objective variable in the outputs. Set `calculation_structure` on a cost to scope it explicitly: a mapping from objective name to the list of variable names that cost should compute, or `None` to compute all of that objective's variables (an empty list computes none).

Objectives you leave out of the mapping are not dropped. Inside a `DataFit` each unscoped objective is bound to all of its variables — the same as mapping it to `None` — so scoping one objective (e.g. `{"ocp": ["Voltage [V]"]}` while a `"cc"` objective also exists) still computes `"cc"` in full.

Use this when one cost should only see a subset of variables — most commonly when you pair a per-variable cost (e.g. `SSE`) with a weighted `Wasserstein`. The Wasserstein owns the dQ/dV variables (whose model and data sides may have different lengths by construction), and the SSE is scoped to skip them so the lengths never collide.

```python theme={null}
import ionworks_schema as iws

fit = iws.DataFit(
    objectives={
        "ocp": iws.objectives.ElectrodeBalancing(
            data_input="file:.../ocv.csv",
            options={
                "objective variables": [
                    "Voltage [V]",
                    "Differential capacity [Ah/V] (model axis)",
                    "Voltage [V] (model axis)",
                ],
                "dQdU model axis": True,
            },
        ),
    },
    parameters={...},
    cost=[
        iws.costs.SSE(
            calculation_structure={"ocp": ["Voltage [V]"]},
        ),
        iws.costs.Wasserstein(
            position_variable="Voltage [V] (model axis)",
            weight_variable="Differential capacity [Ah/V] (model axis)",
            calculation_structure={
                "ocp": [
                    "Voltage [V] (model axis)",
                    "Differential capacity [Ah/V] (model axis)",
                ],
            },
        ),
    ],
)
```

<Note>
  `calculation_structure` replaces the deprecated `objective_names` field (a flat list of objective names with no per-variable control). Specifying both on the same cost raises a validation error.
</Note>

### Length-mismatch warning

Element-wise costs (`SSE`, `MSE`, `RMSE`, `MAE`, `Max`) combine the model and data arrays point-by-point, so a variable whose model and data sides have different lengths almost never gives a meaningful score. At fit setup, `DataFit` checks the shapes of every variable each cost is configured to score and emits a `UserWarning` for each mismatch — for example:

```text theme={null}
UserWarning: variable 'Voltage [V] (model axis)' of objective 'ocp' has mismatched
model/data shapes ((512,) vs (128,)). An element-wise cost will combine them
point-by-point, which is almost never intended. Scope the cost with an explicit
`calculation_structure` so each variable is compared against a matching-length
counterpart.
```

The check runs once at fit setup (not on every objective evaluation), so it has no impact on fit performance. When you see this warning, scope the cost with [`calculation_structure`](#scoping-a-cost-with-calculation_structure) so it only sees variables whose model and data lengths match — and route any model-axis variables to a `Wasserstein` cost (or another distribution metric) instead. Distribution costs like `Wasserstein` are skipped by the check, since unequal-length sample sets are expected there.

## Aligning dQ/dV peaks on the model voltage axis

`iws.objectives.ElectrodeBalancing` can emit dQ/dV on the **model's own full-window voltage axis** in addition to (or instead of) the data voltage grid. Set `dQdU model axis: True` in `options` and add the two model-axis variables — `"Differential capacity [Ah/V] (model axis)"` and `"Voltage [V] (model axis)"` — to `objective variables`.

Use this when you want a weighted cost (typically `Wasserstein` in [point-cloud mode](#wasserstein-weighted-point-cloud-mode)) to *position-shift* — i.e. align dQ/dV peaks in voltage rather than residual-by-residual on the data grid. The model and data sides have different lengths by construction, so only a weighted cost should consume them; pair them with a sibling per-variable cost scoped via `calculation_structure` (see above) to keep the rest of the fit honest.

The existing data-axis variables (`"Differential capacity [Ah/V]"` plus the masked siblings `"Voltage [V] (dQdU)"` / `"Capacity [A.h] (dQdU)"`) remain available — both axes can be requested side by side.

## Available objectives

| Schema | Use for |
| - | - |
| `iws.objectives.CurrentDriven(data_input=..., options={...})` | Time-series voltage vs. current loads (drive cycles, custom loads) |
| `iws.objectives.Pulse(data_input=..., options={...})` | Pulse experiments — GITT, HPPC, ICI — with optional feature-extraction variants |
| `iws.objectives.OCPHalfCell(electrode=..., data_input=...)` | Half-cell OCP curves |
| `iws.objectives.MSMRHalfCell(...)` | Fit MSMR parameters to half-cell data |
| `iws.objectives.MSMRFullCell(...)` | Fit MSMR parameters to full-cell data. Supports `Differential voltage [V/Ah]` and `Differential capacity [Ah/V]` as objective variables |
| `iws.objectives.ElectrodeBalancing(...)` | Stoichiometry windows from full-cell discharge. Supports `Differential voltage [V/Ah]` and `Differential capacity [Ah/V]` as objective variables — see [`ElectrodeBalancing` options](#electrodebalancing-options-for-ocv-fitting) |
| `iws.objectives.EIS(...)` | Electrochemical impedance spectra. Supports refined particle meshes via `simulation_kwargs` and multi-SOC joint fits — see [`EIS` options](#eis-fitting-impedance-spectra) |
| `iws.objectives.Resistance(...)` | DC resistance extracted from pulse data |
| `iws.objectives.CalendarAgeing(...)` / `iws.objectives.CycleAgeing(...)` | Ageing curves |

Combine several by passing a `dict[str, objective]` to `DataFit.objectives`.

The objectives that run a simulation — `CurrentDriven`, `Pulse`, `EIS`, `CalendarAgeing`, and `CycleAgeing` — need a model to simulate against, and construction fails without one. Give it as `options={"model": pybamm.lithium_ion.SPMe()}`, or point at a model stored on the platform with `options={"parameterized_model_id": "<id>"}`. The remaining objectives fit data directly and take no model.

`MSMRHalfCell` and `MSMRFullCell` project fitted host-site fractions `Xj` onto the bounded simplex so they sum to exactly 1 and stay within bounds. `MSMRFullCell` always does this; on `MSMRHalfCell` it is the `"project"` default of the `constrain Xj method` option, and the recommended setting.

`constrain Xj method` also accepts `"reformulate"` and `"explicit"`, which are kept for reproducing older fits. They enforce the constraint more weakly: `"reformulate"` replaces the final `Xj` with the complement of the others, and `"explicit"` applies a soft constraint the optimizer can trade away. Neither can now pass unnoticed — both objectives check the fitted fractions after the fit and fail when `|sum(Xj) - 1|` exceeds 0.05 or any `Xj` is negative, warning above `1e-6`. Set the `DataFit` `"validate"` option to `False` to skip the check and get the unphysical values back. Set `penalize Xj complement bounds` on `MSMRHalfCellOptions` to additionally penalize a reformulated complement that falls outside the final `Xj` bounds.

## `GITTModel`: diffusion-only model for GITT and pulse fits

`GITTModel` is a fitting-only model intended for extracting **solid-phase diffusivities** (and a single lumped ohmic resistance) from GITT or pulse-relaxation measurements. It solves x-averaged spherical particle diffusion in each modelled electrode, with the surface flux set by the applied current, and computes the cell voltage from the electrode open-circuit potentials evaluated at the particle-surface stoichiometries, minus an ohmic drop through a lumped `"Ohmic resistance [Ohm]"` parameter.

There are no reaction kinetics (Butler-Volmer), no electrolyte dynamics, and no thermal effects — all parameters are constant except the OCPs. Use it when you want fast, well-conditioned fits to diffusion-dominated portions of GITT or pulse data, and reach for `SPM` / `SPMe` / `DFN` when you need a full physics simulation.

Select the cell configuration via the `"working electrode"` option:

| `"working electrode"` | Configuration |
| - | - |
| `"both"` (default) | Full cell. Both electrodes are modelled. A positive (discharge) current delithiates the negative electrode and lithiates the positive electrode. |
| `"positive"` | Half-cell against a lithium-metal counter electrode (pybamm half-cell convention). Only the working electrode is modelled. A positive current lithiates the working electrode (discharge for a cathode material, charge for an anode material). Anode-material half cells also use `"positive"` — rename the anode's parameters to the positive convention first. |

Each modelled electrode is parameterised with the standard full-cell parameter names (thickness, active material volume fraction, particle radius, diffusivity, OCP, maximum and initial concentrations), plus the current function, electrode cross-sectional area, initial temperature, and `"Ohmic resistance [Ohm]"`.

### Fitting a full-cell GITT measurement

```python theme={null}
import ionworks_schema as iws

fit = iws.DataFit(
    objectives={
        "gitt": iws.objectives.Pulse(
            data_input="file:.../gitt.csv",
            options={
                "model": iws.models.GITTModel(),
            },
        ),
    },
    parameters={
        "Negative particle diffusivity [m2.s-1]": iws.Parameter(
            "Negative particle diffusivity [m2.s-1]",
            initial_value=2e-14,
            bounds=(1e-15, 1e-12),
        ),
        "Positive particle diffusivity [m2.s-1]": iws.Parameter(
            "Positive particle diffusivity [m2.s-1]",
            initial_value=2e-15,
            bounds=(1e-16, 1e-13),
        ),
        "Ohmic resistance [Ohm]": iws.Parameter(
            "Ohmic resistance [Ohm]",
            initial_value=0.02,
            bounds=(1e-3, 1e-1),
        ),
    },
)
```

### Fitting a half-cell pulse measurement

Pass `"working electrode": "positive"` to model a single electrode against a lithium-metal counter. Only the working-electrode parameters are needed.

```python theme={null}
import ionworks_schema as iws

fit = iws.DataFit(
    objectives={
        "pulse": iws.objectives.Pulse(
            data_input="file:.../half_cell_pulse.csv",
            options={
                "model": iws.models.GITTModel(
                    options={"working electrode": "positive"},
                ),
            },
        ),
    },
    parameters={
        "Positive particle diffusivity [m2.s-1]": iws.Parameter(
            "Positive particle diffusivity [m2.s-1]",
            initial_value=2e-15,
            bounds=(1e-16, 1e-13),
        ),
        "Ohmic resistance [Ohm]": iws.Parameter(
            "Ohmic resistance [Ohm]",
            initial_value=0.02,
            bounds=(1e-3, 1e-1),
        ),
    },
)
```

<Note>
  `"working electrode"` only accepts `"both"` or `"positive"` — anything else fails schema validation. Any other keys in `options` are forwarded to the underlying battery-model options for parameter bookkeeping; they do not change the diffusion-only physics.
</Note>

## `EIS`: fitting impedance spectra

`iws.objectives.EIS` compares model impedance against measured electrochemical impedance spectroscopy (EIS) data in the frequency domain, using PyBaMM's `EISSimulation`. Use it to identify kinetic parameters — the two electrodes' reference exchange-current density (the charge-transfer arc diameters) and double-layer capacity (the arc frequencies) — plus the ohmic offset, which time-domain discharges alone leave degenerate.

The data must contain `Frequency [Hz]`, `Z_Re [Ohm]`, and `Z_Im [Ohm]` columns (capacitive band with `Z_Im < 0`, matching the SDK's [EIS upload validator](/data/format)).

<Note>
  The model must be built with `"surface form": "differential"` — `EISSimulation` rejects the default form, and `"algebraic"` drops the double-layer capacity being fitted.
</Note>

### Refining the mesh via `simulation_kwargs`

`options["simulation_kwargs"]` forwards mesh and discretisation kwargs (`var_pts`, `submesh_types`, `geometry`, `spatial_methods`) straight through to `EISSimulation` — the same shape you already use for `CurrentDriven` / `Pulse` objectives. Use it to refine the particle mesh when the default resolution smears the charge-transfer arc.

Time-domain-only keys (`solver`, `solver_kwargs`, `solve_kwargs`, `output_variables`, `experiment`, …) are silently dropped with an info log, so a `simulation_kwargs` dict shared with a time-domain objective is accepted without raising.

```python theme={null}
import ionworks_schema as iws
import pybamm

model = pybamm.lithium_ion.SPMe(
    options={"surface form": "differential"},
)

fit = iws.DataFit(
    objectives={
        "eis": iws.objectives.EIS(
            data_input="file:.../eis.csv",
            options={
                "model": model,
                "simulation_kwargs": {
                    # Refine the positive-particle mesh — same shape as for
                    # CurrentDriven / Pulse objectives.
                    "var_pts": {"r_p": 40},
                },
            },
        ),
    },
    parameters={...},
)
```

### Fitting across multiple SOC set-points

Fit one `EIS` objective per SOC set-point and combine them in the same `DataFit`. Each objective pins its operating point with an objective-level `Initial SOC [%]` (or `Initial voltage [V]`) parameter, and the shared kinetic parameters are identified jointly across all spectra:

```python theme={null}
import ionworks_schema as iws
import pybamm

model = pybamm.lithium_ion.SPMe(
    options={"surface form": "differential"},
)

fit = iws.DataFit(
    objectives={
        f"eis_soc{int(soc * 100)}": iws.objectives.EIS(
            data_input=f"db:<eis-soc{int(soc * 100)}-measurement-id>",
            options={"model": model},
            parameters={"Initial SOC [%]": soc * 100},
        )
        for soc in (0.25, 0.50, 0.75)
    },
    parameters={
        "Positive electrode reference exchange-current density [A.m-2]": iws.Parameter(
            "j0ref_p", initial_value=1.0, bounds=(0.1, 10.0),
        ),
        "Positive electrode double-layer capacity [F.m-2]": iws.Parameter(
            "C_dl_p", initial_value=0.1, bounds=(0.01, 1.0),
        ),
    },
)
```

## Specifying `data_input`

Every objective's `data_input` (and any other `data` field on a calculation or interpolant) accepts the same set of forms:

* A reference string: `"db:<id>"` to reference an uploaded measurement. `"file:..."` and `"folder:..."` are read from your local machine and inlined into the config by the API client on submit, so they work both locally and when you submit a fit to Ionworks — subject to the same 1,000-row inline limit as a bare `DataFrame`. For larger datasets, upload a measurement and reference it with `"db:<id>"`.
* An `ionworksdata.DataLoader` (local or fetched with `DataLoader.from_db(...)`).
* A bare pandas or polars `DataFrame` of pre-loaded columns.
* An `ionworksdata.AnalysisLoader`, which reads one of a measurement's stored [analyses](/data/analyses) as the data instead of its time series — see [Taking the targets from a stored analysis](#taking-the-targets-from-a-stored-analysis).

```python theme={null}
import pybamm
import ionworks_schema as iws
import pandas as pd

df = pd.DataFrame(
    {
        "Time [s]": [...],
        "Voltage [V]": [...],
        "Current [A]": [...],
    }
)

obj = iws.objectives.CurrentDriven(
    data_input=df,
    options={"model": pybamm.lithium_ion.SPMe()},
)
```

When a bare `DataFrame` is passed, it is auto-wrapped on serialization to match the parser's expected `{"data": <columns>}` shape — so `data_input=df` and `data_input={"data": df}` behave the same. String paths and already-wrapped dicts are left untouched.

<Note>
  A string `data_input` must start with `db:`, `file:`, or `folder:` to say where the data is read from — a bare path such as `"data/1C.csv"` is rejected when the objective is constructed. Write it as `"file:data/1C.csv"`.
</Note>

A dict `data_input` is matched against the payload shapes above by its keys — `{"time_series": ..., "steps": ...}`, `{"data": ..., "options": ...}`, `{"data": "db:<id>", "analysis": ...}`, or `{"data": ..., "metadata": ...}` — and validated strictly, so a misspelt key or loading option is reported when the objective is constructed rather than being ignored. A dict matching none of those shapes is treated as a plain mapping of column names to values. `time_series`, `steps`, `data`, `options`, `time_range`, `analysis`, and `metadata` are therefore reserved: a column literally named one of them is read as a payload key instead of as a column.

<Note>
  Inline DataFrames are capped at 1,000 rows per call. For larger datasets, upload as a measurement and reference it by ID instead. See [inline time series size limit](/data/reading#inline-time-series-size-limit).
</Note>

## Generating a CycleAgeing experiment from data

`iws.objectives.CycleAgeing` normally requires an explicit `pybamm.Experiment` describing the cycling protocol. When the protocol is already encoded in the cycler step information attached to your data, set `experiment="from data"` to skip rebuilding it by hand. The experiment is generated lazily, when the fit starts, by calling `DataLoader.generate_experiment()` on the loaded step table.

Use this when:

* The fitted data carries its own step information (a local `ionworksdata.DataLoader`, or one fetched with `DataLoader.from_db(...)`).
* You want the simulated protocol to track the measurement protocol exactly — the per-step current, power, or voltage setpoint the cycler recorded, and either the duration of each step or the endpoint it reached (see [Ending each step on its endpoint](#ending-each-step-on-its-endpoint)).

```python theme={null}
import pybamm
import ionworks_schema as iws

fit = iws.DataFit(
    objectives={
        "ageing": iws.objectives.CycleAgeing(
            data_input="db:<measurement-id>",
            options={
                "model": pybamm.lithium_ion.SPM(options={"SEI": "ec reaction limited"}),
                "experiment": "from data",
                "objective variables": ["LLI [%]"],
            },
        ),
    },
    parameters={...},
)
```

If the data you are fitting against (for example, a per-cycle summary table) is a different object from the measurement that defines the protocol, pass a separate `DataLoader` as `experiment` instead — the steps come from that loader, while the residuals are still computed against `data_input`:

```python theme={null}
import pybamm
import ionworksdata as iwdata
import ionworks_schema as iws

protocol = iwdata.DataLoader.from_db("<protocol-measurement-id>")

fit = iws.DataFit(
    objectives={
        "ageing": iws.objectives.CycleAgeing(
            data_input="db:<summary-measurement-id>",
            options={
                "model": pybamm.lithium_ion.SPM(),
                "experiment": protocol,
                "objective variables": ["LLI [%]"],
            },
        ),
    },
    parameters={...},
)
```

<Note>
  `experiment="from data"` requires `data_input` to resolve to a `DataLoader` (or a dict whose `"data"` entry is a `DataLoader`) that carries step information. When you pass a separate `DataLoader` as `experiment`, that loader must carry the step information instead. Either way, configurations missing steps fail fast at objective construction with a clear error, before any simulation runs.
</Note>

### Taking the targets from a stored analysis

A per-cycle summary is often already stored on the platform as an [analysis](/data/analyses) of the measurement it was extracted from — for example degradation modes per RPT. Point `data_input` at it with `ionworksdata.AnalysisLoader`, and the fit reads that analysis as its targets while the same measurement's steps give the experiment:

```python theme={null}
import pybamm
import ionworksdata as iwdata
import ionworks_schema as iws
from ionworks import AnalysisType

measurement_id = "<measurement-id>"

fit = iws.DataFit(
    objectives={
        "ageing": iws.objectives.CycleAgeing(
            data_input=iwdata.AnalysisLoader(
                measurement_id, analysis_type=AnalysisType.LAM_LLI_FROM_RPT
            ),
            options={
                "model": pybamm.lithium_ion.SPM(options={"SEI": "ec reaction limited"}),
                "experiment": iwdata.DataLoader.from_db(measurement_id),
                "objective variables": ["LLI [%]", "LAM_ne [%]", "LAM_pe [%]"],
            },
        ),
    },
    parameters={...},
)
```

The same measurement ID appears twice on purpose: its steps define the protocol, and its linked analysis supplies the targets. Both are resolved on the platform when the fit is submitted.

`AnalysisLoader` selects the analysis by `analysis_id`, `analysis_type`, and/or `name`; every field given must match, except that `analysis_id` names one analysis exactly and any other fields are then ignored. Give at least one — an empty selector is rejected — and separate two analyses of the same type by `name`. A selector that matches more than one analysis raises rather than picking one. The equivalent dict form is `{"data": "db:<measurement-id>", "analysis": {"analysis_type": "lam_lli_from_rpt"}}` (`iws.AnalysisSpec` builds the `analysis` part). An analysis is a feature table, not a trace, so it cannot be combined with `options` or `time_range`.

To check which analysis a selector will pick before submitting, use [`client.analysis.find_one`](/data/analyses#finding-one-analysis-on-a-measurement).

### Ending each step on its endpoint

By default each generated step runs for the duration the cycler recorded. Set `termination="events"` to end it where the measured step ended instead, in the variable it was controlling — a constant-current or constant-power step at its final voltage, a voltage hold at its final current. The step then carries no duration, so its length is the model's own:

```python theme={null}
options={
    "model": pybamm.lithium_ion.SPM(options={"SEI": "ec reaction limited"}),
    "experiment": "from data",
    "termination": "events",
    "objective variables": ["LLI [%]"],
}
```

Which one you want follows from what the protocol was driving:

* **`"events"`** for endpoint-driven protocols — rate capability, CC-CV cycling — where the question is *how much capacity to this endpoint?* A model whose capacity differs from the cell's then disagrees in capacity rather than ending a discharge at a dangling voltage.
* **`"duration"`** (the default) for time-driven protocols, where the question is *what voltage at this time?*

<Warning>
  `"events"` replays the endpoint each step reached, not the cycler's intent — the steps table cannot tell a cut-off from a time limit, because a step whose voltage moves in one direction always ends at its own extreme either way. On a step whose voltage barely moves, such as a short pulse, the endpoint can already be satisfied at the model's initial state; PyBaMM warns and skips such a step rather than solving it.
</Warning>

### Simulating a voltage hold instead of replaying its current

A generated experiment emits each constant-voltage step as a current interpolant: the model is driven with the current the cycler recorded. During a hold that current is what the cell's impedance produced, so prescribing it asks the model to reproduce the cell's impedance rather than to simulate the hold. Set `use_cv=True` to emit holds as voltage holds instead — and with `termination="events"`, each one ends at its measured final current, the way the cycler's taper cut-off did:

```python theme={null}
options={
    "model": pybamm.lithium_ion.SPM(options={"SEI": "ec reaction limited"}),
    "experiment": "from data",
    "use_cv": True,
    "termination": "events",
    "objective variables": ["LLI [%]"],
}
```

<Warning>
  Pair `use_cv=True` with `termination="events"`. Under `"duration"` the constant-current leg before a hold stops on the clock, which can leave the model a hundred millivolts or more below the hold's setpoint — and the hold then draws whatever current closes that gap, measured at tens of times the cell's nominal rate. That current is the right answer to the question asked — moving a terminal voltage that far that fast does take it — so the thing to fix is the experiment, not the solve. A voltage-vs-time residual never shows it, because the voltage is exactly where the hold says it should be.
</Warning>

<Note>
  A taper endpoint drifts as the cell ages, so on a long cycling run `use_cv=True` with `"events"` gives almost every hold its own distinct step. `CycleAgeing` defaults `experiment_model_mode="unified"` precisely to build one model for a repeated step, and distinct holds cost that reuse — cut-off voltages, which repeat exactly, do not.
</Note>

Both options apply only to an experiment generated from data. Passing either alongside an explicit `pybamm.Experiment` raises at objective construction — build what you want into that experiment instead.

### Which cycle each data row describes

Your data's `Cycle number` column says which cycle of the simulation each row is compared against. For a generated experiment it is matched against the steps table's `Cycle count` rather than used as a position: the experiment holds one cycle per `Cycle count`, but numbers its own cycles from zero whatever the table says.

`Cycle count` is not the cycler's own label — that is `Cycle from cycler`. It is a counter `ionworksdata` derives, starting at 0 and running contiguously, so on a cycler labelling its cycles 1, 2, 3 you get:

```
Cycle from cycler : [1, 1, 1, 1, 2, 2, 2, 2, 3, 3, 3, 3]
Cycle count       : [0, 0, 0, 0, 1, 1, 1, 1, 2, 2, 2, 2]
```

Across a whole measurement the counts and the experiment's positions therefore coincide, and it makes no difference which you numbered your rows by. They diverge when the loader covers only part of a measurement: slicing with `first_step` / `last_step` keeps the original counts instead of renumbering, so a loader sliced to start at the second cycle has `Cycle count` `[1, 2]` behind experiment cycles 0 and 1. Number those rows `1, 2` — what the steps table says — not `0, 1`.

A per-cycle summary built with `ionworksdata.cycle_metrics.get_cycle_metrics` already carries the right values in its own `Cycle count` column, so renaming that column to `Cycle number` is all the fit needs.

An experiment you pass in directly carries no `Cycle count` to match against, so there a `Cycle number` stays a 0-based position among its cycles.

A `by_cycle` metric's values are compared with the data row for row, over the rows where that variable has a value. Leave `cycles=` unset and the cycles are taken from the data. Set it and it must name the cycles the data has a value of that variable for — a cycle with no such row raises at build, instead of silently comparing simulation cycle 0 against a row describing a different cycle of the cell.

### Targets measured at different cycles

Not every quantity an ageing test produces is measured at every cycle — a resistance sweep might run only on alternate RPTs. Leave the missing values as `NaN`: a `NaN` in a target column means there is no measured value at that cycle, and that row is left out of the fit for that variable alone. The other variables still use it, so only measured points are scored and nothing needs interpolating onto a shared cycle grid.

```python theme={null}
import numpy as np
import pandas as pd

data = pd.DataFrame(
    {
        "Cycle number": [0, 12, 24, 36, 48],
        "SOH [%]": [100.0, 97.8, 94.0, 91.5, 89.8],  # every RPT
        "RI [%]": [0.0, np.nan, 3.1, np.nan, 6.4],  # alternate RPTs only
    }
)
```

A quantity measured later in the same test than the others occupies its own row, naming the cycle it was measured in.

## `OCPHalfCell` options

| Option | Default | Description |
| - | - | - |
| `theta_ref` | direction-dependent | Reference stoichiometry. A positive electrode defaults so that `0 <= theta <= theta_ref`; a negative electrode so that `theta_ref <= theta <= 1`. |
| `stoichiometry limits` | `(0, 1)` | Stoichiometry window fitted over. Narrow it (e.g. `(0.05, 0.95)`) if the fit struggles near the endpoints. |

```python theme={null}
ocp = iws.objectives.OCPHalfCell(
    electrode="positive",
    data_input=data,
    options={"stoichiometry limits": (0.05, 0.95)},
)
```

## Interpolating the driving current

`CurrentDriven` and `Pulse` drive the model with the measured current, held as
an interpolant over your data. By default that interpolant is **compressed** —
samples are dropped wherever doing so stays within tolerance — which keeps the
solve fast on long or densely sampled traces.

| Option | Default | Description |
| - | - | - |
| `interpolant_atol` | the solver's `atol` when one is supplied in `simulation_kwargs`, else `1e-6` | Absolute tolerance for the compression. |
| `interpolant_rtol` | the solver's `rtol` when one is supplied, else `1e-4` | Relative tolerance for the compression. |
| `interpolant_lossless` | `False` | Use the current samples exactly, with no compression. The two tolerances above are then ignored. |

Reach for these when the driving current is being distorted — a short pulse
flattened, or a sharp transition rounded off — because the compression
tolerance is loose relative to the feature you care about. Tighten
`interpolant_atol` / `interpolant_rtol`, or set `interpolant_lossless=True` to
reproduce the trace exactly:

```python theme={null}
pulse = iws.objectives.Pulse(
    data_input=data,
    options={"model": model, "interpolant_lossless": True},
)
```

<Note>
  A lossless interpolant assumes the trace's time column is monotonically
  non-decreasing. It also costs solve time on a dense trace, which is why
  compression is the default — prefer tightening the tolerances first, and
  reserve `interpolant_lossless` for when you need the input reproduced exactly.
</Note>

## Tuning the auto-built solver

Simulation-backed objectives (`CurrentDriven`, `Pulse`, `CalendarAgeing`, `CycleAgeing`, `MSMRFullCell`, …) build an `IonworksSolver` for you when no explicit `solver` is provided. Pass `solver_kwargs` inside `simulation_kwargs` to override individual pieces of that default without restating the rest:

* Nested `options` are merged over the default IDAKLU options. For example, `{"options": {"compile": True}}` flips on model compilation but keeps every other tuned option.
* Other top-level keys (`atol`, `rtol`, `on_extrapolation`, …) override the corresponding default solver kwargs.

`solver_kwargs` is ignored (with a warning) when an explicit `solver` is supplied — configure those on the solver instance directly. It is also ignored when the model's default solver isn't IDAKLU-based.

```python theme={null}
import pybamm
import ionworks_schema as iws

fit = iws.DataFit(
    objectives={
        "1C": iws.objectives.CurrentDriven(
            data_input="file:.../1C.csv",
            options={
                "model": pybamm.lithium_ion.SPMe(),
                "simulation_kwargs": {
                    "solver_kwargs": {
                        "options": {"compile": True},
                        "atol": 1e-8,
                    },
                },
            },
        ),
    },
    parameters={...},
)
```

<Tip>
  Enabling `compile` ahead of time (`{"options": {"compile": True}}`) trades a one-off compilation cost for faster repeated evaluations — useful when the same objective is solved many times during a fit or sweep.
</Tip>

### Forwarding kwargs to the runtime solve

`simulation_kwargs` also accepts `solve_kwargs`, a dict forwarded to the runtime `sim.solve(...)` call on every objective evaluation. Use it for arguments that belong on the solve itself rather than the solver — for example `starting_solution` to warm-start from a previous solution, or any other `pybamm.Simulation.solve` argument.

* `solve_kwargs` is applied regardless of whether the objective auto-built the solver or you supplied an explicit `solver`. It is the recommended way to pass solve-time arguments that work with any solver.
* `solver_kwargs` (above) tunes the auto-built solver at *construction* time; `solve_kwargs` configures each *solve call*. The two are independent and can be combined.
* Keys the objective controls directly — `inputs`, `initial_soc`, `t_eval`, `t_interp`, `frequencies` — are reserved and raise a `ValueError` if passed via `solve_kwargs`.
* For `CycleAgeing`, `save_at_cycles` is derived automatically from the metrics; any value passed via `solve_kwargs` is ignored with a warning so that the cycles required by the metrics are preserved.
* For `CurrentDriven` fits against models with **open-circuit potential hysteresis** enabled, pass `"direction": "charge"` or `"direction": "discharge"` via `solve_kwargs` to seed the initial hysteresis state on the corresponding OCP branch. Without a direction, the model starts on its default branch, which can bias the first few seconds of the predicted voltage — and, for short experiments, the fitted parameters. Match `direction` to whichever half-cycle the dataset represents (typically the sign of the measured current).

```python theme={null}
import pybamm
import ionworks_schema as iws

# SPMe with a current-sigmoid hysteresis submodel on the negative OCP
model = pybamm.lithium_ion.SPMe(
    options={"open-circuit potential": ("current sigmoid", "single")},
)

fit = iws.DataFit(
    objectives={
        "discharge_1C": iws.objectives.CurrentDriven(
            data_input="file:.../1C_discharge.csv",
            options={
                "model": model,
                "simulation_kwargs": {
                    # Start on the discharge branch of the hysteresis loop.
                    "solve_kwargs": {"direction": "discharge"},
                },
            },
        ),
    },
    parameters={...},
)
```

```python theme={null}
import pybamm
import ionworks_schema as iws

fit = iws.DataFit(
    objectives={
        "pulse": iws.objectives.Pulse(
            data_input="file:.../pulse.csv",
            options={
                "model": pybamm.lithium_ion.SPMe(),
                "simulation_kwargs": {
                    # Tunes the auto-built solver (construction time):
                    "solver_kwargs": {"options": {"compile": True}},
                    # Forwarded to every sim.solve(...) call (runtime).
                    # prior_solution is a pybamm.Solution you obtained from an
                    # earlier sim.solve(...) — substitute your own:
                    "solve_kwargs": {"starting_solution": prior_solution},
                },
            },
        ),
    },
    parameters={...},
)
```

### `CycleAgeing`: automatic `store_first_last` for first/last-only metrics

`CycleAgeing` lets you supply `metrics` — a mapping from each objective variable to a `.by_cycle()` metric that pulls the value of interest out of the simulation. Defaults are provided for `"LLI [%]"`, `"LAM_ne [%]"`, and `"LAM_pe [%]"`, all of which read a single per-step sample.

When *every* metric in that mapping reads only the first or last sample of a step — i.e. the defaults, or any `First`/`Last` `.by_cycle()` metric — `CycleAgeing` now defaults `solver_kwargs["store_first_last"]` to `True`. The solver then stores only the endpoints of each step, which is far more memory-light for long cycling solves and produces identical results for these metrics.

The flag is only auto-set when it is safe to do so:

* Metrics that read interior points (e.g. `Mean(...).by_cycle()`) leave the default off so no samples are dropped.
* Composed metrics (arithmetic of `First`/`Last`) are conservatively left alone.
* An explicit `store_first_last` in `solver_kwargs` is always respected.
* Supplying your own `solver` skips solver-kwargs injection entirely (as elsewhere).

```python theme={null}
import pybamm
import ionworks_schema as iws

fit = iws.DataFit(
    objectives={
        "ageing": iws.objectives.CycleAgeing(
            data_input="file:.../ageing.csv",
            options={
                "model": pybamm.lithium_ion.SPM(),
                "experiment": "from data",
                "objective variables": ["LLI [%]", "LAM_ne [%]"],
                # Defaults already read first/last only, so store_first_last
                # is enabled automatically. Override explicitly when needed:
                # "simulation_kwargs": {
                #     "solver_kwargs": {"store_first_last": False},
                # },
            },
        ),
    },
    parameters={...},
)
```

### `CycleAgeing`: unified experiment model for cheaper cycling

Long cycling protocols repeat the same handful of steps thousands of times. By default pybamm builds a separate switching model per step, which is wasteful when every cycle is the same shape. `CycleAgeing` now defaults `simulation_kwargs["experiment_model_mode"]` to `"unified"`, so a single switching model covers the whole experiment — much cheaper to build and solve for repeated cycling, with identical results.

The default is applied whenever an experiment is available (passed as the `experiment` option, or generated from data via [`experiment="from data"`](#generating-a-cycleageing-experiment-from-data)). It is only a default: any explicit `experiment_model_mode` you pass in `simulation_kwargs` is respected.

```python theme={null}
import pybamm
import ionworks_schema as iws

fit = iws.DataFit(
    objectives={
        "ageing": iws.objectives.CycleAgeing(
            data_input="file:.../ageing.csv",
            options={
                "model": pybamm.lithium_ion.SPM(options={"SEI": "ec reaction limited"}),
                "experiment": "from data",
                "objective variables": ["LLI [%]"],
                # "unified" is applied automatically. Override to fall back to
                # the legacy per-step model:
                # "simulation_kwargs": {"experiment_model_mode": "legacy"},
            },
        ),
    },
    parameters={...},
)
```

<Note>
  Only `CycleAgeing` sets this default — other simulation-backed objectives keep pybamm's usual `experiment_model_mode`. If you need the same behaviour on a different objective, pass `experiment_model_mode="unified"` in `simulation_kwargs` explicitly.
</Note>

<Tip>
  For most optimisers, `SSE` is the safest choice — it has both a residual-array form and a scalar form, so it's compatible with every algorithm. Use `MSE` or `RMSE` when you need scale-independent reporting.
</Tip>

<CardGroup cols={2}>
  <Card title="Objective Functions (theory)" icon="bullseye" href="/guide/data-fitting/objective-functions">
    Residual vs. canonical form, MLE interpretation.
  </Card>

  <Card title="Data Fitting overview" icon="chart-line" href="/build/parameterize/data-fitting/overview">
    Putting objectives, parameters, and optimisers together.
  </Card>
</CardGroup>
