Save and load results#
PyPhi results are expensive to compute, so you will usually want to store them
and read them back later. Two functions do this: pyphi.save writes a result
to a file, and pyphi.load reads it back into a live object. The same
operations are also available as .save() and .load() methods on each
serializable result type.
import pyphi
pyphi.config.progress_bars = False
First, compute something worth saving. Here pyphi.analyze returns an
Analysis, from which we take the cause-effect structure and the
system-irreducibility analysis.
from pyphi import examples
system = examples.iit4_2023_fig1a_system()
analysis = pyphi.analyze(system.substrate, system.state)
ces = analysis.ces # cause-effect structure (distinctions + relations)
sia = analysis.sia # system-irreducibility analysis (φ_s)
print(f"φ_s = {float(sia.phi):.4f} | {len(ces.distinctions)} distinctions")
φ_s = 0.1339 | 6 distinctions
Save and load#
pyphi.save takes the object and a path; pyphi.load takes the path and
reconstructs the object. Equality is value-based, so a round-trip compares
equal to the original.
from pathlib import Path
from tempfile import mkdtemp
out = Path(mkdtemp())
pyphi.save(ces, out / "ces.json")
restored = pyphi.load(out / "ces.json")
print("type preserved:", type(restored).__name__)
print("round-trips :", restored == ces)
type preserved: CauseEffectStructure
round-trips : True
Every serializable type also has these operations as methods. obj.save(path)
is the same as pyphi.save(obj, path), and the classmethod Type.load(path)
is the same as pyphi.load(path):
ces.save(out / "ces2.json")
CauseEffectStructure = type(ces)
CauseEffectStructure.load(out / "ces2.json") == ces
True
File formats#
The wire format is inferred from the file extension:
.json— a human-readable JSON document..mpk— msgpack, a compact binary encoding of the same document. Prefer it for large structures.a trailing
.gzon either (.json.gz,.mpk.gz) is transparently gzip-compressed on save and decompressed on load.
for name in ["ces.json", "ces.mpk", "ces.json.gz"]:
pyphi.save(ces, out / name)
size = (out / name).stat().st_size
print(f"{name:<12} {size:>7,} bytes")
ces.json 36,316 bytes
ces.mpk 27,496 bytes
ces.json.gz 4,441 bytes
All three round-trip to the same object:
all(pyphi.load(out / name) == ces for name in ["ces.json", "ces.mpk", "ces.json.gz"])
True
To choose the format explicitly instead of inferring it from the extension,
pass format="json" or format="msgpack" to either function.
What is serializable#
Most result types serialize: Substrate, System, cause-effect structures,
system-irreducibility analyses, and actual-causation results all round-trip.
for obj in [system.substrate, system, sia, ces]:
pyphi.save(obj, out / "obj.mpk")
round_tripped = pyphi.load(out / "obj.mpk") == obj
print(f"{type(obj).__name__:<30} {round_tripped}")
Substrate True
System True
SystemIrreducibilityAnalysis True
CauseEffectStructure True
Batch-run results round-trip too: a pyphi.sweep table with its raw
results, and an optimize outcome including the winning substrate and its
analysis. Their DataFrames are embedded in the document as
parquet, so dtypes and NaN values survive
exactly.
import pandas as pd
result = pyphi.sweep(system.substrate, states=[system.state], progress=False)
pyphi.save(result, out / "sweep.json")
loaded = pyphi.load(out / "sweep.json")
pd.testing.assert_frame_equal(loaded.df, result.df)
loaded.df
| phi | normalized_phi | is_irreducible | partition_margin | cause_state_margin | effect_state_margin | effectively_tied | substrate | formalism | subset | state | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 0.133873 | 0.066937 | True | 0.026941 | 0.003492 | 0.030059 | False | 0 | IIT_4_0_2026 | (0, 1, 2) | (0, 1, 1) |
The Analysis returned by pyphi.analyze saves as a whole, carrying its
system, system-irreducibility analysis, and cause-effect structure together, so
you do not have to save the components separately:
pyphi.save(analysis, out / "analysis.json")
reloaded = pyphi.load(out / "analysis.json")
print(f"φ_s = {reloaded.phi:.4f} | Φ = {reloaded.big_phi:.4f}")
φ_s = 0.1339 | Φ = 4.3657
Experiment provenance writers#
Experiment scripts need two things beyond pyphi.save: output files that
never overwrite earlier runs, and a record of how each file was produced.
The writers in pyphi.provenance provide both. Parameters are encoded
into the filename, a repeated save lands in a _v2 file instead of
clobbering the first, and every file embeds a full provenance record —
pyphi version, git commit, timestamp, and seed.
from pyphi import provenance
path = provenance.save_json(
{"phi": 0.133873},
out,
"sweep_study",
params={"seed": 42, "trials": 60},
)
path.name
'sweep_study_seed42_trials60.json'
provenance.save_json(
{"phi": 0.5}, out, "sweep_study", params={"seed": 42, "trials": 60}
).name
'sweep_study_seed42_trials60_v2.json'
save_npz does the same for arrays of raw per-trial data, and
save_dataframe writes a DataFrame as parquet — the format used for
DataFrame outputs throughout PyPhi — with the metadata embedded in the
parquet schema. read_metadata recovers the provenance and parameters
from any of the three formats:
metadata = provenance.read_metadata(path)
{key: metadata["provenance"][key] for key in ("seed", "pyphi_version")}
{'seed': 42, 'pyphi_version': '2.0.0rc3.dev6+g2b98c1d5b'}
Compatibility note#
This serializer is a format break from the jsonify layer used in
PyPhi 1.x. Files written by the old layer cannot be read by pyphi.load, and
files written here are not readable by the old pyphi.jsonify machinery.
Cause-effect structures are also stored compactly: each distinction is
written once in a table, and relations reference their members by index rather
than embedding a full copy of every distinction they contain.
Where to go next#
Export results converts results to DataFrames, NetworkX graphs, and files other tools read.
Cache results persists results automatically across runs.
Run a sweep as an HTCondor campaign saves and collects the pieces of a run sharded across machines.