IO β Saving & Loading Arrays
Definition
NumPy has its own fast, exact binary formats (
.npysingle-array,.npzmulti-array archive) alongside text-based options (savetxt/genfromtxt/loadtxt) for interoperating with CSV-like files.
Binary Format β .npy (Single Array, Recommended for NumPy-to-NumPy)
np.save("array.npy", arr) # save a single array
loaded = np.load("array.npy") # load it back β exact dtype/shape preservedPrefer
.npy/.npzover text formats for pure NumPy workflowsBinary formats preserve dtype and shape exactly, and are dramatically faster to read/write than text-based CSV for large arrays.
.npz β Multiple Arrays in One Archive
np.savez("data.npz", features=X, labels=y) # uncompressed
np.savez_compressed("data.npz", features=X, labels=y) # compressed, smaller file, slower to read/write
loaded = np.load("data.npz")
loaded["features"] # access by the keyword name used when saving
loaded["labels"]
list(loaded.keys()) # see what arrays are storedText Formats
np.savetxt("out.csv", arr, delimiter=",", fmt="%.4f", header="col1,col2,col3")
np.loadtxt("out.csv", delimiter=",", skiprows=1) # simple, fast, but no missing-value handling
np.genfromtxt("out.csv", delimiter=",", skip_header=1, filling_values=np.nan) # handles missing values| Function | Use case |
|---|---|
np.loadtxt | clean, fully rectangular numeric data, no missing values |
np.genfromtxt | messier data β handles missing values, mixed types, comments |
np.savetxt | write an array to a plain-text delimited file |
For anything CSV-heavy with mixed dtypes or missing data, pandas' read_csv is almost always the better tool β NumPy's text I/O is best reserved for clean, purely numeric data.
Reading a Structured/Record Array
data = np.genfromtxt(
"people.csv",
delimiter=",",
names=True, # use header row as field names
dtype=None, # infer per-column dtype
encoding="utf-8"
)
data["age"] # access a named field like a dict keyMemory-Mapped Files (For Arrays Too Large to Fit in RAM)
mmap_arr = np.memmap("big.dat", dtype="float32", mode="r", shape=(100000, 100))
mmap_arr[0:10] # reads only the requested slice from disk, not the whole file
memmapfor out-of-core processingUseful when working with arrays larger than available RAM β data is read from disk on demand rather than loaded entirely upfront.
Pickle (General Python Objects, Including Arrays)
import pickle
with open("arr.pkl", "wb") as f:
pickle.dump(arr, f)
with open("arr.pkl", "rb") as f:
arr = pickle.load(f)Never unpickle files from untrusted sources β pickle can execute arbitrary code.