MultiIndex β Hierarchical Indexing
Definition
A
MultiIndexlets aSeries/DataFrameaxis have multiple levels of labels β useful for representing higher-dimensional data (e.g. year + month, or region + product) in a 2D structure.
Creating a MultiIndex
# From groupby with multiple keys (most common origin)
s = df.groupby(["region", "quarter"])["sales"].sum() # s.index is a MultiIndex
# Explicitly, from tuples
index = pd.MultiIndex.from_tuples(
[("Mumbai", 2025), ("Mumbai", 2026), ("Pune", 2025)],
names=["city", "year"]
)
df2 = pd.DataFrame({"sales": [100, 120, 90]}, index=index)
# From the cartesian product of levels
index = pd.MultiIndex.from_product(
[["Mumbai", "Pune"], [2025, 2026]],
names=["city", "year"]
)
# From arrays
index = pd.MultiIndex.from_arrays(
[["Mumbai", "Mumbai", "Pune"], [2025, 2026, 2025]],
names=["city", "year"]
)
# By set_index on multiple columns
df.set_index(["city", "year"])Inspecting Levels
df.index.names # names of each level
df.index.levels # unique values per level
df.index.get_level_values("city") # values of one level as an array
df.index.nlevels # number of levelsSelecting Data
df.loc["Mumbai"] # all rows for outer level "Mumbai"
df.loc[("Mumbai", 2025)] # specific (outer, inner) combination
df.loc[("Mumbai", 2025), "sales"] # + column selection
df.loc[["Mumbai", "Pune"]] # multiple outer-level values
df.xs("Mumbai", level="city") # cross-section: fix one level, keep rest
df.xs(2025, level="year") # fix the inner level instead
df.xs(("Mumbai", 2025), level=["city", "year"]) # fix multiple levels at once
.xs()for level-agnostic selection
.xs()is especially useful for selecting on an inner level without needing to specify the outer level(s) β something plain.loc[]canβt do directly.
Reordering & Sorting Levels
df.swaplevel("city", "year") # swap two levels (order changes, values unchanged)
df.sort_index(level="year") # sort by a specific level
df.sort_index() # sort by all levels, outer to inner
df.reorder_levels(["year", "city"]) # arbitrary reorderFlattening a MultiIndex
df.reset_index() # levels become regular columns
df.columns = ["_".join(map(str, c)) for c in df.columns] # flatten MultiIndex COLUMNS to single strings
groupby().agg()with multiple functions often produces a MultiIndex on the columns too β flattening with a list comprehension like above is the standard fix.
Stack / Unstack with MultiIndex
See 08-Reshaping β stack()/unstack() move a level between the row index and the columns.
s.unstack() # move innermost row-index level to columns
s.unstack(level="year") # move a specific named levelMultiIndex on Columns
df.columns = pd.MultiIndex.from_tuples([
("sales", "2025"), ("sales", "2026"), ("profit", "2025")
])
df["sales"] # select the whole "sales" top-level group
df[("sales", "2025")] # select a specific leaf columnNotes & Gotchas
Performance: sort the index for fast slicing
Slicing operations (
df.loc["Mumbai":"Pune"]) on an unsorted MultiIndex can raiseUnsortedIndexErroror silently be slow. Calldf.sort_index()after building a MultiIndex if you plan to slice it.