pandas.Series.filter#
- Series.filter(arg=None, /, like=None, regex=None, axis=None, *, items=None, cond=None, na=False)[source]#
Subset the rows or columns according to a boolean mask or the labels.
Rows (or columns with
axis=1) are kept where a boolean mask is True, or selected by their labels withitems,like, orregex.- Parameters:
- argcallable, expression, or list-like, optional
Positional-only. A callable or an expression created with
pandas.col()is a boolean mask, seecond. Any other list-like selects labels, seeitems. A list-like of booleans also selects labels, but since it is likely intended as a mask a warning is issued; passitemsorcondinstead to be explicit.- likestr
Keep labels from axis for which “like in label == True”. This will be deprecated in a future version; use
obj.filter(lambda obj: obj.columns.astype(str).str.contains(like, regex=False), axis=1)instead.- regexstr (regular expression)
Keep labels from axis for which re.search(regex, label) == True. This will be deprecated in a future version; use
obj.filter(lambda obj: obj.columns.astype(str).str.contains(regex), axis=1)instead.- axis{0 or ‘index’, 1 or ‘columns’, None}, default None
The axis to filter on, expressed either as an index (int) or axis name (str). Defaults to the index for a boolean mask, and to the info axis (‘columns’ for
DataFrame) when selecting labels. An expression only supports the index. ForSeriesthis parameter is unused and defaults toNone.- itemslist-like, optional
Keep labels from axis which are in
items. This will be deprecated in a future version; useDataFrame.select()when all ofitemsare present in the columns, orobj.loc[:, pd.Index(items).intersection(obj.columns)](or the equivalent for the index) otherwise.- condarray-like of bool, callable, or expression, optional
A boolean mask selecting the entries to keep. A
Seriesis aligned with the labels of the filtered axis; any other array-like must have the same length as that axis. A callable is called with the object and must return a boolean mask. An expression such aspd.col("a") > 1is evaluated against the DataFrame and is only supported withaxis=0.- na{“raise”, True, False}, default False
How to treat missing values in a boolean mask.
TrueorFalsetreats missing values as that value, matchingobj[mask]for a mask with nullable boolean dtype;"raise"raises aValueError. Ignored when selecting labels.
- Returns:
- Same type as caller
The filtered subset of the DataFrame or Series.
- Raises:
- TypeError
If none or more than one of the positional argument,
items,cond,like, andregexis passed, or ifcondis not a one-dimensional boolean mask.- ValueError
If a mask contains missing values and
na="raise", if a boolean array is not one-dimensional, or if an expression is passed withaxis=1.- IndexError
If a mask that is not a Series has a different length than the filtered axis.
- IndexingError
If a Series mask cannot be aligned with the filtered axis.
- Warns:
- UserWarning
If a list-like of booleans is passed positionally.
See also
DataFrame.locAccess a group of rows and columns by label(s) or a boolean array.
DataFrame.whereReplace values where the condition is False.
Notes
The positional argument is a boolean mask only when it is a callable or an expression; any other value, including a list-like of booleans, selects labels. Use
condto filter with a boolean array orSeries.Selecting labels with
items,like, orregexwill be deprecated in a future version.Examples
>>> df = pd.DataFrame( ... {"one": [1, 4], "two": [2, 5], "three": [3, 6]}, ... index=["mouse", "rabbit"], ... ) >>> df one two three mouse 1 2 3 rabbit 4 5 6
Filter rows with a boolean Series.
>>> df.filter(cond=df["two"] > 2) one two three rabbit 4 5 6
The same using an expression or a callable, which are convenient in method chains and may be passed positionally.
>>> df.filter(pd.col("two") > 2) one two three rabbit 4 5 6 >>> df.filter(lambda df: df["two"] > 2) one two three rabbit 4 5 6
Filter columns with a boolean array.
>>> df.filter(cond=df.columns.str.endswith("e"), axis=1) one three mouse 1 3 rabbit 4 6
Missing values in the mask are treated as False by default; pass
na="raise"to raise instead, orna=Trueto keep them.>>> mask = pd.array([True, None], dtype="boolean") >>> df.filter(cond=mask) one two three mouse 1 2 3 >>> df.filter(cond=mask, na=True) one two three mouse 1 2 3 rabbit 4 5 6
Select columns by their labels.
>>> df.filter(items=["one", "three"]) one three mouse 1 3 rabbit 4 6 >>> df.filter(regex="e$", axis=1) one three mouse 1 3 rabbit 4 6 >>> df.filter(like="bbi", axis=0) one two three rabbit 4 5 6