pandas.DataFrame.select_dtypes#

DataFrame.select_dtypes(include=None, exclude=None)[source]#

Return a subset of the DataFrame’s columns based on the column dtypes.

This method allows for filtering columns based on their data types. It is useful when working with heterogeneous DataFrames where operations need to be performed on a specific subset of data types.

Parameters:
include, excludescalar or list-like

A selection of dtypes or strings to be included/excluded. At least one of these parameters must be supplied.

Returns:
DataFrame

The subset of the frame including the dtypes in include and excluding the dtypes in exclude.

Raises:
ValueError
  • If both of include and exclude are empty

  • If include and exclude have overlapping elements

  • If a datetime64/timedelta64 spec, or an interval spec’s subtype, names a resolution no column can have, e.g. 'datetime64[10s]'

  • If an IntervalDtype or CategoricalDtype spec gives some of its attributes but not all, or leaves an interval subtype’s resolution unset, e.g. pd.IntervalDtype('int64') or 'interval[datetime64]'

TypeError
  • If a numpy string or bytes dtype is passed in, e.g. np.str_, '<U8' or bytes

See also

DataFrame.dtypes

Return Series with the data type of each column.

Notes

  • To select all numeric types, use np.number or 'number'

  • To select strings, use the builtin str, which selects pandas.StringDtype columns and pandas.ArrowDtype string/large_string columns; the string spec 'str' selects only the pandas.StringDtype ones

  • See the numpy dtype hierarchy

  • A dtype instance (e.g. np.dtype("int32") or pd.CategoricalDtype(["a", "b"])) selects only columns with exactly that dtype, whereas a class or string selects a family of dtypes. A bare pd.CategoricalDtype() or pd.IntervalDtype() names the family too, but an instance that gives some of its attributes and not others raises, since no column has such a dtype: pd.IntervalDtype("int64") leaves closed unset, pd.CategoricalDtype(ordered=True) the categories

  • To select datetimes, use np.datetime64, 'datetime' or 'datetime64'

  • To select timedeltas, use np.timedelta64, 'timedelta' or 'timedelta64'

  • To select datetimes or timedeltas of a specific resolution, pass a unit-qualified dtype such as 'datetime64[us]' or 'timedelta64[ms]'; this matches only columns with exactly that resolution, whereas an unqualified spec matches every resolution

  • To select Pandas categorical dtypes, use 'category'

  • To select all timezone-aware datetime dtypes, use 'datetimetz' or pandas.DatetimeTZDtype; a string such as 'datetime64[ns, US/Eastern]' selects only that exact dtype

  • To select all period dtypes, use pandas.PeriodDtype; a string such as 'period[D]' selects only that frequency

  • An ExtensionDtype subclass matches every instance of that subclass regardless of parametrization, e.g. pd.ArrowDtype selects all pyarrow-backed columns and pd.CategoricalDtype selects all categorical columns

Examples

>>> df = pd.DataFrame(
...     {"a": [1, 2] * 3, "b": [True, False] * 3, "c": [1.0, 2.0] * 3}
... )
>>> df
        a      b  c
0       1   True  1.0
1       2  False  2.0
2       1   True  1.0
3       2  False  2.0
4       1   True  1.0
5       2  False  2.0
>>> df.select_dtypes(include="bool")
   b
0  True
1  False
2  True
3  False
4  True
5  False
>>> df.select_dtypes(include=["float64"])
   c
0  1.0
1  2.0
2  1.0
3  2.0
4  1.0
5  2.0
>>> df.select_dtypes(exclude=["int64"])
       b    c
0   True  1.0
1  False  2.0
2   True  1.0
3  False  2.0
4   True  1.0
5  False  2.0