Categories
Python Answers

How to shuffle python Pandas DataFrame rows?

To shuffle python Pandas DataFrame rows, we call the data frame sample method.

For instance, we write

df.sample(frac=1)

to call sample on the df data frame.

The frac keyword argument specifies the fraction of rows to return in the random sample, so frac=1 means to return all rows in random order.

Categories
Python Answers

How to count the NaN values in a column in Python Pandas DataFrame?

To count the NaN values in a column in Python Pandas DataFrame, we call isna and sum.

For instance, we write

s = pd.Series([1,2,3, np.nan, np.nan])
count = s.isna().sum()

to create a series witth pd.Series.

The we call isna to return the isna values in the series.

And then we call sum to get the count of them.

Categories
Python Answers

How to convert index of a Python Pandas dataframe into a column?

To convert index of a Python Pandas dataframe into a column, we cann reset_index.

For instance, we write

df = df.reset_index(level=0)

to call reset_index on the df data frame to convert index of a Python Pandas dataframe into a column

Categories
Python Answers

How to filter Python Pandas DataFrame by substring criteria?

To filter Python Pandas DataFrame by substring criteria, we can call the str.contains method.

For instance, we write

df[df['A'].str.contains("hello")]

to return a data frame with rows in column 'A' that has strings containing 'hello' with str.contains.

Categories
Python Answers

How to create an empty Python Pandas DataFrame, then filling it?

To create an empty Python Pandas DataFrame, then filling it, we can append the new data to a list and put the list data in the data frame.

For instance, we write

data = []
for a, b, c in some_function_that_yields_data():
    data.append([a, b, c])

df = pd.DataFrame(data, columns=['A', 'B', 'C'])

to call data.append to append the data in the data list.

Then we create the data frame from data using pd.DataFrame with data and set columns to an array with the column names.