applied · Level 3

Data analysis with Python: investigate and explain a dataset

Load, inspect, clean, group, join and chart a small invented dataset with pandas, then say honestly how sure you are.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 3Building
  • 90 min
  • 6 chapters
  • Free PDF, no account
The Assemblyapplied / 03

Start with the essentials

The short answer

Data analysis with Python starts with a written question. Load the table with pandas, inspect what is really there, and clean a copy in a script that logs every change. Compare like with like, check what each share divides by, and put an interval around your result. Recompute key numbers a second way, and never treat a correlation as proof of a cause.

What you will learn

  • You will be able to load a CSV file with pandas and inspect its shape, types, summaries and category counts.
  • You will be able to clean a copy of a dataset with a script that logs every change and never edits the original.
  • You will be able to compare groups fairly, join a lookup table safely and choose an honest denominator.
  • You will be able to save a chart with matplotlib, read it critically and explain why correlation is not causation.
  • You will be able to put a bootstrap interval around a result and say what your sample size allows you to claim.
  • You will be able to brief an AI assistant without sharing real data and check the analysis code it returns.

Who it is for

Anyone who can run a Python script and wants to answer a real question from a table of data, then explain the answer honestly. You need a terminal, Python 3.12 or later and permission to install packages. No statistics background is needed.

Before you start

  • Comfort running a Python script from a terminal, and with variables, loops, functions, files and CSV. The Python: from first script to a useful automation workbook covers all of this.

Keep learning

The complete workbook

This workbook follows one question through one invented, deliberately messy table of support tickets. You will inspect, clean, group, join and chart it with pandas and matplotlib, measure your uncertainty with a simple bootstrap and make the analysis reproducible. It ends with safe habits for using an AI assistant to write analysis code.

  1. 01
    Start with a question, then meet the data

    Write the question down, set up, make the practice data and look before changing anything.

    In the workbook · 1 exercise
  2. 02
    Clean a copy and log every change

    Cleaning decides what the data means: do it in a script, on a copy, with every decision logged.

    In the workbook · 1 exercise
  3. 03
    Group, compare and choose the denominator

    Split rows into groups, summarise each one and compare like with like.

    In the workbook · 1 exercise
  4. 04
    Join a lookup table and chart it over time

    Bring in a second table safely, then watch the data month by month.

    In the workbook · 1 exercise
  5. 05
    How sure can you be?

    A difference in your sample is not automatically a difference in the world.

    In the workbook · 1 exercise
  6. 06
    Make it reproducible, and use AI safely

    An analysis nobody can rerun or check is only a story.

    In the workbook · 1 exercise

Also inside: a 10-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

Do I need statistics training to analyse data with Python?

Not to start. Counting, grouping, comparing medians and reading a chart need care rather than advanced maths. You do need honesty about limits: check sample sizes, show uncertainty and never treat a correlation as a cause. For decisions that matter, ask someone with statistical training to review your method.

Why pandas rather than the csv module?

The csv module is fine for reading rows and for small checks, as check.py shows. pandas adds typed columns, missing-value handling, grouping, joins, dates and plotting in a few lines. Using both is a good habit: pandas for the analysis, the standard library for an independent second calculation.

Which versions does this workbook use?

Python 3.12.10 on Windows 11 with pandas 3.0.6, matplotlib 3.11.2 and numpy 2.5.3, installed in a fresh virtual environment. These pins need Python 3.12 or later. Pin the same versions to reproduce every printed number: another Python version may invent different data, and the macOS and Linux commands were not run.

Should I delete rows with missing values?

Not by default. dropna() removes a row if any column is missing, so here it would throw away 86 of 300 tickets, mostly for a missing score that has nothing to do with timing. Drop only what your question requires, name the column, and log how many rows you lost and why.

When should I use the median instead of the mean?

When the data is skewed, such as times, prices or incomes, where a few very large values pull the mean up. The median is the middle value and resists that pull. Report the counts behind either one, and consider showing both.

Why did chat look much faster than it really was?

Chat handled a larger share of quick login problems, while email handled more slow delivery problems. Topic affected both which channel was used and how long a ticket took. Comparing within each topic removed most of the gap. Always ask what else differs between the groups you compare.

What does a 95% bootstrap interval tell me?

It shows how much your result might vary between samples of the same size. If it sits well away from zero, the gap is unlikely to be chance alone. If it crosses zero, the data cannot confirm the direction. It cannot fix a biased sample or a confounder.

Can I paste my company's data into an AI assistant?

Do not paste real data. Give the assistant the column names, types and a few invented rows instead. The ICO's guidance, under review in 2026, says pseudonymised data, with names replaced by reference numbers, is still personal data under the UK GDPR, so follow your organisation's rules before any real data goes near any tool.

How do I check AI-written analysis code?

Read it before you run it. Look for silent dropna() calls, inner joins, joins on the wrong key, invented column names and shares with no counts. Print row counts after each step. Then recompute one key number a different way, and make sure the two routes agree.

How do I make my analysis reproducible?

Pin package versions, seed anything random, keep the raw file untouched, and let scripts write a cleaned copy and a log. Save the generated files too: Python's documentation promises repeatable output across versions only for random() itself, not for methods such as choice or gauss.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

DataFrame
pandas' table: named columns of equal length, each with one data type.
dtype
The data type of a column, such as str (text), float64 (decimal numbers) or datetime64 (dates and times).
Exploratory data analysis
Looking at data with summaries and charts to find its structure and problems before assuming anything about it.
Missing value
A value that is absent, shown by pandas as NaN. Summaries skip it by default, so always check how many there are.
Median
The middle value once the values are sorted. Long tails pull it far less than they pull the mean.
groupby
Split rows into groups by a column, summarise each group, then combine the results.

6 of the workbook's 12 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

Examples in this workbook were run on: Python 3.12.10 on Windows 11, in a fresh virtual environment with pandas 3.0.6, matplotlib 3.11.2 and numpy 2.5.3 (installed by pip 25.0.1) (2026-09-26).

These workbooks use AI assistance. See how the workbooks are made.

  1. What's new in 3.0.0 (default string data type)pandas 3.0.6 documentation
  2. NumPy 2.5.0 release notes (drops Python 3.11; supports 3.12 to 3.14)NumPy
  3. pandas.read_csvpandas 3.0.6 documentation
  4. pandas.DataFrame.to_csv (mode 'x')pandas 3.0.6 documentation
  5. pandas.DataFrame.describepandas 3.0.6 documentation
  6. pandas.Series.value_countspandas 3.0.6 documentation
  7. pandas.DataFrame.duplicatedpandas 3.0.6 documentation
  8. pandas.DataFrame.drop_duplicatespandas 3.0.6 documentation
  9. pandas.DataFrame.dropnapandas 3.0.6 documentation
  10. Working with missing datapandas 3.0.6 documentation
  11. pandas.Series.medianpandas 3.0.6 documentation
  12. Group by: split-apply-combine (including named aggregation)pandas 3.0.6 documentation
  13. pandas.crosstabpandas 3.0.6 documentation
  14. pandas.pivot_tablepandas 3.0.6 documentation
  15. pandas.DataFrame.mergepandas 3.0.6 documentation
  16. pandas.Series.dt.to_periodpandas 3.0.6 documentation
  17. pandas.Series.corrpandas 3.0.6 documentation
  18. pandas.DataFrame.plotpandas 3.0.6 documentation
  19. matplotlib.figure.Figure.savefigMatplotlib 3.11.2 documentation
  20. matplotlib.axes.Axes.set_ylimMatplotlib 3.11.2 documentation
  21. matplotlib.pyplot.subplotsMatplotlib 3.11.2 documentation
  22. matplotlib.figure.Figure.tight_layoutMatplotlib 3.11.2 documentation
  23. venv: creation of virtual environmentsPython Software Foundation
  24. random: generate pseudo-random numbers (notes on reproducibility)Python Software Foundation
  25. statistics: mathematical statistics functionsPython Software Foundation
  26. csv: CSV file reading and writingPython Software Foundation
  27. Built-in functions: open() and mode 'x'Python Software Foundation
  28. Modules: the module search path (The Python Tutorial)Python Software Foundation
  29. 1.1.1 What is EDA?NIST/SEMATECH e-Handbook of Statistical Methods
  30. 1.3.5.1 Measures of locationNIST/SEMATECH e-Handbook of Statistical Methods
  31. Dataplot reference manual: CORRELATIONNational Institute of Standards and Technology
  32. 1.3.3.26 Scatter plot (causality is not proved by association)NIST/SEMATECH e-Handbook of Statistical Methods
  33. 1.3.3.4 Bootstrap plotNIST/SEMATECH e-Handbook of Statistical Methods
  34. 1.3.5.2 Confidence limits for the meanNIST/SEMATECH e-Handbook of Statistical Methods
  35. What is personal data?Information Commissioner's Office

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 26 September 2026.

NextKeep going

Where to go next