Introduction
People new to ML often spend a lot of time because of the messy installation environment: many versions of Python, notebooks that work today but fail tomorrow, packages installed haphazardly, project files with no structure. This article helps you build a stable learning and working environment so that from the next lesson you can only focus on learning ML instead of fixing installation errors.
Lesson objectives
- Install a clean Python environment for ML
- Know how to organize project folders to learn and reuse code
- Run first notebook with NumPy, Pandas and scikit-learn
1. Which Python version to choose?
Python 3.11 is recommended.
Reason:
- Most popular ML libraries support it well.
- Faster and more stable than 3.9/3.10.
- Less risk of incompatibility than newer versions.
Check version:
python --version
python3 --version
2. Use venv or Conda?
For newbies, there are two popular options:
venv
Advantages:
- Available in Python.
- Light, simple.
Disadvantages:
- Less convenient when working with packages with complex native dependencies.
Conda / Miniconda
Advantages:
- Good environment management for data/ML.
- Easier to install scientific package.
Disadvantages:
- Slightly heavier.
In this course, if you are completely new, you can use it venv. If you plan on deep learning long term, maybe switch to Conda.
3. Create the first project
For example with venv:
mkdir ml-course
cd ml-course
python -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install numpy pandas scikit-learn matplotlib jupyter
On Windows PowerShell:
.venv\Scripts\Activate.ps1
4. Suggested folder structure
ml-course/
├── notebooks/
├── data/
│ ├── raw/
│ └── processed/
├── src/
│ ├── features/
│ ├── models/
│ └── utils/
├── outputs/
│ ├── figures/
│ └── models/
├── requirements.txt
└── README.md
Meaning:
notebooks/: a place to experiment and learn.data/raw/: original data, not edited directly.data/processed/: cleaned data.src/: reusable code.outputs/: model, chart, artifact.
5. Tools should be pre-installed
Required
- Python
- VS Code
- Jupyter
numpy,pandas,scikit-learn
Should have
matplotlib,seabornipykernelblackor similar formatter
pip install seaborn ipykernel
6. Run the first notebook
jupyter notebook
Or use VS Code to open the file .ipynb.
Try running the following:
import numpy as np
import pandas as pd
from sklearn.datasets import load_iris
iris = load_iris(as_frame=True)
df = iris.frame
print(df.head())
print(df.shape)
print(df['target'].value_counts())
If the notebook works, your environment is sufficient for most of the basic lessons in the series.
7. The files should be there from the beginning
requirements.txt
numpy
pandas
scikit-learn
matplotlib
seaborn
jupyter
README.md
The minimum README should contain:
- project goal
- how to install the environment
- how to run notebook or script
8. Common installation errors
Error: installed package but notebook does not recognize it
The reason is usually that the notebook is running a different kernel from the environment you just installed.
How to handle:
python -m ipykernel install --user --name ml-course
Then select the correct kernel in VS Code/Jupyter.
Error: ModuleNotFoundError
Check:
- Has the environment been activated yet?
- Are you using Python correctly?
- Which environment does the package install into?
Error: notebook is too messy
Solution:
- notebook is only used for exploration
- Reusable code switched
src/ - don't put everything in a file 1000 lines long
9. Standard work from the beginning
Three small but extremely important principles:
- Each project uses a separate environment.
- Do not edit the original data directly.
- Record how to run the project in README.
If you do these three things from the beginning, you will save a lot of headaches as the number of projects increases.
Practice exercises
- Create a separate environment for this series.
- Create a folder structure according to the sample above.
- Run the first notebook and save the snapshot or output.
Common mistakes
- Install the package into the global environment.
- Use one environment for every project.
- I don't know which kernel the notebook is using.
Completion criteria
- Create your own ML environment
- Run the first notebook
- Have neat, reusable project folders