Chuyển đến nội dung chính

Lesson 2: Set up a production standard ML learning environment

Install Python, Jupyter, VS Code, NumPy/Pandas/scikit-learn; create project templates, manage dependencies and notebook workflows.

🧠 AI & ML — Lesson 1 Lesson 2: Set up a standard ML learning environment production. production

Machine Learning: From Basics to Advanced

Part 0: Getting started for newbies (Week 0)

xdev.asia

Introduction

People new to ML often spend a lot of time because of the messy installation environment: many versions of Python, notebooks that work today but fail tomorrow, packages installed haphazardly, project files with no structure. This article helps you build a stable learning and working environment so that from the next lesson you can only focus on learning ML instead of fixing installation errors.

Lesson objectives

  • Install a clean Python environment for ML
  • Know how to organize project folders to learn and reuse code
  • Run first notebook with NumPy, Pandas and scikit-learn

1. Which Python version to choose?

Python 3.11 is recommended.

Reason:

  • Most popular ML libraries support it well.
  • Faster and more stable than 3.9/3.10.
  • Less risk of incompatibility than newer versions.

Check version:

python --version
python3 --version

2. Use venv or Conda?

For newbies, there are two popular options:

venv

Advantages:

  • Available in Python.
  • Light, simple.

Disadvantages:

  • Less convenient when working with packages with complex native dependencies.

Conda / Miniconda

Advantages:

  • Good environment management for data/ML.
  • Easier to install scientific package.

Disadvantages:

  • Slightly heavier.

In this course, if you are completely new, you can use it venv. If you plan on deep learning long term, maybe switch to Conda.

3. Create the first project

For example with venv:

mkdir ml-course
cd ml-course

python -m venv .venv
source .venv/bin/activate

pip install --upgrade pip
pip install numpy pandas scikit-learn matplotlib jupyter

On Windows PowerShell:

.venv\Scripts\Activate.ps1

4. Suggested folder structure

ml-course/
├── notebooks/
├── data/
│   ├── raw/
│   └── processed/
├── src/
│   ├── features/
│   ├── models/
│   └── utils/
├── outputs/
│   ├── figures/
│   └── models/
├── requirements.txt
└── README.md

Meaning:

  • notebooks/: a place to experiment and learn.
  • data/raw/: original data, not edited directly.
  • data/processed/: cleaned data.
  • src/: reusable code.
  • outputs/: model, chart, artifact.

5. Tools should be pre-installed

Required

  • Python
  • VS Code
  • Jupyter
  • numpy, pandas, scikit-learn

Should have

  • matplotlib, seaborn
  • ipykernel
  • black or similar formatter
pip install seaborn ipykernel

6. Run the first notebook

jupyter notebook

Or use VS Code to open the file .ipynb.

Try running the following:

import numpy as np
import pandas as pd
from sklearn.datasets import load_iris

iris = load_iris(as_frame=True)
df = iris.frame

print(df.head())
print(df.shape)
print(df['target'].value_counts())

If the notebook works, your environment is sufficient for most of the basic lessons in the series.

7. The files should be there from the beginning

requirements.txt

numpy
pandas
scikit-learn
matplotlib
seaborn
jupyter

README.md

The minimum README should contain:

  • project goal
  • how to install the environment
  • how to run notebook or script

8. Common installation errors

Error: installed package but notebook does not recognize it

The reason is usually that the notebook is running a different kernel from the environment you just installed.

How to handle:

python -m ipykernel install --user --name ml-course

Then select the correct kernel in VS Code/Jupyter.

Error: ModuleNotFoundError

Check:

  • Has the environment been activated yet?
  • Are you using Python correctly?
  • Which environment does the package install into?

Error: notebook is too messy

Solution:

  • notebook is only used for exploration
  • Reusable code switched src/
  • don't put everything in a file 1000 lines long

9. Standard work from the beginning

Three small but extremely important principles:

  1. Each project uses a separate environment.
  2. Do not edit the original data directly.
  3. Record how to run the project in README.

If you do these three things from the beginning, you will save a lot of headaches as the number of projects increases.

Practice exercises

  1. Create a separate environment for this series.
  2. Create a folder structure according to the sample above.
  3. Run the first notebook and save the snapshot or output.

Common mistakes

  • Install the package into the global environment.
  • Use one environment for every project.
  • I don't know which kernel the notebook is using.

Completion criteria

  • Create your own ML environment
  • Run the first notebook
  • Have neat, reusable project folders