Add numpy, pandas, scikit-learn and matplotlib to your course folder, run notebooks inside VS Code, and prove it all works. About 30 minutes, no cost, no new accounts.
Unit 2 teaches how machines learn from examples. To try that, learners need four standard Python libraries for working with data and one way to run code in small pieces and see the result straight away.
The libraries are numpy (fast arithmetic on long lists of numbers), pandas (tables, like a spreadsheet you drive with code), scikit-learn (ready-made machine learning methods) and matplotlib (charts). The "small pieces" tool is a notebook: a file that mixes code, results and charts on one page. VS Code, which learners installed in Unit 1, opens notebooks with one free extension.
Setup takes about 30 minutes. It costs nothing and needs no new accounts. All data in Unit 2 is made up to look like SAP data, so nothing leaves the learner's laptop.
These four libraries are the common language of data work. Data scientists at your company, at SAP partners and at SAP itself use them to explore data before any model reaches production. When a team says "we looked at the data in a notebook", this is what they mean.
For a leader, two things follow:
Notebooks are for exploring, not for running the business. A notebook that predicts late deliveries is a good way to test an idea. It is not a system anyone should rely on at month-end. Moving from notebook to production is covered in later units.
Notebooks keep their results inside the file. A table printed in a notebook is saved in it. If someone opens real customer data in a notebook and shares the file, the data goes with it. The course uses made-up data for exactly this reason.
SAP's data scientists use the same open tools, and SAP also offers its own Python library, hana-ml, for machine learning inside SAP HANA. As of September 2026 it is published on PyPI and works with pandas tables, so skills from this unit carry over. It needs access to an SAP HANA Cloud system, which this unit doesn't. The deep layer explains the difference.
Time: about 30 minutes, most of it waiting for downloads. The libraries are a few hundred megabytes in total.
Money: nothing.
One version rule: as of September 2026, the latest releases of these libraries need Python 3.11 or newer, and numpy's newest release needs 3.12. A learner who installed Python in Unit 1 already meets this.
Can learners install Python packages from PyPI (pypi.org and files.pythonhosted.org), or do we have an internal package mirror to use instead?
Can learners install VS Code extensions from the Marketplace, or is the list controlled?
What is our rule for opening company data in notebooks on laptops, and for sharing notebook files?
If a learner later wants to try SAP's own Python library for machine learning inside SAP HANA, who can give access to a HANA Cloud system for learning?
"We need Anaconda or a data science platform." Not for this course. Plain Python with a virtual environment and pip is enough, and it is the same setup learners already have.
"Notebooks are a separate program." VS Code opens them directly with the Jupyter extension. No browser tool is needed.
"Machine learning needs a GPU." Not for Unit 2. These models train in seconds on any laptop.
"A good notebook result means the model is ready." It means the idea is worth testing further. Unit 8 covers how to evaluate properly.
Pick one answer for each question. The explanation appears after you choose.
1What does Unit 2 add to a learner's laptop, and what does it cost?
Answer: C. Four free Python libraries for data work (numpy, pandas, scikit-learn and matplotlib), plus ipykernel and VS Code's Jupyter extension to run notebooks. About 30 minutes, no cost and no new accounts.
2What is a notebook, and what is it good for?
Answer: B. A page of code cells, each followed by its output, where Python stays running between cells. It is good for exploring data and testing an idea, not for running a business process that people rely on.
3Why is sharing a notebook a data-protection risk?
Answer: A. Notebooks save their results inside the file. If someone opens real customer data and shares or commits the notebook, the data goes with it. The course uses made-up data for exactly this reason.
4Does Unit 2 need a GPU, Anaconda or a data science platform?
Answer: D. No. Plain Python with the course's virtual environment and pip is enough, and the Unit 2 models train in seconds on any laptop.
5A team shows a good result from a notebook. Is the model ready for production?
Answer: C. No. A good notebook result means the idea is worth testing further. Proper evaluation comes in Unit 8, and moving from notebook to production in later units.
6What should you ask IT before learners start Unit 2?
Answer: B. Whether learners can install packages from PyPI or must use an internal mirror, whether VS Code extensions are allowed, and what the rule is for opening company data in notebooks and sharing notebook files.
7How do these open tools connect to SAP?
Answer: A. SAP's data scientists use the same tools, and SAP publishes hana-ml, a Python library that works with pandas-style tables and runs machine learning inside SAP HANA. It isn't needed for Unit 2, but the skills carry over.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Deep layer · 30 min read
#Mental model: same folder, same Python, a new way to run it
Nothing about your course folder changes. You still have one .venv, one requirements.txt and one Git repository. This unit adds five lines to requirements.txt and one new way to run code: a notebook, where Python stays running between cells.
flowchart LR
R[requirements.txt<br/>+5 lines] -->|pip install| V[.venv<br/>course Python]
V -->|runs| S[hello_data.py<br/>in the terminal]
V -->|kernel via ipykernel| N[first_notebook.ipynb<br/>in VS Code]
S -->|writes| D[sample_deliveries.csv]
D -->|read by| N
A script runs top to bottom and exits. A notebook's kernel stays alive, so a variable you made in cell 1 is still there in cell 5. That is why notebooks are good for exploring data, and also why they confuse people: the result depends on which cells you ran and in what order.
When you open a notebook in VS Code, the Jupyter extension starts a kernel: a Python process that waits for code. You choose which Python it uses with the Select Kernel button. If you pick your course .venv, the kernel sees every library you installed there.
VS Code's documentation is clear on what the environment needs: not the full Jupyter package, only ipykernel, the small library that lets a Python environment act as a kernel. That is why ipykernel is in this unit's list.
The four data libraries build on each other:
Library
Builds on
What you do with it in Unit 2
numpy
Nothing (compiled code)
Make random sample data, do arithmetic on whole columns
pandas
numpy
Hold deliveries in a table, group by plant, read and write CSV
scikit-learn
numpy and SciPy
Split data, train a model, score it
matplotlib
numpy
Draw charts in scripts and notebooks
pip installs the extra libraries each one depends on, such as SciPy for scikit-learn, without you asking.
#Build it yourself: add the data tools and prove they work
You will add the Unit 2 libraries to your course folder, run a short script that makes SAP-shaped sample deliveries and trains a first model, then open the same data in a notebook. Finally, a check script confirms everything.
Before you start: complete Set up your computer for this course. It installs Python, VS Code and Git, and creates your orchestrate-course folder with its .venv virtual environment and requirements.txt. This walkthrough doesn't repeat those steps.
Open VS Code, choose File > Open Folder, and open orchestrate-course.
Open a terminal: Terminal > New Terminal.
Run:
python --version
(On macOS or Linux, before .venv is active, use python3.)
You need 3.11 or newer. Python 3.12 or later is best: as of September 2026 the newest numpy release needs 3.12, and on 3.11 pip quietly installs a slightly older numpy, which is fine for this course.
If you see 3.10 or older, install a newer Python (Unit 1, Step 1), then delete the .venv folder in VS Code's file list and recreate it:
In VS Code, open requirements.txt and add these five lines below the existing ones:
numpy
pandas
scikit-learn
matplotlib
ipykernel
Save the file (Ctrl+S, or Cmd+S on Mac).
With (.venv) showing, run:
pip install -r requirements.txt
Wait for a line starting Successfully installed. This can take a few minutes; the libraries are large. Libraries you already had show Requirement already satisfied, which is fine.
This script proves all four libraries work together. It makes 500 made-up outbound deliveries with fields shaped like SAP data (a delivery number, a plant, a shipping type, items, weight, lead time) and a Late flag. Then it prints a summary, draws one chart and trains a first model. Unit 2's later topics explain what the model is doing; here it only needs to run.
In VS Code's file list, right-click unit02, choose New File and name it hello_data.py.
Paste the code below and save.
In the terminal, still inside unit02, run:
python hello_data.py
"""Unit 2 smoke test: make SAP-shaped sample deliveries, look at them, fit a tiny model.
How to run (from the unit02 folder, with the course .venv turned on):
python hello_data.py # 500 made-up deliveries
python hello_data.py --rows 2000
All data is made up. It writes sample_deliveries.csv and late_by_plant.png next to this file.
"""
import argparse
from pathlib import Path
import matplotlib
matplotlib.use("Agg") # draw into a file, no window needed
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
HERE = Path(__file__).parent
PLANTS = ["1010", "1020", "1710"] # made-up plant codes
CARRIERS = ["ROAD", "RAIL", "EXPRESS"] # made-up shipping types
def make_deliveries(rows: int, seed: int = 42) -> pd.DataFrame:
"""Build a table that looks like outbound deliveries, with a 'late' flag."""
rng = np.random.default_rng(seed) # same seed, same data, every run
df = pd.DataFrame({
"Delivery": [str(80000000 + i) for i in range(rows)],
"Plant": rng.choice(PLANTS, size=rows),
"ShippingType": rng.choice(CARRIERS, size=rows, p=[0.6, 0.3, 0.1]),
"Items": rng.integers(1, 12, size=rows),
"WeightKg": rng.gamma(2.0, 150.0, size=rows).round(1),
"LeadTimeDays": rng.integers(1, 15, size=rows),
})
# A made-up rule plus noise decides which deliveries were late.
risk = (
0.25 * df["Items"]
+ 0.004 * df["WeightKg"]
- 0.35 * df["LeadTimeDays"]
+ np.where(df["Plant"] == "1710", 1.0, 0.0)
+ np.where(df["ShippingType"] == "EXPRESS", -1.5, 0.0)
+ rng.normal(0, 1.0, size=rows)
)
df["Late"] = (risk > 0).astype(int)
return df
def main() -> None:
parser = argparse.ArgumentParser(description="Unit 2 smoke test on made-up deliveries.")
parser.add_argument("--rows", type=int, default=500, help="how many deliveries to make")
args = parser.parse_args()
# 1. numpy + pandas: make and inspect the data
df = make_deliveries(args.rows)
df.to_csv(HERE / "sample_deliveries.csv", index=False)
print(f"Made {len(df)} deliveries. First three:")
print(df.head(3).to_string(index=False))
print("\nShare of late deliveries by plant:")
by_plant = df.groupby("Plant")["Late"].mean().round(2)
print(by_plant.to_string())
# 2. matplotlib: one chart, saved as a file
ax = by_plant.plot(kind="bar", color="#4a6fa5", rot=0)
ax.set_title("Share of late deliveries by plant (made-up data)")
ax.set_ylabel("share late")
plt.tight_layout()
plt.savefig(HERE / "late_by_plant.png")
plt.close()
# 3. scikit-learn: a first model, just to prove the library works
features = pd.get_dummies(df[["Items", "WeightKg", "LeadTimeDays", "Plant", "ShippingType"]])
x_train, x_test, y_train, y_test = train_test_split(
features, df["Late"], test_size=0.25, random_state=0
)
model = LogisticRegression(max_iter=1000).fit(x_train, y_train)
accuracy = model.score(x_test, y_test)
print(f"\nModel accuracy on deliveries it has not seen: {accuracy:.2f}")
print("Wrote sample_deliveries.csv and late_by_plant.png")
if __name__ == "__main__":
main()
What success looks like:
Made 500 deliveries. First three:
Delivery Plant ShippingType Items WeightKg LeadTimeDays Late
80000000 1010 RAIL 6 132.1 3 1
80000001 1710 RAIL 3 256.8 1 1
80000002 1020 ROAD 3 230.0 9 0
Share of late deliveries by plant:
Plant
1010 0.46
1020 0.53
1710 0.63
Model accuracy on deliveries it has not seen: 0.85
Wrote sample_deliveries.csv and late_by_plant.png
Your numbers should match, because the seed fixes the random data. If the last decimal of the accuracy differs, that is fine: small differences between library versions can do that. Two new files appear in unit02. Click late_by_plant.png to see the chart.
What each part of the script does:
Part
What it does
matplotlib.use("Agg")
Draws charts into files, so the script needs no window
make_deliveries
Uses numpy's random generator with a fixed seed to build a pandas table of made-up deliveries, then marks some late with a made-up rule plus noise
df.to_csv(...)
Saves the table as sample_deliveries.csv for the notebook in Step 7
groupby("Plant")["Late"].mean()
Share of late deliveries per plant, the kind of question you ask first
by_plant.plot(...), savefig
Turns that summary into a bar chart file
pd.get_dummies
Turns text columns such as Plant into 0/1 columns a model can use
train_test_split, LogisticRegression, score
Holds back a quarter of the rows, trains on the rest, and reports how often the model is right on the held-back rows
--rows
Optional: make more or fewer deliveries, for example --rows 2000
Click the Variables icon in the notebook toolbar. A panel opens at the bottom listing df. Use Show variable in data viewer next to it to scroll through all 500 rows.
In VS Code, create a new file in the course folder (not in unit02), paste the script below and save it as check_unit02.py.
Run it:
python check_unit02.py
"""Check that your computer is ready for Unit 2 (data science tools).
Run it from your course folder: python check_unit02.py
It only reads your setup; it installs nothing and sends nothing anywhere.
"""
import importlib.metadata
import importlib.util
import os
import shutil
import subprocess
import sys
problems = 0
def report(ok: bool, label: str, fix: str = "", optional: bool = False) -> None:
"""Print one line: OK, MISSING (must fix) or LATER (optional for now)."""
global problems
if ok:
print(f" OK {label}")
elif optional:
print(f" LATER {label} -> {fix}")
else:
problems += 1
print(f" MISSING {label} -> {fix}")
def version_of(package: str) -> str:
try:
return importlib.metadata.version(package)
except importlib.metadata.PackageNotFoundError:
return ""
print("\n1. Python")
v = sys.version_info
report(v >= (3, 11), f"Python {v.major}.{v.minor}.{v.micro}",
"Unit 2 libraries need Python 3.11 or newer; install a newer Python and recreate .venv (Step 1)")
report(sys.prefix != sys.base_prefix, "virtual environment is active", "activate .venv (Step 2)")
print("\n2. Libraries")
LIBRARIES = [ # (import name, package name on PyPI)
("numpy", "numpy"),
("pandas", "pandas"),
("sklearn", "scikit-learn"),
("matplotlib", "matplotlib"),
("ipykernel", "ipykernel"),
]
for module, package in LIBRARIES:
found = importlib.util.find_spec(module) is not None
label = f"{package} {version_of(package)}".strip()
report(found, label, "pip install -r requirements.txt (Step 3)")
print("\n3. Libraries work together")
# Import all of them in a fresh Python and do one tiny calculation each.
# This catches broken or half-finished installs that find_spec can't see.
test = (
"import matplotlib; matplotlib.use('Agg'); import matplotlib.pyplot as plt;"
"import numpy as np, pandas as pd;"
"from sklearn.linear_model import LinearRegression;"
"df = pd.DataFrame({'x': np.arange(5), 'y': np.arange(5) * 2.0});"
"m = LinearRegression().fit(df[['x']], df['y']);"
"assert abs(m.coef_[0] - 2.0) < 1e-6;"
"plt.plot(df['x'], df['y']); plt.close()"
)
try:
result = subprocess.run([sys.executable, "-c", test], capture_output=True, text=True, timeout=180)
ok = result.returncode == 0
detail = (result.stderr.strip().splitlines() or ["unknown error"])[-1] if not ok else ""
except subprocess.TimeoutExpired:
ok, detail = False, "timed out"
report(ok, "import, fit a line, draw a chart",
f"{detail}; reinstall with pip install --force-reinstall -r requirements.txt (Step 3)")
print("\n4. Editor")
code = shutil.which("code")
if code:
try:
exts = subprocess.run([code, "--list-extensions"], capture_output=True, text=True, timeout=60).stdout.lower()
except Exception:
exts = ""
report("ms-toolsai.jupyter" in exts, "VS Code Jupyter extension",
"install Jupyter by Microsoft in VS Code (Step 4)", optional=True)
else:
report(False, "VS Code 'code' command", "can't check from here; confirm Step 4 in VS Code by eye", optional=True)
print("\n5. Course folder")
report(os.path.isdir("unit02"), "unit02 folder", "create it (Step 5)", optional=True)
report(os.path.exists(os.path.join("unit02", "hello_data.py")), "unit02/hello_data.py",
"create it (Step 6)", optional=True)
print()
if problems:
print(f"{problems} item(s) to fix. Fix them in order, then run this again.")
sys.exit(1)
print("All set. Your computer is ready for Unit 2.")
What success looks like:
1. Python
OK Python 3.11.15
OK virtual environment is active
2. Libraries
OK numpy 2.4.6
OK pandas 3.0.6
OK scikit-learn 1.9.1
OK matplotlib 3.11.2
OK ipykernel 7.4.0
3. Libraries work together
OK import, fit a line, draw a chart
4. Editor
LATER VS Code 'code' command -> can't check from here; confirm Step 4 in VS Code by eye
5. Course folder
OK unit02 folder
OK unit02/hello_data.py
All set. Your computer is ready for Unit 2.
This run used Python 3.11, so pip chose numpy 2.4; on Python 3.12 or newer you'll see a newer numpy. Your version numbers may differ; that is fine as long as each line says OK. The LATER line under Editor is also fine: it means the script couldn't ask VS Code directly because the code command isn't on your PATH, which is common on macOS. Step 7 working already proves the extension is there. If code is on your PATH, you see OK VS Code Jupyter extension instead.
What each part of the check does:
Part
What it checks
Python
Version 3.11 or newer, and that .venv is active
Libraries
That each library can be found, and which version is installed
Libraries work together
Starts a fresh Python, imports all four, fits a straight line and draws a chart. This catches broken installs a simple "is it there?" test misses
Editor
Asks VS Code, through its code command, whether the Jupyter extension is installed
Course folder
That unit02 and hello_data.py exist
Like check_setup.py, it uses only Python's built-in modules, so it runs even before Step 3. It changes nothing and sends nothing anywhere.
pandas DataFrames are also the meeting point with SAP's own Python tools. SAP publishes hana-ml, the Python machine learning client for SAP HANA, on PyPI. As described there in September 2026, it gives you a HANA DataFrame that queries data without copying it to your laptop, plus Python access to two in-database libraries: the Predictive Analysis Library (PAL) for classic algorithms and the Automated Predictive Library (APL) for automated modelling. It can turn a HANA DataFrame into a pandas DataFrame and back.
The difference matters. In this unit, data comes to the model on your laptop. With hana-ml, the model runs inside the database, next to the data. That is often the right choice for large SAP tables and for data that must not leave the system.
hana-ml is not needed for Unit 2. It needs an SAP HANA Cloud instance with the script server, PAL and APL enabled, according to SAP Learning's setup lesson, and it is licensed under the SAP Developer License Agreement. SAP Learning has a free journey, "Developing AI workflows with the Python Machine Learning Client for SAP HANA", for when you have access. The concepts you learn in Unit 2 with scikit-learn (training, test splits, metrics) carry over directly.
Running cells out of order. A notebook remembers everything. If a result looks wrong, restart the kernel from the notebook toolbar and run all cells from the top.
The wrong kernel. The kernel name shows at the top right. If it isn't your .venv, imports fail or use other library versions.
Codes read as numbers. Plants, company codes and document numbers are text. Set dtype when reading CSV files, or leading zeros disappear.
Real data in notebooks. Outputs are saved inside the .ipynb file. Opening real customer data and committing the notebook puts that data in your Git history. Use sample data, and clear all outputs (from the notebook toolbar) before sharing.
Installing "just one more library" outside .venv. Always check for (.venv) before pip install, and add the library to requirements.txt so the setup can be rebuilt.
The next topics use sample_deliveries.csv. Make a version that fits the business you know.
Copy hello_data.py to my_deliveries.py in unit02.
In make_deliveries, add a column Customer with three made-up customer numbers, chosen at random like Plant is.
Add one line to the risk calculation so deliveries to one of those customers are more often late.
Change the output file names to my_deliveries.csv and my_late_by_plant.png.
Run python my_deliveries.py. Check that the new CSV has a Customer column.
In first_notebook.ipynb, add a cell that reads my_deliveries.csv and shows the share of late deliveries by Customer. Check that your chosen customer has the highest share.
Save your work in Git:
git add unit02/my_deliveries.py unit02/first_notebook.ipynb
git commit -m "Add my own sample deliveries"
Done when:check_unit02.py prints All set, my_deliveries.csv has a Customer column, the notebook shows your chosen customer as the most often late, and git log shows the commit. Keep the script: Unit 2's late-delivery prediction topic trains on data like this.
Pick one answer for each question. The explanation appears after you choose.
1Which package name do you install, and which name do you import, for scikit-learn?
Answer: D. You install scikit-learn and write import sklearn. Install names and import names don't always match.
2What does a notebook's kernel remember between cells, and how can that give you a wrong result?
Answer: C. Every variable and import from cells already run. Results then depend on which cells ran and in what order. If something looks wrong, restart the kernel and run all cells from the top.
3Why does the notebook need ipykernel in .venv, but not the full Jupyter package?
Answer: B. VS Code's Jupyter extension provides the notebook interface. The environment only needs ipykernel so it can act as the kernel, the Python process that runs the cells with your course libraries.
4Why read Plant and Delivery as text rather than numbers?
Answer: A. They are SAP codes, not quantities. Read as numbers, leading zeros disappear and codes can be summed or sorted wrongly. Set dtype to str when reading the CSV.
5Your check script says OK for every library, but a notebook cell fails with No module named 'pandas'. What is the likely cause?
Answer: D. The notebook uses a different Python. Click the kernel name at the top right and choose the one inside your course .venv.
6How do the four data libraries depend on each other?
Answer: C. numpy is the base for arithmetic on arrays. pandas builds tables on numpy, scikit-learn builds models on numpy and SciPy, and matplotlib draws charts from numpy data. pip installs the extra dependencies for you.
7Why do the numbers from hello_data.py match the page?
Answer: B. It makes 500 made-up, SAP-shaped deliveries, summarizes late deliveries by plant, draws a chart, and trains and scores a first model. A fixed random seed makes the data identical on every run; small version differences may change the last decimal.
8When would you prefer SAP's hana-ml, which runs the model inside SAP HANA, over pandas and scikit-learn on a laptop?
Answer: A. For large SAP tables and for data that must not leave the system, because the model runs in the database next to the data. It needs an SAP HANA Cloud instance with the right features enabled, and isn't needed for Unit 2.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Sources
Jupyter Notebooks in VS Code (VS Code docs)— Jupyter extension, Create New Jupyter Notebook command, kernel picker, run-cell shortcuts, Variables view, Data Viewer, saving plots