Orchestrate

Set up for Unit 2: data science tools

Add numpy, pandas, scikit-learn and matplotlib to your course folder, run notebooks inside VS Code, and prove it all works. About 30 minutes, no cost, no new accounts.

Updated Sep 29, 2026Foundational 5 minDeep 30 min
Foundational layer · 5 min read

The 60-second version

Unit 2 teaches how machines learn from examples. To try that, learners need four standard Python libraries for working with data and one way to run code in small pieces and see the result straight away.

The libraries are numpy (fast arithmetic on long lists of numbers), pandas (tables, like a spreadsheet you drive with code), scikit-learn (ready-made machine learning methods) and matplotlib (charts). The "small pieces" tool is a notebook: a file that mixes code, results and charts on one page. VS Code, which learners installed in Unit 1, opens notebooks with one free extension.

Setup takes about 30 minutes. It costs nothing and needs no new accounts. All data in Unit 2 is made up to look like SAP data, so nothing leaves the learner's laptop.

Why it matters to the business

These four libraries are the common language of data work. Data scientists at your company, at SAP partners and at SAP itself use them to explore data before any model reaches production. When a team says "we looked at the data in a notebook", this is what they mean.

For a leader, two things follow:

  • Notebooks are for exploring, not for running the business. A notebook that predicts late deliveries is a good way to test an idea. It is not a system anyone should rely on at month-end. Moving from notebook to production is covered in later units.
  • Notebooks keep their results inside the file. A table printed in a notebook is saved in it. If someone opens real customer data in a notebook and shares the file, the data goes with it. The course uses made-up data for exactly this reason.

How SAP does it

SAP's data scientists use the same open tools, and SAP also offers its own Python library, hana-ml, for machine learning inside SAP HANA. As of September 2026 it is published on PyPI and works with pandas tables, so skills from this unit carry over. It needs access to an SAP HANA Cloud system, which this unit doesn't. The deep layer explains the difference.

What this unit adds

Item What it is Cost Used in
numpy Fast maths on arrays of numbers Free Every Unit 2 topic, and most later ones
pandas Tables of data: load, filter, group, join Free Machine learning in plain terms, predicting late deliveries
scikit-learn Ready-made models and metrics Free Classification and the metrics that matter, predicting late deliveries
matplotlib Charts Free Gradient descent by hand, classification metrics
ipykernel Lets VS Code run notebook code with your course Python Free All notebooks from now on
Jupyter extension for VS Code Opens and runs notebooks in the editor Free All notebooks from now on

No SAP system, cloud account or model key is needed for this unit.

Time and money

  • Time: about 30 minutes, most of it waiting for downloads. The libraries are a few hundred megabytes in total.
  • Money: nothing.
  • One version rule: as of September 2026, the latest releases of these libraries need Python 3.11 or newer, and numpy's newest release needs 3.12. A learner who installed Python in Unit 1 already meets this.

Questions to ask IT

  • Can learners install Python packages from PyPI (pypi.org and files.pythonhosted.org), or do we have an internal package mirror to use instead?
  • Can learners install VS Code extensions from the Marketplace, or is the list controlled?
  • What is our rule for opening company data in notebooks on laptops, and for sharing notebook files?
  • If a learner later wants to try SAP's own Python library for machine learning inside SAP HANA, who can give access to a HANA Cloud system for learning?

Common misconceptions

  • "We need Anaconda or a data science platform." Not for this course. Plain Python with a virtual environment and pip is enough, and it is the same setup learners already have.
  • "Notebooks are a separate program." VS Code opens them directly with the Jupyter extension. No browser tool is needed.
  • "Machine learning needs a GPU." Not for Unit 2. These models train in seconds on any laptop.
  • "A good notebook result means the model is ready." It means the idea is worth testing further. Unit 8 covers how to evaluate properly.

Key terms

  • Library (package): ready-made code you install with pip.
  • DataFrame: a pandas table with named columns, the main thing you work with in Unit 2.
  • Notebook (.ipynb file): a page of code cells, each followed by its output.
  • Cell: one block of code in a notebook, run on its own.
  • Kernel: the running Python behind a notebook. It remembers what earlier cells did.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1What does Unit 2 add to a learner's laptop, and what does it cost?

    Answer: C. Four free Python libraries for data work (numpy, pandas, scikit-learn and matplotlib), plus ipykernel and VS Code's Jupyter extension to run notebooks. About 30 minutes, no cost and no new accounts.
  2. 2What is a notebook, and what is it good for?

    Answer: B. A page of code cells, each followed by its output, where Python stays running between cells. It is good for exploring data and testing an idea, not for running a business process that people rely on.
  3. 3Why is sharing a notebook a data-protection risk?

    Answer: A. Notebooks save their results inside the file. If someone opens real customer data and shares or commits the notebook, the data goes with it. The course uses made-up data for exactly this reason.
  4. 4Does Unit 2 need a GPU, Anaconda or a data science platform?

    Answer: D. No. Plain Python with the course's virtual environment and pip is enough, and the Unit 2 models train in seconds on any laptop.
  5. 5A team shows a good result from a notebook. Is the model ready for production?

    Answer: C. No. A good notebook result means the idea is worth testing further. Proper evaluation comes in Unit 8, and moving from notebook to production in later units.
  6. 6What should you ask IT before learners start Unit 2?

    Answer: B. Whether learners can install packages from PyPI or must use an internal mirror, whether VS Code extensions are allowed, and what the rule is for opening company data in notebooks and sharing notebook files.
  7. 7How do these open tools connect to SAP?

    Answer: A. SAP's data scientists use the same tools, and SAP publishes hana-ml, a Python library that works with pandas-style tables and runs machine learning inside SAP HANA. It isn't needed for Unit 2, but the skills carry over.
Deep layer · 30 min read

Mental model: same folder, same Python, a new way to run it

Nothing about your course folder changes. You still have one .venv, one requirements.txt and one Git repository. This unit adds five lines to requirements.txt and one new way to run code: a notebook, where Python stays running between cells.

flowchart LR
  R[requirements.txt<br/>+5 lines] -->|pip install| V[.venv<br/>course Python]
  V -->|runs| S[hello_data.py<br/>in the terminal]
  V -->|kernel via ipykernel| N[first_notebook.ipynb<br/>in VS Code]
  S -->|writes| D[sample_deliveries.csv]
  D -->|read by| N

A script runs top to bottom and exits. A notebook's kernel stays alive, so a variable you made in cell 1 is still there in cell 5. That is why notebooks are good for exploring data, and also why they confuse people: the result depends on which cells you ran and in what order.

How it works

When you open a notebook in VS Code, the Jupyter extension starts a kernel: a Python process that waits for code. You choose which Python it uses with the Select Kernel button. If you pick your course .venv, the kernel sees every library you installed there.

VS Code's documentation is clear on what the environment needs: not the full Jupyter package, only ipykernel, the small library that lets a Python environment act as a kernel. That is why ipykernel is in this unit's list.

The four data libraries build on each other:

Library Builds on What you do with it in Unit 2
numpy Nothing (compiled code) Make random sample data, do arithmetic on whole columns
pandas numpy Hold deliveries in a table, group by plant, read and write CSV
scikit-learn numpy and SciPy Split data, train a model, score it
matplotlib numpy Draw charts in scripts and notebooks

pip installs the extra libraries each one depends on, such as SciPy for scikit-learn, without you asking.

Build it yourself: add the data tools and prove they work

You will add the Unit 2 libraries to your course folder, run a short script that makes SAP-shaped sample deliveries and trains a first model, then open the same data in a notebook. Finally, a check script confirms everything.

Before you start: complete Set up your computer for this course. It installs Python, VS Code and Git, and creates your orchestrate-course folder with its .venv virtual environment and requirements.txt. This walkthrough doesn't repeat those steps.

What you need

  • Your course folder from Unit 1.
  • About 30 minutes and an internet connection.
  • No accounts and no cost.

Step 1: Check your Python version

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. Run:

    python --version

    (On macOS or Linux, before .venv is active, use python3.)

  4. You need 3.11 or newer. Python 3.12 or later is best: as of September 2026 the newest numpy release needs 3.12, and on 3.11 pip quietly installs a slightly older numpy, which is fine for this course.

If you see 3.10 or older, install a newer Python (Unit 1, Step 1), then delete the .venv folder in VS Code's file list and recreate it:

  • Windows (PowerShell):

    python -m venv .venv
    .venv\Scripts\Activate.ps1
    pip install -r requirements.txt
  • macOS / Linux:

    python3 -m venv .venv
    source .venv/bin/activate
    pip install -r requirements.txt

Step 2: Turn on the virtual environment

If your prompt doesn't start with (.venv), turn it on:

  • Windows (PowerShell):

    .venv\Scripts\Activate.ps1
  • macOS / Linux:

    source .venv/bin/activate

Step 3: Add the Unit 2 libraries

  1. In VS Code, open requirements.txt and add these five lines below the existing ones:

    numpy
    pandas
    scikit-learn
    matplotlib
    ipykernel
  2. Save the file (Ctrl+S, or Cmd+S on Mac).

  3. With (.venv) showing, run:

    pip install -r requirements.txt
  4. Wait for a line starting Successfully installed. This can take a few minutes; the libraries are large. Libraries you already had show Requirement already satisfied, which is fine.

Step 4: Add the Jupyter extension to VS Code

  1. Click the Extensions icon on the left (four small squares).
  2. Search for Jupyter and install the one published by Microsoft.
  3. If VS Code asks to reload, click the button it shows.

Step 5: Make the Unit 2 folder

In the terminal, from your course folder:

  • Windows (PowerShell):

    New-Item -ItemType Directory -Force unit02
    cd unit02
  • macOS / Linux:

    mkdir -p unit02
    cd unit02

Step 6: Run the smoke-test script

This script proves all four libraries work together. It makes 500 made-up outbound deliveries with fields shaped like SAP data (a delivery number, a plant, a shipping type, items, weight, lead time) and a Late flag. Then it prints a summary, draws one chart and trains a first model. Unit 2's later topics explain what the model is doing; here it only needs to run.

  1. In VS Code's file list, right-click unit02, choose New File and name it hello_data.py.

  2. Paste the code below and save.

  3. In the terminal, still inside unit02, run:

    python hello_data.py
"""Unit 2 smoke test: make SAP-shaped sample deliveries, look at them, fit a tiny model.

How to run (from the unit02 folder, with the course .venv turned on):
    python hello_data.py            # 500 made-up deliveries
    python hello_data.py --rows 2000

All data is made up. It writes sample_deliveries.csv and late_by_plant.png next to this file.
"""
import argparse
from pathlib import Path

import matplotlib

matplotlib.use("Agg")  # draw into a file, no window needed
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

HERE = Path(__file__).parent
PLANTS = ["1010", "1020", "1710"]           # made-up plant codes
CARRIERS = ["ROAD", "RAIL", "EXPRESS"]      # made-up shipping types


def make_deliveries(rows: int, seed: int = 42) -> pd.DataFrame:
    """Build a table that looks like outbound deliveries, with a 'late' flag."""
    rng = np.random.default_rng(seed)  # same seed, same data, every run
    df = pd.DataFrame({
        "Delivery": [str(80000000 + i) for i in range(rows)],
        "Plant": rng.choice(PLANTS, size=rows),
        "ShippingType": rng.choice(CARRIERS, size=rows, p=[0.6, 0.3, 0.1]),
        "Items": rng.integers(1, 12, size=rows),
        "WeightKg": rng.gamma(2.0, 150.0, size=rows).round(1),
        "LeadTimeDays": rng.integers(1, 15, size=rows),
    })
    # A made-up rule plus noise decides which deliveries were late.
    risk = (
        0.25 * df["Items"]
        + 0.004 * df["WeightKg"]
        - 0.35 * df["LeadTimeDays"]
        + np.where(df["Plant"] == "1710", 1.0, 0.0)
        + np.where(df["ShippingType"] == "EXPRESS", -1.5, 0.0)
        + rng.normal(0, 1.0, size=rows)
    )
    df["Late"] = (risk > 0).astype(int)
    return df


def main() -> None:
    parser = argparse.ArgumentParser(description="Unit 2 smoke test on made-up deliveries.")
    parser.add_argument("--rows", type=int, default=500, help="how many deliveries to make")
    args = parser.parse_args()

    # 1. numpy + pandas: make and inspect the data
    df = make_deliveries(args.rows)
    df.to_csv(HERE / "sample_deliveries.csv", index=False)
    print(f"Made {len(df)} deliveries. First three:")
    print(df.head(3).to_string(index=False))
    print("\nShare of late deliveries by plant:")
    by_plant = df.groupby("Plant")["Late"].mean().round(2)
    print(by_plant.to_string())

    # 2. matplotlib: one chart, saved as a file
    ax = by_plant.plot(kind="bar", color="#4a6fa5", rot=0)
    ax.set_title("Share of late deliveries by plant (made-up data)")
    ax.set_ylabel("share late")
    plt.tight_layout()
    plt.savefig(HERE / "late_by_plant.png")
    plt.close()

    # 3. scikit-learn: a first model, just to prove the library works
    features = pd.get_dummies(df[["Items", "WeightKg", "LeadTimeDays", "Plant", "ShippingType"]])
    x_train, x_test, y_train, y_test = train_test_split(
        features, df["Late"], test_size=0.25, random_state=0
    )
    model = LogisticRegression(max_iter=1000).fit(x_train, y_train)
    accuracy = model.score(x_test, y_test)
    print(f"\nModel accuracy on deliveries it has not seen: {accuracy:.2f}")
    print("Wrote sample_deliveries.csv and late_by_plant.png")


if __name__ == "__main__":
    main()

What success looks like:

Made 500 deliveries. First three:
Delivery Plant ShippingType  Items  WeightKg  LeadTimeDays  Late
80000000  1010         RAIL      6     132.1             3     1
80000001  1710         RAIL      3     256.8             1     1
80000002  1020         ROAD      3     230.0             9     0

Share of late deliveries by plant:
Plant
1010    0.46
1020    0.53
1710    0.63

Model accuracy on deliveries it has not seen: 0.85
Wrote sample_deliveries.csv and late_by_plant.png

Your numbers should match, because the seed fixes the random data. If the last decimal of the accuracy differs, that is fine: small differences between library versions can do that. Two new files appear in unit02. Click late_by_plant.png to see the chart.

What each part of the script does:

Part What it does
matplotlib.use("Agg") Draws charts into files, so the script needs no window
make_deliveries Uses numpy's random generator with a fixed seed to build a pandas table of made-up deliveries, then marks some late with a made-up rule plus noise
df.to_csv(...) Saves the table as sample_deliveries.csv for the notebook in Step 7
groupby("Plant")["Late"].mean() Share of late deliveries per plant, the kind of question you ask first
by_plant.plot(...), savefig Turns that summary into a bar chart file
pd.get_dummies Turns text columns such as Plant into 0/1 columns a model can use
train_test_split, LogisticRegression, score Holds back a quarter of the rows, trains on the rest, and reports how often the model is right on the held-back rows
--rows Optional: make more or fewer deliveries, for example --rows 2000

Step 7: Open the data in a notebook

  1. Press Ctrl+Shift+P (Cmd+Shift+P on Mac) and run Create: New Jupyter Notebook.

  2. Save it straight away (Ctrl+S or Cmd+S) as first_notebook.ipynb inside unit02.

  3. Click Select Kernel at the top right of the notebook. Choose Python Environments, then the environment inside your course folder's .venv.

  4. In the first cell, paste the code below. Run it with Shift+Enter, which runs the cell and moves to a new one.

    import pandas as pd
    
    df = pd.read_csv("sample_deliveries.csv", dtype={"Delivery": str, "Plant": str})
    df.head()

    You see a table of the first five deliveries under the cell.

  5. In the next cell, paste and run:

    df.groupby("Plant")["Late"].mean().plot(kind="bar", rot=0, title="Share late by plant")

    A bar chart appears under the cell.

  6. In a third cell, paste and run:

    df.groupby("ShippingType")["Late"].mean().round(2)
    ShippingType
    EXPRESS    0.31
    RAIL       0.59
    ROAD       0.55
    Name: Late, dtype: float64
  7. Click the Variables icon in the notebook toolbar. A panel opens at the bottom listing df. Use Show variable in data viewer next to it to scroll through all 500 rows.

  8. Save the notebook.

Step 8: Run the Unit 2 check

  1. Go back to your course folder in the terminal:

    cd ..
  2. In VS Code, create a new file in the course folder (not in unit02), paste the script below and save it as check_unit02.py.

  3. Run it:

    python check_unit02.py
"""Check that your computer is ready for Unit 2 (data science tools).

Run it from your course folder:  python check_unit02.py
It only reads your setup; it installs nothing and sends nothing anywhere.
"""
import importlib.metadata
import importlib.util
import os
import shutil
import subprocess
import sys

problems = 0


def report(ok: bool, label: str, fix: str = "", optional: bool = False) -> None:
    """Print one line: OK, MISSING (must fix) or LATER (optional for now)."""
    global problems
    if ok:
        print(f"  OK       {label}")
    elif optional:
        print(f"  LATER    {label}  ->  {fix}")
    else:
        problems += 1
        print(f"  MISSING  {label}  ->  {fix}")


def version_of(package: str) -> str:
    try:
        return importlib.metadata.version(package)
    except importlib.metadata.PackageNotFoundError:
        return ""


print("\n1. Python")
v = sys.version_info
report(v >= (3, 11), f"Python {v.major}.{v.minor}.{v.micro}",
       "Unit 2 libraries need Python 3.11 or newer; install a newer Python and recreate .venv (Step 1)")
report(sys.prefix != sys.base_prefix, "virtual environment is active", "activate .venv (Step 2)")

print("\n2. Libraries")
LIBRARIES = [  # (import name, package name on PyPI)
    ("numpy", "numpy"),
    ("pandas", "pandas"),
    ("sklearn", "scikit-learn"),
    ("matplotlib", "matplotlib"),
    ("ipykernel", "ipykernel"),
]
for module, package in LIBRARIES:
    found = importlib.util.find_spec(module) is not None
    label = f"{package} {version_of(package)}".strip()
    report(found, label, "pip install -r requirements.txt (Step 3)")

print("\n3. Libraries work together")
# Import all of them in a fresh Python and do one tiny calculation each.
# This catches broken or half-finished installs that find_spec can't see.
test = (
    "import matplotlib; matplotlib.use('Agg'); import matplotlib.pyplot as plt;"
    "import numpy as np, pandas as pd;"
    "from sklearn.linear_model import LinearRegression;"
    "df = pd.DataFrame({'x': np.arange(5), 'y': np.arange(5) * 2.0});"
    "m = LinearRegression().fit(df[['x']], df['y']);"
    "assert abs(m.coef_[0] - 2.0) < 1e-6;"
    "plt.plot(df['x'], df['y']); plt.close()"
)
try:
    result = subprocess.run([sys.executable, "-c", test], capture_output=True, text=True, timeout=180)
    ok = result.returncode == 0
    detail = (result.stderr.strip().splitlines() or ["unknown error"])[-1] if not ok else ""
except subprocess.TimeoutExpired:
    ok, detail = False, "timed out"
report(ok, "import, fit a line, draw a chart",
       f"{detail}; reinstall with pip install --force-reinstall -r requirements.txt (Step 3)")

print("\n4. Editor")
code = shutil.which("code")
if code:
    try:
        exts = subprocess.run([code, "--list-extensions"], capture_output=True, text=True, timeout=60).stdout.lower()
    except Exception:
        exts = ""
    report("ms-toolsai.jupyter" in exts, "VS Code Jupyter extension",
           "install Jupyter by Microsoft in VS Code (Step 4)", optional=True)
else:
    report(False, "VS Code 'code' command", "can't check from here; confirm Step 4 in VS Code by eye", optional=True)

print("\n5. Course folder")
report(os.path.isdir("unit02"), "unit02 folder", "create it (Step 5)", optional=True)
report(os.path.exists(os.path.join("unit02", "hello_data.py")), "unit02/hello_data.py",
       "create it (Step 6)", optional=True)

print()
if problems:
    print(f"{problems} item(s) to fix. Fix them in order, then run this again.")
    sys.exit(1)
print("All set. Your computer is ready for Unit 2.")

What success looks like:

1. Python
  OK       Python 3.11.15
  OK       virtual environment is active

2. Libraries
  OK       numpy 2.4.6
  OK       pandas 3.0.6
  OK       scikit-learn 1.9.1
  OK       matplotlib 3.11.2
  OK       ipykernel 7.4.0

3. Libraries work together
  OK       import, fit a line, draw a chart

4. Editor
  LATER    VS Code 'code' command  ->  can't check from here; confirm Step 4 in VS Code by eye

5. Course folder
  OK       unit02 folder
  OK       unit02/hello_data.py

All set. Your computer is ready for Unit 2.

This run used Python 3.11, so pip chose numpy 2.4; on Python 3.12 or newer you'll see a newer numpy. Your version numbers may differ; that is fine as long as each line says OK. The LATER line under Editor is also fine: it means the script couldn't ask VS Code directly because the code command isn't on your PATH, which is common on macOS. Step 7 working already proves the extension is there. If code is on your PATH, you see OK VS Code Jupyter extension instead.

What each part of the check does:

Part What it checks
Python Version 3.11 or newer, and that .venv is active
Libraries That each library can be found, and which version is installed
Libraries work together Starts a fresh Python, imports all four, fits a straight line and draws a chart. This catches broken installs a simple "is it there?" test misses
Editor Asks VS Code, through its code command, whether the Jupyter extension is installed
Course folder That unit02 and hello_data.py exist

Like check_setup.py, it uses only Python's built-in modules, so it runs even before Step 3. It changes nothing and sends nothing anywhere.

Step 9: Save your work in Git

  1. Tell Git to ignore notebook scratch files. Open .gitignore in your course folder and add this line:

    .ipynb_checkpoints/
  2. Save everything with Git, from the course folder:

    git add requirements.txt .gitignore check_unit02.py unit02/hello_data.py unit02/first_notebook.ipynb
    git commit -m "Set up Unit 2 data tools"

The CSV and PNG are left out on purpose: the script recreates them in a second.

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Unit 1, Step 1, then open a new terminal. macOS/Linux: use python3 until .venv is active
MISSING Python 3.10... or pip says No matching distribution found for pandas Your Python is too old for current releases Install a newer Python and recreate .venv (Step 1)
ModuleNotFoundError: No module named 'sklearn' (or pandas, numpy) The library isn't installed in the Python you're using Check for (.venv) in the prompt, then repeat Step 3.3
MISSING import, fit a line, draw a chart with an error name A library is only half installed, often after a dropped download Run pip install --force-reinstall -r requirements.txt
pip shows ProxyError, SSLError or Could not fetch URL Your network or company proxy blocks PyPI Try a home network, or ask IT for access to pypi.org and files.pythonhosted.org, or for the internal package mirror
No .venv in the notebook's kernel list VS Code hasn't found the environment, or ipykernel isn't in it Finish Step 3, then close and reopen the course folder in VS Code and try Select Kernel again
Notebook cell fails with No module named 'pandas' though the check says OK The notebook uses a different Python Click the kernel name at the top right and choose the .venv one (Step 7.3)
Notebook can't find sample_deliveries.csv The notebook isn't in unit02, or Step 6 wasn't run Save the notebook in unit02 and run python hello_data.py there
Script runs but no chart file appears You ran it from another folder and are looking in the wrong place The script writes next to itself, in unit02. Look there
Windows: .venv\Scripts\Activate.ps1 cannot be loaded PowerShell blocks scripts Run Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope CurrentUser, answer Y, and try again

Where this shows up in SAP

pandas DataFrames are also the meeting point with SAP's own Python tools. SAP publishes hana-ml, the Python machine learning client for SAP HANA, on PyPI. As described there in September 2026, it gives you a HANA DataFrame that queries data without copying it to your laptop, plus Python access to two in-database libraries: the Predictive Analysis Library (PAL) for classic algorithms and the Automated Predictive Library (APL) for automated modelling. It can turn a HANA DataFrame into a pandas DataFrame and back.

The difference matters. In this unit, data comes to the model on your laptop. With hana-ml, the model runs inside the database, next to the data. That is often the right choice for large SAP tables and for data that must not leave the system.

hana-ml is not needed for Unit 2. It needs an SAP HANA Cloud instance with the script server, PAL and APL enabled, according to SAP Learning's setup lesson, and it is licensed under the SAP Developer License Agreement. SAP Learning has a free journey, "Developing AI workflows with the Python Machine Learning Client for SAP HANA", for when you have access. The concepts you learn in Unit 2 with scikit-learn (training, test splits, metrics) carry over directly.

Pitfalls

  • Running cells out of order. A notebook remembers everything. If a result looks wrong, restart the kernel from the notebook toolbar and run all cells from the top.
  • The wrong kernel. The kernel name shows at the top right. If it isn't your .venv, imports fail or use other library versions.
  • Codes read as numbers. Plants, company codes and document numbers are text. Set dtype when reading CSV files, or leading zeros disappear.
  • Real data in notebooks. Outputs are saved inside the .ipynb file. Opening real customer data and committing the notebook puts that data in your Git history. Use sample data, and clear all outputs (from the notebook toolbar) before sharing.
  • Installing "just one more library" outside .venv. Always check for (.venv) before pip install, and add the library to requirements.txt so the setup can be rebuilt.

Exercise: make the sample data yours

The next topics use sample_deliveries.csv. Make a version that fits the business you know.

  1. Copy hello_data.py to my_deliveries.py in unit02.

  2. In make_deliveries, add a column Customer with three made-up customer numbers, chosen at random like Plant is.

  3. Add one line to the risk calculation so deliveries to one of those customers are more often late.

  4. Change the output file names to my_deliveries.csv and my_late_by_plant.png.

  5. Run python my_deliveries.py. Check that the new CSV has a Customer column.

  6. In first_notebook.ipynb, add a cell that reads my_deliveries.csv and shows the share of late deliveries by Customer. Check that your chosen customer has the highest share.

  7. Save your work in Git:

    git add unit02/my_deliveries.py unit02/first_notebook.ipynb
    git commit -m "Add my own sample deliveries"

Done when: check_unit02.py prints All set, my_deliveries.csv has a Customer column, the notebook shows your chosen customer as the most often late, and git log shows the commit. Keep the script: Unit 2's late-delivery prediction topic trains on data like this.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Which package name do you install, and which name do you import, for scikit-learn?

    Answer: D. You install scikit-learn and write import sklearn. Install names and import names don't always match.
  2. 2What does a notebook's kernel remember between cells, and how can that give you a wrong result?

    Answer: C. Every variable and import from cells already run. Results then depend on which cells ran and in what order. If something looks wrong, restart the kernel and run all cells from the top.
  3. 3Why does the notebook need ipykernel in .venv, but not the full Jupyter package?

    Answer: B. VS Code's Jupyter extension provides the notebook interface. The environment only needs ipykernel so it can act as the kernel, the Python process that runs the cells with your course libraries.
  4. 4Why read Plant and Delivery as text rather than numbers?

    Answer: A. They are SAP codes, not quantities. Read as numbers, leading zeros disappear and codes can be summed or sorted wrongly. Set dtype to str when reading the CSV.
  5. 5Your check script says OK for every library, but a notebook cell fails with No module named 'pandas'. What is the likely cause?

    Answer: D. The notebook uses a different Python. Click the kernel name at the top right and choose the one inside your course .venv.
  6. 6How do the four data libraries depend on each other?

    Answer: C. numpy is the base for arithmetic on arrays. pandas builds tables on numpy, scikit-learn builds models on numpy and SciPy, and matplotlib draws charts from numpy data. pip installs the extra dependencies for you.
  7. 7Why do the numbers from hello_data.py match the page?

    Answer: B. It makes 500 made-up, SAP-shaped deliveries, summarizes late deliveries by plant, draws a chart, and trains and scores a first model. A fixed random seed makes the data identical on every run; small version differences may change the last decimal.
  8. 8When would you prefer SAP's hana-ml, which runs the model inside SAP HANA, over pandas and scikit-learn on a laptop?

    Answer: A. For large SAP tables and for data that must not leave the system, because the model runs in the database next to the data. It needs an SAP HANA Cloud instance with the right features enabled, and isn't needed for Unit 2.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in