FIXME
Overview
Time estimation: 3H
Version: main
Last update: 2025-11-03
Questions:Objectives:
How to create a synthetic datasets for future inference?
How to serve model with FastAPI?
Learn best practices to serve model with FastAPI and get predictions as a table.
Deploying and maintaining machine learning (ML) models can be challenging, especially in academia where ….
Machine Learning Operations (MLOps) addresses these challenges by streamlining and automating the ML lifecycle - from development to deployment and ongoing management. In this hands-on tutorial, you will learn some basic MLOps practices and tools through an end-to-end project. You will gain practical experience with cloud infrastructure, containerization, orchestration, and model serving—core skills that make ML projects more scalable, reproducible, and production-ready.
Specifically, you will learn how to (1) deploy a trained model on the cloud; (2) containerize it with Docker for portability, (3) scale it using Kubernetes, and (4) serve the model with FastAPI via an API interface.
As a proof of concept, we will use the Breast Cancer Wisconsin dataset from sklearn.datasets module. We will focus on deployment and serving rather than preprocessing, so you can gain hands-on experience with moving models into production.
Note
This tutorial is intended for beginners in MLOps and focuses on the deployment and serving of machine learning workflows in the cloud. It does not cover the theoretical foundations of machine learning or the full range of MLOps tools and best practices. We assume you already have basic to intermediate knowledge of machine learning, including model training and the use of libraries such as scikit-learn.
Before you start, ensure you meet the following requirements:
If you are already a member of a SimpleVM or Openstack project in de.NBI Cloud, you can proceed to create a VM instance. If you are unfamiliar with this process, we recommend reviewing our SimpleVM tutorial and the SimpleVM and Openstack Wikis for step-by-step guidance.
For this hands-on tutorial, we need to install the following packages:
Install
venvandpip3sudo apt update sudo apt install python3-venv python3-pip -y pip3 --version
Activate virtual environment and upgrade
pippython3 -m venv ~/mlenv source ~/mlenv/bin/activate
Install the required packages:
pip install pandas pip3 install -U scikit-learn pip install sdv pip install "fastapi[standard]" pip install uvicorn
sklearn.datasets has lots of datasets for benchmarking and the utility to generate the one for the task of classifition (sklearn.dataset.make_classification) and regression (sklearn.datasets.make_regression).
Load the breast cancer Wisconsin dataset and inspect the features and target values which are benign (0) and malignant (1).
Dataset description
The dataset consists of 569 samples, 30 features, and 2 classes (benign, malignant). Features are extracted from digitized FNA images of breast masses, capturing key characteristics of the cell nuclei.
10 features are computed for each cell nucleus:
- radius (mean of distances from center to points on the perimeter)
- texture (standard deviation of gray-scale values)
- perimeter
- area
- smoothness (local variation in radius lengths)
- compactness (perimeter² / area — 1.0)
- concavity (severity of concave portions of the contour)
- concave points (number of concave portions of the contour)
- symmetry
- fractal dimension (“coastline approximation” — 1)
The mean, standard error and “worst” or largest (mean of the three largest values) of these features were computed for each image, resulting in 30 features.
Load the dataset
from sklearn.datasets import load_breast_cancer # returns dictionary with features and target values load_breast_cancer() # load dataset into pandas DataFrame and target Series X, y = load_breast_cancer(return_X_y=True, as_frame=True)
Here,
return_X_y function returns a tuple (data, target) instead of a Bunch object. Default is False.as_frame function returns the data and target as pandas.DataFrame and pandas.Series, making it easier to inspect features.Inspect the dataset
- Check the dataset for missing values.
- Identify features that are highly correlated with one another.
- Identify features with low variance across samples (i.e., uninformative features) using
VarianceThresholdclass fromsklearn.feature_selectionmodule.- Is the
targetcolumn balanced in the dataset?Solution
import numpy as np from sklearn.feature_selection import VarianceThreshold # check features with missing values X.isnull().any() # compute the absolute correlation matrix corr_matrix = X.corr().abs() # select the upper or lower triangular part of the correlation matrix upper_triangle = corr_matrix.where(np.triu(np.ones(corr_matrix.shape), k=1).astype(bool)) # select the features with an absolute correlation greater than arbitraty 0.80 to_drop = [column for column in upper_triangle.columns if any(upper_triangle[column] > 0.80)] print(len(to_drop)) # select features with non-zero variance selector = VarianceThreshold(threshold=0) selector.fit(X) # features with low variance low_variance_features = X.columns[~selector.get_support()] print(low_variance_features.tolist()) # check class distribution as percentages/probabilities print(y.value_counts(normalize=True))
Next, we will generate a synthetic dataset that mimics the statistical properties and correlations of the breast cancer Wisconsin dataset. For this, we will use SDV - Synthetic Data Vault - a library specifically designed for generating realistic synthetic data. It supports single-table and multi-table datasets, allowing you to create synthetic data for testing, development, or privacy-sensitive applications. This will allow us to get an inference without relying only on the original dataset. If you would like to dive deeper into how SDV works, you can check out the official documentation or read the research paper.
We will use the Gaussian Copula algorithm because it is both fast and straightforward. The GaussianCopulaSynthesizer generates synthetic data by applying classical statistical methods to learn the structure of your dataset. You need to provide a Metadata object as the first argument to GaussianCopulaSynthesizer.
Metadata object describes the dataset you want to synthesize. All other parameters are optional and can be used to customize the behavior of the synthesizer. The metadata includes essential information about your dataset, such as:
Creating Metadata
Create metadata for the breast cancer Wisconsin dataset using
Metadataclass fromsdv.metadatamodule.Solution
# create metadata from sdv.metadata import Metadata metadata = Metadata.detect_from_dataframe( data=real_data, table_name='breast_cancer') # inspect metadata metadata.visualize()Mark personally-identifiable information
If your dataset contains sensitive information, you can mark personally identifiable information (PII) in the metadata to ensure it is handled appropriately. As a proof of concept, we will generate fake names using the faker.
pip install Faker# generate fake but realistic-looking names from faker import Faker fake = Faker() # create fake names for all samples real_data['Patient'] = [fake.name() for _ in range(len(real_data))] # regenerate metadata metadata.detect_from_dataframe(data=real_data, table_name="breast_cancer") # update metadata and mark sensitive information metadata.update_column(column_name="Patient", sdtype="name", pii=True)Create Metadata for multiple tables
Though it is not the scope of this tutorial, it would be useful for you to also learn how to create a metadata for multiple tables and further generate a synthetic datasets.
Create synthetic dataset for future inference
from sdv.single_table import GaussianCopulaSynthesizer import pandas as pd # combine features and targets real_data = pd.concat([X, y], axis=1) # initiate sdv synthesizer with created metadata synthesizer = GaussianCopulaSynthesizer(metadata) # learn patterns from the data synthesizer.fit(data=real_data) # generate 100 samples synthetic_data = synthesizer.sample(num_rows=100) # visualize the real vs. synthetic data from sdv.evaluation.single_table import get_column_plot fig = get_column_plot( real_data=real_data, synthetic_data=synthetic_data, column_name='mean radius', # chose any column ) fig.show()
You can also compare the distribution of the target values in the real and synthetic datasets to verify that they are similar. Once satisfied, save the synthetic data (without target column) as external_test.tsv for use in future inference tasks.
At this stage, you have explored the dataset and generated synthetic data for future inference. Now, we will train a classifier and prepare it for deployment via an API. Create a python script train.py that loads data, trains a model, evaluates performance, and saves the trained model. We will train a Random Forest classifier with the default settings as its robust, easy to use, and work well for structured tabular data.
Train Random Forest with the default parameters
- Divide dataset into into training and test sets. Use
train_test_splitfromsklearn.model_selectionwith a test size of 10% and stratification to preserve the class distribution.- Train Random Forest with the default parameters. Use
RandomForestClassifierfromsklearn.ensemblemodule.- Generate predictions for the test set and compute a classification report.
- Save the trained model
joblib.Solution
from sklearn.model_selection import train_test_split from sklearn.ensemble import RandomForestClassifier from sklearn.metrics import * import joblib # split data X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.1, stratify=y, shuffle=True, random_state=42) # train classifier clf = RandomForestClassifier(random_state=42) clf.fit(X_train, y_train) # evaluate the model test_pred = clf.predict(X_test) # get classification report report = classification_report(y_test, test_pred, output_dict=True) report = pd.DataFrame(report).transpose() report.reset_index(inplace=True) print(report) # save the trained model joblib.dump(clf, "model.joblib")
Confusion matrix report
To further evaluate your model’s performance, you can use pycm - python confusion matrix - a library that provides detailed statistics and metrics for classification models.
pip install pycmfrom pycm import * # create a confusion matrix cm = ConfusionMatrix(actual_vector=y_test.to_numpy(), predict_vector=test_pred) # print the matrix cm.print_matrix() # alternatively save a detailed HTML report cm.save_html("breast_cancer_pycm_report", summary=True, color=(0, 64, 64))The detailed HTML report with overall statistics can be opened in your browser.
Since the classifier is performing well, we can proceed to the next step: creating an API using FastAPI. For this hands-on tutorial, we will focus only on the essential features of FastAPI needed to serve our trained model, keeping the implementation simple and beginner-friendly. If you would like to learn more, you can refer to the official documentation.
Create app folder in your working directory and place the trained model in there because this is our server and everything required for this server is inside that folder. Then create a new python file main.py whithin app folder, import the required libraries, and load the trained model.
main.pyfrom fastapi import FastAPI, UploadFile, File # to handle HTTP requests and file uploads from fastapi.responses import StreamingResponse # to send large files (e.g., CSV/TSV) back to the client efficiently import pandas as pd import joblib # to load pre-trained model import io # for in-memory file operations # load the trained model model = joblib.load("app/model.joblib") # define classes class_names = {"benign": 0, "malignant": 1} # maps numeric predictions to meaningful class labels
Initialize an app and define root endpoint
app = FastAPI() # creates the main API application object that will handle all routes and endpoints # root endpoint @app.get("/") def read_root(): """ Root endpoint for checking if the API is running. Returns a simple JSON message. """ return {"message": "Breast cancer inference API"}
Next, define prediction endpoint that accepts a tab-separated CVS/TSV file with features and make predictions using our trained model.
Define prediction endpoint
@app.post("/predict/") async def predict_from_csv(file: UploadFile = File(...)): """ Endpoint for making predictions from a tab-separated CSV/TSV file. Args: file (UploadFile): Uploaded tab-separated CSV file with feature columns. Returns: StreamingResponse: A downloadable CSV file with an added `prediction` column. """ # read uploaded file into pandas DataFrame contents = await file.read() data = pd.read_csv(io.BytesIO(contents), sep="\t") print(data.sample(5)) # optional: inspect sample rows # make predictions using loaded model predictions = model.predict(data) data["prediction"] = class_names[predictions] # save predictions to an in-memory CSV buffer output = io.StringIO() data.to_csv(output, sep="\t", index=False) output.seek(0) # reset buffer pointer to the beginning # return the predictions as downloadable TSV file return StreamingResponse( output, media_type="text/csv", headers={"Content-Disposition": "attachment; filename=predictions.tsv"} )
After creating main.py, you can run the FastAPI API and test it with a sample dataset.
To start the server, open a new terminal, run the command, and click the server link to check the root endpoint:
fastapi dev app/main.py
When you are running the API on a remote VM, establish an SSH tunnel to access it from your local browser:
ssh -N -L localhost:XXXX:localhost:XXXX -i PRIVATE_SSH_KEY INSTANCE@FLOATING_IP
Here,
XXXX is the port number, e.g., 8000.-i PRIVATE_SSH_KEY path to your private SSH key for authentication.INSTANCE corresponds to your username on the remote VM (you can verify it by running whoami in your previous terminal session).FLOATING_IP is public IP of your remote VM.Once the ssh tunnel is established, open your browser at http://127.0.0.1:8000. You should see the JSON response:
{"message":"Breast cancer inference API"}
FastAPI provides interactive documentation for testing endpoints:
To test the /predict/ endpoint, click Try it out, upload your features table (e.g., external_test.tsv, the synthetic dataset created earlier), click Execute, and download the resulting file with predictions.
You can also test the API from the command line:
curl -X 'POST' 'http://127.0.0.1:8000/predict/' -F "file=@external_test.tsv" --output predictions.tsv
Here,
-X POST sends a POST request.-F "file=@external_test.tsv" uploads the test file.-o predictions.tsv saves the API response (predictions) as a local file.Add data processing endpoint
Create a FastAPI POST endpoint
/preprocess/normalize/that does the following:
- accepts a tab-separated CSV/TSV file;
- normalizes all numeric features to the range [0,1] using
MinMaxScalerfromsklearn.preprocessingmodule;- returns the processed table.
Solution
from fastapi import FastAPI, UploadFile, File from fastapi.responses import StreamingResponse import pandas as pd import io from sklearn.preprocessing import MinMaxScaler app = FastAPI() @app.post("/preprocess/normalize/") async def normalize_features(file: UploadFile = File(...)): # read uploaded file into pandas DataFrame contents = await file.read() data = pd.read_csv(io.BytesIO(contents), sep="\t") print(data.head(5)) # optional: inspect the dataset # identify numeric columns numeric_cols = data.select_dtypes(include='number').columns # apply scaling scaler = MinMaxScaler() data[numeric_cols] = scaler.fit_transform(data[numeric_cols]) # save processed data output = io.StringIO() data.to_csv(output, sep="\t", index=False) output.seek(0) return StreamingResponse( output, media_type="text/csv", headers={"Content-Disposition": "attachment; filename=normalized.tsv"} )
At this point, you have created a basic FastAPI application to serve your trained model. Next, we will wrap the code, dependencies, and API in a Docker container for easy deployment and portability.
To deploy our classifier across different environments (to solve the so-called “It works on my machine” issue 😉), we need to containerize the application. Docker allows us to package the code, dependencies, and runtime environment into a single, portable container.
Create a requirements.txt file in the root of your project with the exact versions of libraries used for training your model. This ensures compatibility, especially for scikit-learn, which must match the version used to train the model.
scikit-learn==1.6.1
fastapi
pandas
joblib
uvicorn
Create Dockerfile in your project root. The Dockerfile specifies the operating system, python version, and dependencies your application needs.
touch Dockerfile
Add the following code:
# base image
FROM python:3.12
# set work directory
WORKDIR /app
# copy all files to app in container
COPY . /app
# install dependencies listed in the file
RUN pip install --no-cache-dir -r /app/requirements.txt
# expose port
EXPOSE 8000
# run the FastAPI app using uvicorn
CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000"]
The Dockerfile does the following:
uvicorn handles web requests and serves the FastAPI app.app.main:app points to the FastAPI instance in main.py under the app directory.--host 0.0.0.0 makes the app accessible outside the container.--port 8000 binds the server to port 8000 (matching the exposed port).Build the container image with the following command:
docker build -t breast-cancer-inference .
Here,
-t names the Docker image.This step takes some time usually when run for the first time.
Run the container locally and map port 8000 from the container to your host machine:
docker run -p 8000:8000 breast-cancer-inference
Here,
-p 8000:8000 maps host port 8000 to container port 8000.The FastAPI app is now running inside the container and accessible at http://127.0.0.1:8000.
You can stop a running Docker container using container ID or name:
# show running containers
docker ps
docker stop <container_id_or_name>
You have a containerized ML model served via FastAPI, ready to be deployed consistently across any environment.
The next step is deploying and scaling this container using Kubernetes for production-grade availability and management.
Kubernetes (K8s) is a container orchestration platform to deploy, manage, and scale containers. It would allow us to run our FastAPI container reliably in production, enable horizontal scaling for handling multiple requests, and automate deployment and management of ML APIs.
If you are already familiar with kubernetes, you can skip the introduction section and proceed with the next steps.
Briefly explain what Kubernetes (K8s) is …
Kubernetes manages containerized applications by scaling resourcing, managing container health and availability. And it is particularly helpful for machine learning applications.
Kubernetes improves ML scalability by dynamically adjusting resources based on demand. This ensures optimal performance under varying loads and reduces the need for manual intervention.
For now, let’s understand 3 components: pods, services, and deployments.
Before starting, ensure you have:
We have already built the Docker image breast-cancer-inference.
What to learn next?
ML model deployment is not the last stage in model management. Learn how to continuously validate/monitor data and train model:
Key Points
FIXME
Author(s): Dilfuza Djamalova
Supported by: