<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Machine Learning on Zack's Blog</title><link>https://zackblog.work/categories/machine-learning/</link><description>Recent content in Machine Learning on Zack's Blog</description><generator>Hugo</generator><language>en-au</language><lastBuildDate>Sun, 01 Mar 2026 05:27:11 +0000</lastBuildDate><atom:link href="https://zackblog.work/categories/machine-learning/index.xml" rel="self" type="application/rss+xml"/><item><title>Can AI Actually Build a Complex Project? A Brutal RAG System Post-Mortem</title><link>https://zackblog.work/posts/can-ai-actually-build-a-complex-project-a-brutal-rag-system-post-mortem/</link><pubDate>Sun, 01 Mar 2026 05:27:11 +0000</pubDate><guid>https://zackblog.work/posts/can-ai-actually-build-a-complex-project-a-brutal-rag-system-post-mortem/</guid><description>&lt;p&gt;Last month, I did something that, looking back, was probably the most educational practice of my year. I tried using Claude and GPT to build a complete multi-source RAG system from scratch. This wasn&amp;rsquo;t a &amp;ldquo;toy demo&amp;rdquo; for social media; it was a real internal tool for my lab designed to handle messy document data, query rewriting, hybrid search, and deployment for my team.&lt;/p&gt;
&lt;p&gt;I set a strict rule for myself: &lt;strong&gt;Let the AI write almost all the code, while I handle the requirements and decision-making.&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Improve Billing &amp; Health Analyzer with Bedrock v2</title><link>https://zackblog.work/posts/improve-billing-health-analyzer-with-bedrock-v2/</link><pubDate>Wed, 21 Jan 2026 04:50:08 +0000</pubDate><guid>https://zackblog.work/posts/improve-billing-health-analyzer-with-bedrock-v2/</guid><description>&lt;p&gt;I constantly rely on AWS Billing and Cost Management and Trusted Advisor to monitor costs and security. However, manual billing reviews are time-consuming, and raw Cost Explorer data doesn&amp;rsquo;t provide the full picture. What if an intelligent program could proactively call the Billing and Trusted Advisor APIs, retrieve the raw data, and pass it to an LLM for in-depth analysis of account health—then deliver actionable insights and AI-powered recommendations? That would save a significant amount of time and effort.&lt;/p&gt;</description></item><item><title>Serverless Billing &amp; Health Analyzer with Bedrock</title><link>https://zackblog.work/posts/serverless-billing-health-analyzer-with-bedrock/</link><pubDate>Sat, 18 Oct 2025 11:42:05 +0000</pubDate><guid>https://zackblog.work/posts/serverless-billing-health-analyzer-with-bedrock/</guid><description>&lt;p&gt;I constantly rely on AWS Billing and Cost Management and Trusted Advisor to monitor costs and security. However, manual billing reviews are time-consuming, and raw Cost Explorer data doesn’t provide the full picture. What if an intelligent program could proactively call the Billing and Trusted Advisor APIs, retrieve the raw data, and pass it to an LLM for in-depth analysis of account health—then deliver actionable insights and AI-powered recommendations? That would save a significant amount of time and effort.&lt;/p&gt;</description></item><item><title>Boost EKS RAG with LangChain</title><link>https://zackblog.work/posts/boost-eks-rag-with-langchain/</link><pubDate>Sat, 11 Oct 2025 05:21:03 +0000</pubDate><guid>https://zackblog.work/posts/boost-eks-rag-with-langchain/</guid><description>&lt;p&gt;After building a &lt;a href="https://zackblog.work/posts/bedrock-powered-rag-on-eks/"&gt;custom RAG system on Amazon EKS&lt;/a&gt;, my son asked why the application could only handle one question at a time, while ChatGPT allows users to continue conversations through follow-up questions with contextual memory. That made me realize there was an opportunity to leverage the power of the LangChain framework for better abstractions, conversational memory, and access to the broader LangChain ecosystem.&lt;/p&gt;
&lt;p&gt;This post documents the journey of migrating to a LangChain-powered solution — the challenges faced, and the new capabilities and benefits gained.&lt;/p&gt;</description></item><item><title>AWS Bedrock with Open WebUI</title><link>https://zackblog.work/posts/aws-bedrock-with-open-webui/</link><pubDate>Tue, 16 Sep 2025 14:56:30 +0000</pubDate><guid>https://zackblog.work/posts/aws-bedrock-with-open-webui/</guid><description>&lt;p&gt;Recently Data team reached out, trying to build an EC2 running Open WebUI, connecting to AWS Bedrock, to offer team member AI application. This guide provides step-by-step instructions for deploying and integrating with AWS Bedrock models. .&lt;/p&gt;
&lt;p&gt;The architecture consists of two Docker containers operating on a single EC2 instance: Open WebUI serves as the user-facing chat interface, while a Bedrock Access Gateway acts as middleware. This gateway securely forwards requests from Open WebUI to the AWS Bedrock API using the EC2 instance&amp;rsquo;s attached IAM role. The containers communicate over a private Docker network, isolating traffic between them.&lt;/p&gt;</description></item><item><title>AI Coding - Build an Electricity and Gas Analytic Platform</title><link>https://zackblog.work/posts/ai-coding-build-an-electricity-and-gas-analytic-platform/</link><pubDate>Thu, 14 Aug 2025 07:45:20 +0000</pubDate><guid>https://zackblog.work/posts/ai-coding-build-an-electricity-and-gas-analytic-platform/</guid><description>&lt;p&gt;In the past few days, I got an idea from my recent electricity and gas bills—thinking about how to run data analysis to examine usage patterns and trends using the Python approach: employing Pandas for reading data, using Numpy for normalisation, and Matplotlib to create a few visualisations.&lt;/p&gt;
&lt;p&gt;However, this time I wanted to try talking to Gemini CLI via the terminal to see if it could generate a web-based analytical platform for me, as someone with no developer background.&lt;/p&gt;</description></item><item><title>Exploratory Data Analysis (EDA) with Peppa Pig</title><link>https://zackblog.work/posts/exploratory-data-analysis-eda-with-peppa-pig/</link><pubDate>Sun, 27 Jul 2025 05:56:03 +0000</pubDate><guid>https://zackblog.work/posts/exploratory-data-analysis-eda-with-peppa-pig/</guid><description>&lt;p&gt;This was a linguistic analysis project where the primary goal was not just to count words, but to evaluate the language, themes, and emotional tone of the children&amp;rsquo;s show &amp;ldquo;Peppa Pig&amp;rdquo; (specifically, the first four seasons) to determine its suitability for a pre-kindergarten audience. In this study, I will try to answer a broader question:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;Beyond a simple word list, what can a multi-faceted data analysis tell us about the show&amp;rsquo;s true educational and emotional value?&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>MLOPS - Enhancing Oscar Model with LightGBM</title><link>https://zackblog.work/posts/mlops-enhancing-oscar-model-with-lightgbm/</link><pubDate>Fri, 21 Mar 2025 13:36:27 +0000</pubDate><guid>https://zackblog.work/posts/mlops-enhancing-oscar-model-with-lightgbm/</guid><description>&lt;p&gt;In a previous post - &lt;a href="https://zackblog.work/posts/mlops-build-a-oscar-best-picture-winner-model/"&gt;MLOps - Build a Oscar Best Picture Winner Model&lt;/a&gt;, I was able to establish a baseline model using a &lt;code&gt;RandomForestClassifier&lt;/code&gt; to predict the Oscar for Best Picture. This classic workflow involved data cleaning, training, and prediction, providing a solid starting point.&lt;/p&gt;
&lt;p&gt;However, a deeper look at the results revealed critical weaknesses that an experienced machine learning engineer would immediately flag:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Inadequate Model Choice:&lt;/strong&gt; The initial model wasn&amp;rsquo;t powerful enough for the task. The classification report showed a &lt;strong&gt;recall of 0.00 for the &amp;ldquo;winner&amp;rdquo; class&lt;/strong&gt;. This is a major red flag, indicating the model completely failed to identify any actual winners, likely due to the severe class imbalance.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Misleading Evaluation Metrics:&lt;/strong&gt; I think I relied too heavily on accuracy. On an imbalanced dataset, a model can achieve high accuracy simply by always predicting the majority class. Better to shift our focus to more robust metrics like the &lt;strong&gt;F1-score&lt;/strong&gt;, &lt;strong&gt;ROC AUC&lt;/strong&gt;, and &lt;strong&gt;Precision-Recall AUC&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This analysis led to idea to enhance this Oscar prediction with a more sophisticated &lt;strong&gt;LightGBM (LGBM) classifier&lt;/strong&gt; model, known for its high performance, speed, and efficiency on tabular data.&lt;/p&gt;</description></item><item><title>MLOps - Build a Oscar Best Picture Winner Model</title><link>https://zackblog.work/posts/mlops-build-a-oscar-best-picture-winner-model/</link><pubDate>Sun, 23 Feb 2025 09:37:43 +0000</pubDate><guid>https://zackblog.work/posts/mlops-build-a-oscar-best-picture-winner-model/</guid><description>&lt;p&gt;In this post, we will continue to build a basic machine learning model to predict the &lt;strong&gt;Best Picture&lt;/strong&gt; winner at the Academy Awards (Oscar).&lt;/p&gt;
&lt;p&gt;We will use our previous processed dataset that includes information about the nominees and winners from the 72nd to the 96th Oscar ceremonies. The goal is to predict the winner based on various features like IMDb ratings, Metascore, Tomatometer percentage, Golden Globe and BAFTA wins/nominations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Step 1: Understand the features and model&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>MLOps - Data Processing for Oscar Winner Model</title><link>https://zackblog.work/posts/mlops-data-processing-for-oscar-winner-model/</link><pubDate>Sun, 23 Feb 2025 05:35:54 +0000</pubDate><guid>https://zackblog.work/posts/mlops-data-processing-for-oscar-winner-model/</guid><description>&lt;p&gt;After downloading a dataset from Kaggle about historical Oscar nominations and winners, it is time to move on to data cleansing and enrichment.&lt;/p&gt;
&lt;p&gt;The original data contained seven headers, but only a few were useful: the film name, ceremony year, and winner status. Additionally, the dataset spanned from 1927 to 2024 and included all Oscar categories, making it quite complex—perhaps too overwhelming for me as a beginner. So, I decided to focus on &lt;strong&gt;Best Picture&lt;/strong&gt; as it is always the most important award among all the other categories.&lt;/p&gt;</description></item><item><title>MLOps - How About Predict Oscar Winner</title><link>https://zackblog.work/posts/mlops-how-about-predict-oscar-winner/</link><pubDate>Fri, 21 Feb 2025 11:50:12 +0000</pubDate><guid>https://zackblog.work/posts/mlops-how-about-predict-oscar-winner/</guid><description>&lt;p&gt;&lt;strong&gt;The Idea&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The 97th Academy Awards ceremony, presented by the Academy of Motion Picture Arts and Sciences (AMPAS), will take place on March 2, 2025, at the Dolby Theatre in Hollywood, Los Angeles.
Last year 2024 I had some greate experience with some great movies like Dune Part2 and Wicked, not sure if some of my faviourate actors can grab a Oscar.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://zackblog.work/images/mlops120.png"&gt;&lt;img alt="image tooltip here" loading="lazy" src="https://zackblog.work/images/mlops120.png"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;so Why not go and use historical Oscar nominations dataset for the past 20 years, to create and train a model by feeding categories and results, to predict the winners each year, then input this years&amp;rsquo; nominations to get a prediction ??&lt;/p&gt;</description></item><item><title>MLOps - Get Started with KubeRay</title><link>https://zackblog.work/posts/mlops-get-started-with-kuberay/</link><pubDate>Tue, 11 Feb 2025 04:35:22 +0000</pubDate><guid>https://zackblog.work/posts/mlops-get-started-with-kuberay/</guid><description>&lt;p&gt;Ray is an open-source framework designed for scalable and distributed ML workloads, including training, tuning, and inference. It provides a simple API for scaling Python applications across multiple nodes.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;Distributed Training&lt;/code&gt;: Easily scale PyTorch, TensorFlow, and other ML jobs across multiple GPUs/instances.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Hyperparameter Tuning&lt;/code&gt;: Integrates with Optuna and Ray Tune for distributed hyperparameter optimization.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Parallel Inference&lt;/code&gt;: Supports inference pipelines that scale out dynamically based on demand.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Fault Tolerance&lt;/code&gt;: If a node fails, Ray can reschedule tasks on other available nodes.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Together with EKS, Karpenter, and Ray, a modern ML team can achieve dynamic Auto-Scaling and GPU resource allocation from local deployment to Running Distributed ML Jobs in Cloud Production environment:&lt;/p&gt;</description></item><item><title>MLOps - Deploy ML workload into EKS with Karpenter</title><link>https://zackblog.work/posts/mlops-deploy-ml-workload-into-eks-with-karpenter/</link><pubDate>Sat, 01 Feb 2025 04:35:11 +0000</pubDate><guid>https://zackblog.work/posts/mlops-deploy-ml-workload-into-eks-with-karpenter/</guid><description>&lt;p&gt;Finally, machine learning workload into AWS EKS with Karpenter&lt;/p&gt;
&lt;p&gt;In the previous post, I was able to complete both local Pytorch ML and AWS SageMaker practice, and containerize and deploy ML docker image locally.&lt;/p&gt;
&lt;p&gt;In this post, I will explore the model deployment to EKS cluster with Karpenter to simulate a more scalable and real-life production-ready environment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Challenges when moving to Cloud deployment&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Here are key differences between local PyTorch model vs AWS SageMaker artifacts, and how they impact the Docker image size and Performance considerations for LLM images in EKS deployment.&lt;/p&gt;</description></item><item><title>MLOps - Deploy Classifier with AWS Serverless</title><link>https://zackblog.work/posts/mlops-deploy-classifier-with-aws-serverless/</link><pubDate>Sat, 25 Jan 2025 04:34:58 +0000</pubDate><guid>https://zackblog.work/posts/mlops-deploy-classifier-with-aws-serverless/</guid><description>&lt;p&gt;Continue Pneumonia Classifier by moving to AWS Serverless&lt;/p&gt;
&lt;p&gt;Go AWS Serverless Deployment&lt;/p&gt;
&lt;p&gt;Deploying machine learning models in production requires additional considerations to address latency, scalability, cost-efficiency, and monitoring.&lt;/p&gt;
&lt;p&gt;A modern approach to hosting an ML application in AWS can be considered as a &lt;code&gt;serverless architecture&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;This allows users to upload images via a S3 static web page and send them to API Gateway. The API Gateway receives the HTTP POST request and forwards it to the Lambda function, which handles the image preprocessing, inference, and postprocessing logic. The Lambda function sends the image payload to the SageMaker endpoint, by calling the SageMaker endpoint using the SageMaker Runtime SDK (&lt;code&gt;invoke_endpoint&lt;/code&gt;) which hosts my trained model, and then retrieves the prediction. The prediction result is sent back to the frontend for display.&lt;/p&gt;</description></item><item><title>MLOps - Move Image Classifier to AWS SageMaker</title><link>https://zackblog.work/posts/mlops-move-image-classifier-to-aws-sagemaker/</link><pubDate>Sat, 25 Jan 2025 04:34:41 +0000</pubDate><guid>https://zackblog.work/posts/mlops-move-image-classifier-to-aws-sagemaker/</guid><description>&lt;p&gt;Continue Pneumonia Classifier by moving to AWS SageMaker&lt;/p&gt;
&lt;p&gt;In this post, we will continue Pneumonia Classifier practice by leveraging AWS managed AI service SageMaker. We will:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Provision AWS Sagemaker AI domain and workspace.&lt;/li&gt;
&lt;li&gt;Run a JupiterLab to download, process, and upload chest X-ray dataset from Kaggle to S3 for SageMaker access.&lt;/li&gt;
&lt;li&gt;Use SageMaker Pre-built image classification algorithm, to define SageMaker Estimator to configure the SageMaker training job, including compute resources, training duration, and data input method.&lt;/li&gt;
&lt;li&gt;Count the number of training samples and set the hyperparameters, perform hyperparameter tuning to find the best configuration for the model.&lt;/li&gt;
&lt;li&gt;Launching the Hyperparameter Tuning Job.&lt;/li&gt;
&lt;li&gt;Using CloudWatch and SageMaker Train Job to monitor and troubleshoot.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;In the AWS SageMaker AI JupiterLab&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>MLOps - Containerize Classifier Application</title><link>https://zackblog.work/posts/mlops-containerize-classifier-application/</link><pubDate>Sat, 25 Jan 2025 04:34:14 +0000</pubDate><guid>https://zackblog.work/posts/mlops-containerize-classifier-application/</guid><description>&lt;ul&gt;
&lt;li&gt;Last post we were able to build and train the pneumonia classifier model.&lt;/li&gt;
&lt;li&gt;In this post we are moving a step forward to: &lt;strong&gt;Containerizing the model&lt;/strong&gt; and creating a &lt;strong&gt;frontend-backend application&lt;/strong&gt; to allow users to upload images and get predictions.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Overview of the Application Architecture&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A high-level diagram or description of the system architecture:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Backend&lt;/strong&gt;: A Dockerized Flask/FastAPI application serving the trained model.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Frontend&lt;/strong&gt;: A Dockerized React/Streamlit app for image upload and displaying predictions.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Interaction&lt;/strong&gt;: Frontend sends images to the backend, which processes them and returns predictions.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Step 1: Containerizing the Backend&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>MLOps - Build a Image Classifier Model</title><link>https://zackblog.work/posts/mlops-build-a-image-classifier-model/</link><pubDate>Sat, 25 Jan 2025 04:34:02 +0000</pubDate><guid>https://zackblog.work/posts/mlops-build-a-image-classifier-model/</guid><description>&lt;p&gt;Pneumonia causes over 2.5 million deaths annually worldwide. Traditional diagnosis through chest X-ray analysis is time-consuming and requires expert radiologists. This project aims to develop an automated deep learning system to classify pneumonia from chest X-rays with high accuracy, potentially assisting healthcare professionals in faster diagnosis.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Purpose&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Here I will use a local environment with Pytorch and Jupyter Notebook to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Demonstrate end-to-end development of a medical image classifier&lt;/li&gt;
&lt;li&gt;Using Kaggle dataset for NIH Chest X-Ray image processing&lt;/li&gt;
&lt;li&gt;Create a baseline model for pneumonia detection (Normal vs Pneumonia)&lt;/li&gt;
&lt;li&gt;Local GPU acceleration for model training and prediction&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Typical ML steps to achieve image classifier:&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>MLOps - Deploy ML workload to K8S</title><link>https://zackblog.work/posts/mlops-deploy-ml-workload-to-k8s/</link><pubDate>Mon, 20 Jan 2025 04:33:48 +0000</pubDate><guid>https://zackblog.work/posts/mlops-deploy-ml-workload-to-k8s/</guid><description>&lt;p&gt;Machine Learning workload deployed in K8S and minikube!!&lt;/p&gt;
&lt;p&gt;So far the local ML practice is just the beginning, in a real world production environment, ML projects typically follow a structured lifecycle and often deployed in scalable, cloud-based environments. Cloud Providers like AWS with Managed Kubernetes Services (EKS) provide orchestration, scaling, and fault tolerance to handle ML workloads.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Path for deploying MLOps workload in K8S&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Transition from local ML practice to a K8S-based deployment (&lt;strong&gt;This post&lt;/strong&gt;)&lt;/li&gt;
&lt;li&gt;Start with MLOps tools like &lt;strong&gt;Kubeflow&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Shift from local to cloud platforms (AWS) to deploy ML workload on EKS.&lt;/li&gt;
&lt;li&gt;Practice deploying models using REST APIs (local) and APT Gateway (AWS).&lt;/li&gt;
&lt;li&gt;Try local data engineering (ETL pipelines, data lakes, etc.), then move to AWS data services and solutions for ML workload.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Setting up Minikube with GPU on WSL Ubuntu&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>MLOps - Build a Knowledge Base with DeepSeek R1</title><link>https://zackblog.work/posts/mlops-build-a-knowledge-base-with-deepseek-r1/</link><pubDate>Sun, 19 Jan 2025 04:33:34 +0000</pubDate><guid>https://zackblog.work/posts/mlops-build-a-knowledge-base-with-deepseek-r1/</guid><description>&lt;p&gt;&amp;lsquo;Deep Seek is really hot at the moment!&lt;/p&gt;
&lt;p&gt;In this post, I want to build a local Knowledge Base using WSL, Docker, Ollama, Open WebUI and DeepSeek R1 7b.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Ollama&lt;/em&gt; is a desktop application designed to run and interact with large language models (LLMs) locally on a machine. It provides an easy interface for downloading, managing, and using various LLMs, ensuring privacy and local execution.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Open WebUI&lt;/em&gt; is a web-based user interface for interacting with AI models. It is often used in conjunction with locally hosted or remote LLMs, providing a customizable and user-friendly platform to input queries and manage model interactions.&lt;/p&gt;</description></item><item><title>MLOps - Explore ML tools</title><link>https://zackblog.work/posts/mlops-explore-ml-tools/</link><pubDate>Wed, 02 Oct 2024 04:32:58 +0000</pubDate><guid>https://zackblog.work/posts/mlops-explore-ml-tools/</guid><description>&lt;p&gt;In the last post &lt;a href="https://zackblog.work/posts/mlops-setup-a-home-machine-learning-lab/"&gt;MLOPS - Lab Setup&lt;/a&gt;, I was able to set the local ML lab environment, and run validation in Jupyter Notebook to test the CODA device and performance on my local PC.&lt;/p&gt;
&lt;p&gt;Although &lt;em&gt;Jupyter Notebooks&lt;/em&gt; can be user-friendly tools for ML practice, offering easy interaction and immediate feedback, which simplifies testing and debugging, it has limitations such as reproducibility issues, challenges in collaboration and version control, scalability concerns for larger projects, and a lack of automation for tasks like retraining.&lt;/p&gt;</description></item><item><title>MLOps - Setup a Home Machine Learning Lab</title><link>https://zackblog.work/posts/mlops-setup-a-home-machine-learning-lab/</link><pubDate>Tue, 01 Oct 2024 04:32:35 +0000</pubDate><guid>https://zackblog.work/posts/mlops-setup-a-home-machine-learning-lab/</guid><description>&lt;p&gt;Transitioning from DevOps to MLOps can be achieved by leveraging existing DevOps expertise by adding new layers specific to machine learning.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Key Differences:&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;code&gt;Model Lifecycle Management&lt;/code&gt;: MLOps handles model training, deployment, and retraining.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;code&gt;Data Versioning&lt;/code&gt;: Tools like DVC ensure dataset version control.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;code&gt;Experiment Tracking&lt;/code&gt;: MLflow and Weights &amp;amp; Biases track model training parameters and results.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;code&gt;Model Serving&lt;/code&gt;: Deploy models with TensorFlow Serving or TorchServe.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;code&gt;Model Drift&lt;/code&gt;: Monitor data changes over time to trigger retraining.&lt;/p&gt;</description></item></channel></rss>