conference-paper

Cost-Effective Scaling of Machine Learning Model using Serverless Architecture

Research footprint

At a glance

Citations
0
References
14
Comments
0
Paper overview

Abstract

Machine learning algorithms are resource intensive processes that require large computing resources. The increasing demand for these resources has highlighted the need for a cost-effective solution in the domain of cloud computing. Traditional systems require manual monitoring and updating of servers and have fixed resources allocated to them that leads to higher costs and underutilization of resources. Serverless architectures like AWS Lambda offer an efficient way for the dynamic scaling of machine learning model deployments without the need for a dedicated team managing servers. Servers are scaled automatically based on traffic and load that results in efficient resource utilization. The study demonstrates how serverless architecture can be utilized for cost-efficient scaling of machine learning deployments, especially for the MNIST dataset. The goal is to compare the performance and cost-efficiency of traditional deployment techniques with serverless architecture. By analyzing key performance metrics like latency, cost per invocation, and resource utilization, the study aims to highlight the advantages of a serverless architecture in handling dynamic user loads. Serverless functions have been found to reduce costs significantly by scaling models layers independently while maintaining the performance of the model. The research also identifies the scenarios where serverless architecture outperforms traditional deployment methods in terms of cost and flexibility.

Record transparency

Publication details

DOI
10.1109/icaiss61471.2025.11041847
OpenAlex
W4411600879
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.