Scalable AI Inference: Empirical Optimization of BentoML-Based Model Serving Under Realistic Workloads
arXiv·low signal
Systematic performance analysis of BentoML-based inference serving across three realistic workload scenarios using RoBERTa. While most research focuses on model design, this study addresses the deployment gap—benchmarking throughput, latency, and scaling behavior under production conditions. Developed in collaboration with graphworks.ai for practical serving optimization.