Skip to content
Tech News
← Back to articles

Finding zombies in our systems: A real-world story of CPU bottlenecks

read original more articles
Why This Matters

This story highlights the importance of performance profiling and system optimization in large-scale machine learning workflows. Identifying CPU bottlenecks and network issues at Pinterest underscores how critical infrastructure tuning is for maintaining reliable, efficient AI training processes, directly impacting industry standards and consumer experiences.

Key Takeaways

Finding zombies in our systems: A real-world story of CPU bottlenecks Pinterest Engineering 13 min read · Apr 15, 2026 -- 4 Listen Share

Vaibhav Shankar; Staff Software Engineer | Raymond Lee; Staff Software Engineer | Chia-Wei Chen; Staff Software Engineer | Shunyao Li; Sr. Software Engineer | Yi Li; Staff Software Engineer | Ambud Sharma; Principal Engineer | Saurabh Vishwas Joshi; Principal Engineer | Charles-A. Francisco; Senior Engineer | Karthik Anantha Padmanabhan; Director, Engineering | David Westbrook; Sr. Manager, Engineering

One day in early 2025, the Kubernetes platform team at Pinterest (PinCompute) got a ping from our partners on the ML platform team. Their Ray-based training jobs , which often take hours of computation on expensive GPU hardware, were crashing. Not every time, but often enough that it was becoming noticeable. Their logs indicated that their distributed training jobs were seeing intermittent loss of network connectivity, and that ultimately caused their jobs to crash. Their ask was simple:

Why is this happening? Can you please make it stop?

What started there led to a more than three-month-long investigation and a great lesson in profiling performance bottlenecks. Read on to learn from our fun story about CPU bottlenecks, AWS network drivers, and yes, how we discovered Zombies in our system!

Background: Ray at Pinterest

At Pinterest, Ray has risen as the backbone of our next-gen ML training and inference. Over the past few years, it has enabled us to scale systems, accelerate experimentation, and significantly boost the performance of models powering our diverse ML workloads.

We have previously shared deep dives on our progress, including: Ray Infrastructure (provisioning ray cluster on in-house K8s clusters at scale [blog]), Batch Inference with Ray (scaling to hundreds of nodes [blog][talk]), Ray for Training (distributed dataloaders and throughput optimization [talk]), and Last-Mile Data Processing (reducing experimentation cycles [blog 1][blog 2]).

Today, we run more than half of the offline ML workload company-wide on Ray, provisioning tens of thousands of Ray clusters per month, a feat made possible only by a robust Kubernetes environment.

Network Model & Challenges

... continue reading