The best candidates would be top 1% at multiple parts of the inference stack.
work on PD disaggregation research
Morph builds the inference infrastructure behind the fastest open models. Our stack spans kernels, model serving, routing, autoscaling, and capacity. We are hiring a performance engineer to make the entire system faster, cheaper, and more reliable.
What you’ll do
Find the gap between theoretical hardware performance and production performance
Trace latency and throughput regressions from the API layer down to individual kernels
Optimize batching, scheduling, routing, quantization, and distributed execution
Build benchmarks and observability that make bottlenecks obvious
Validate that every optimization preserves model quality and correctness
Stack-rank opportunities and ship the highest-impact fixes yourself
... continue reading