GPU Scheduling Is a Business Policy Problem: Designing NVIDIA Run:ai Quotas, Fairness, and Preemption
Introduction A GPU cluster does not know which product launch is contractually committed, which research experiment can wait until tomorrow, or which inference endpoint supports a revenue-producing application. Kubernetes sees pods, resource requests, labels, and scheduling constraints. The business sees customers, deadlines, budgets, risk, and service commitments. That gap is where many shared GPU platforms […]










