Provision capacity for Kueue batch jobs with the cluster autoscaler on AKS

In this article, you pair Kueue admission control with the Azure Kubernetes Service (AKS) cluster autoscaler so that batch jobs get the nodes they need before they run. Create jobs with suspend: true and a queue label. Kueue creates a Kubernetes ProvisioningRequest on their behalf. The cluster autoscaler atomically scales up the node pool to satisfy it, and only then does Kueue admit the job.

Component Role
Kueue Queues the job and gates admission against quota
ProvisioningRequest AdmissionCheck Asks the cluster autoscaler for capacity before admission
Cluster autoscaler Atomically scales up the node pool to satisfy the request

Important

Open-source software is mentioned throughout AKS documentation and samples. Software that you deploy is excluded from AKS service-level agreements, limited warranty, and Azure support. As you use open-source technology alongside AKS, consult the support options available from the respective communities and project maintainers to develop a plan.

Microsoft takes responsibility for building the open-source packages that we deploy on AKS. That responsibility includes having complete ownership of the build, scan, sign, validate, and hotfix process, along with control over the binaries in container images. For more information, see Vulnerability management for AKS and AKS support coverage.

Prerequisites

  • This article assumes basic knowledge of Kueue. For more information, see Learn about Kueue for batch scheduling.
  • An AKS cluster with a node pool that has the cluster autoscaler enabled. The examples use a pool named scalepool with --min-count 1 --max-count 5.
  • The Azure CLI version 2.70 or later. Run az --version to find the version. If you need to install or upgrade, see Install Azure CLI.
  • kubectl connected to the cluster (az aks get-credentials).
  • Helm 3.12 or later to install the Kueue controller.

You can find the example manifests used in this article in the kueue-and-ray-on-aks directory of the Azure/AKS repository. Clone the repository so you can run them:

git clone https://github.com/Azure/AKS.git
cd AKS/examples/kueue-and-ray-on-aks

Install Kueue

Install the Kueue controller by using Helm:

helm install kueue oci://k8sgcr.azk8s.cn/kueue/charts/kueue \
  --version 0.17.1 \
  --namespace kueue-system \
  --create-namespace \
  --wait

Verify the controller is running:

kubectl -n kueue-system get pods

Create the queue configuration

Apply the autoscale queue configuration. The file creates its own cas-kueue-demo namespace along with the queue objects, so you don't need a separate namespace step:

kubectl apply -f 2-kueue-queues/manifests/40-autoscale-queue.yaml

The ResourceFlavor targets nodes labeled agentpool=scalepool, which AKS applies automatically to that node pool:

apiVersion: kueue.x-k8s.io/v1beta2
kind: ResourceFlavor
metadata:
  name: scalepool
spec:
  nodeLabels:
    agentpool: scalepool

Verify:

kubectl get resourceflavor scalepool

Configure the provisioning gate

The same file connects the provisioning gate with three more objects:

  • ProvisioningRequestConfig cas-provreq-config selects the best-effort-atomic-scale-up.autoscaling.x-k8s.io provisioning class, so the autoscaler adds the requested capacity as a single atomic increase.
  • AdmissionCheck cas-provisioning uses the kueue.x-k8s.io/provisioning-request controller and points at that config.
  • ClusterQueue cas-cluster-queue gates admission on cas-provisioning through admissionChecksStrategy, so every workload it admits goes through the provisioning gate first.

Verify the objects:

kubectl get admissioncheck cas-provisioning
kubectl get clusterqueue cas-cluster-queue
kubectl -n cas-kueue-demo get localqueue cas-local-queue

Expected output:

NAME               AGE
cas-provisioning   1m

NAME                COHORT   PENDING WORKLOADS
cas-cluster-queue            0

NAME              CLUSTERQUEUE        PENDING WORKLOADS   ADMITTED WORKLOADS
cas-local-queue   cas-cluster-queue   0                   0

Submit a workload

Submit a suspended Job routed through the queue. It requests three pods that can't all fit on the single starting node, which forces a scale-up:

kubectl apply -f 3-workloads/cas-batch-job/manifests/job.yaml

The Job carries the kueue.x-k8s.io/queue-name: cas-local-queue label and suspend: true, so Kueue takes over admission:

apiVersion: batch/v1
kind: Job
metadata:
  name: kueue-cas-job
  namespace: cas-kueue-demo
  labels:
    kueue.x-k8s.io/queue-name: cas-local-queue
spec:
  parallelism: 3
  completions: 3
  suspend: true
  template:
    spec:
      nodeSelector:
        agentpool: scalepool
      containers:
        - name: worker
          image: mcr.azk8s.cn/azurelinux/busybox:1.36
          command: ["sh", "-c", "echo running on $(hostname); sleep 30"]
          resources:
            requests:
              cpu: "1800m"
              memory: "256Mi"
      restartPolicy: Never

Watch the provisioning flow

Kueue creates a Workload, then a ProvisioningRequest. The cluster autoscaler satisfies the request and the pool grows:

# Kueue creates a Workload and a ProvisioningRequest
kubectl -n cas-kueue-demo get workloads
kubectl -n cas-kueue-demo get provisioningrequest

# The autoscaler marks the request Provisioned and adds nodes
kubectl get nodes -l agentpool=scalepool -w

# The Job runs to completion once nodes are Ready
kubectl -n cas-kueue-demo get job kueue-cas-job -w

Expected end state:

NAME            STATUS     COMPLETIONS   DURATION   AGE
kueue-cas-job   Complete   3/3           34s        2m

Troubleshooting

Symptom Cause Fix
Job stays suspended, no ProvisioningRequest Wrong queue-name label Verify kueue.x-k8s.io/queue-name: cas-local-queue matches the LocalQueue name.
ProvisioningRequest created but status.conditions empty Autoscaler hasn't processed the request yet Confirm the node pool has the cluster autoscaler enabled, then recheck after a minute
Provisioned=False with reason CapacityIsNotFound The pool can't reach the requested size, usually because existing workloads occupy it or the request exceeds --max-count Free capacity or raise --max-count. The autoscaler keeps retrying, and the job stays suspended rather than partially scheduling.
Pods Pending after admission Node label mismatch Confirm scalepool nodes carry the agentpool=scalepool label with kubectl get nodes --show-labels
Kueue controller not running Helm release issue Check with kubectl -n kueue-system logs deploy/kueue-controller-manager

For the full manifests, see the kueue-and-ray-on-aks directory in the Azure/AKS repository.

Clean up resources

Remove the workload and queue configuration:

kubectl delete -f 3-workloads/cas-batch-job/manifests/job.yaml
kubectl delete -f 2-kueue-queues/manifests/40-autoscale-queue.yaml

Next steps

To learn more about the components used in this article, see: