A Kube-Scheduler Plugin for Tekton Capacity Reservation

Cover image credit: Photo by Hitesh Dewasi on Unsplash

A Kube-Scheduler Plugin for Tekton Capacity Reservation

An out-of-tree kube-scheduler plugin that reserves node capacity across a Tekton PipelineRun's full lifetime, closing a gap none of Tekton's own coschedule modes cover on their own.

Share
Table of Contents

Tekton exposes four coschedule settings: workspaces, pipelineruns, isolate-pipelinerun, and disabled. workspaces co-locates tasks within one PipelineRun that share a workspace. pipelineruns co-locates every task in that PipelineRun, regardless of workspace sharing. Neither decides whether a second PipelineRun may land on the node between one task’s pod exiting and the next one starting.

isolate-pipelinerun closes that gap by excluding every other PipelineRun from the node entirely. TEP-0135 requires operators to configure whether multiple PipelineRuns run concurrently on the same node. isolate-pipelinerun makes that policy binary: a node is reserved for one PipelineRun, or isolation is disabled. It can’t admit a second PipelineRun only when capacity remains.

I built a kube-scheduler plugin that reserves each workspace-sharing group’s peak CPU and memory need while that group is active, on the node it lands on. A second PipelineRun can share that node when there’s room and gets pushed elsewhere when there isn’t.

The placement problem is small enough to state before the implementation. PipelineRun A has a 200m task running and a later workspace-sharing task that needs 3800m. PipelineRun B needs 2000m. Current Pod requests make B look admissible on A’s node; A’s future locality requirement does not. The plugin makes that future CPU and memory requirement visible to scheduling.

What the plugin does

tkn-scheduler is a kube-scheduler instance that handles Pods carrying schedulerName: tkn-scheduler. It’s built on k8s.io/kubernetes/cmd/kube-scheduler/app, registering a WorkspaceGroup plugin at the PreFilter, Filter, and Reserve extension points via app.WithPlugin:

app.WithPlugin(schedplugin.Name, schedplugin.New(store, scheme)),

There’s no admission webhook or separate controller deployment for reservation creation. PreFilter looks up the pod’s Reservation Record (RR) by <namespace>/<pipelineRun> and creates it lazily and idempotently on the first qualifying pod. It joins tasks that share a workspace into a group. In the sequential task-a → task-b runs exercised here, each group’s CPU and memory are the maximum of its member tasks rather than their sum. That is valid because runAfter prevents those tasks from overlapping. The prototype does not yet derive a safe peak for arbitrary parallel tasks in the same workspace group.

Filter does the actual admission check. If the group already has a node, every other candidate is rejected outright. For an unassigned group, or for the pod’s own request once the group has a node:

// allocatable - scheduledRequests - foreignReservations - required >= 0
//
// scheduledRequests is the node's currently-allocated pod requests MINUS the requests of any
// other pods already on the node that belong to THIS SAME workspace group
func Evaluate(c CapacityCheck, foreignCPU, foreignMemory resource.Quantity) Result {
scheduledMinusSameGroupCPU := subtractClampToZero(c.NodeScheduledCPU, c.SameGroupPodsCPU)
scheduledMinusSameGroupMemory := subtractClampToZero(c.NodeScheduledMemory, c.SameGroupPodsMemory)
availableCPU := subtractClampToZero(c.NodeAllocatableCPU, scheduledMinusSameGroupCPU)
availableCPU = subtractClampToZero(availableCPU, foreignCPU)
availableMemory := subtractClampToZero(c.NodeAllocatableMemory, scheduledMinusSameGroupMemory)
availableMemory = subtractClampToZero(availableMemory, foreignMemory)
fits := availableCPU.Cmp(c.RequiredCPU) >= 0 && availableMemory.Cmp(c.RequiredMemory) >= 0
...
}

The current foreignCPU/foreignMemory value is the full reservation total from every other group. That prevents a group’s own pending pods from blocking each other, but it also means a foreign group’s running pod is counted once in scheduled requests and again in its full reservation. The intended value is each foreign group’s unrepresented reservation headroom: its reserved peak minus requests from that group’s pods already included in the node total. The same-group subtraction solves a different problem. Without it, a task pod already running on the node would be double-counted against a later task in the same group. A legitimate back-to-back schedule would then look like a capacity shortfall.

Reserve fires on the first pod for a group, atomically setting its node, its reservation state, and a label marking the record active. Later pods in an assigned group can reserve only that node. A scheduling cycle that finds the record already claimed by a different node returns a conflict, not a retry. If a bind fails after Reserve succeeds, Unreserve rolls the record back to unclaimed.

The auxiliary controller-runtime manager feeds terminal-pod events and PipelineRun updates to a release worker. It releases a group when all task pods are terminal or when the PipelineRun itself is done. It’s retry-aware, so an active retry keeps the reservation held after an earlier attempt of the same task fails. An optional resize worker shrinks a TaskRun’s pod to a configurable floor after the TaskRun reaches a terminal state. Its schedule is separate from the reservation’s own release.

This diagram separates the scheduler-framework path from the auxiliary release and resize workers:

flowchart TB
  subgraph K8s["Kubernetes API (kind cluster, v1.37.0)"]
    PRobj["PipelineRun"]
    TR["TaskRun"]
    RR["Reservation Record (CRD)"]
    Pod["TaskRun Pod<br/>schedulerName: tkn-scheduler"]
  end

  subgraph SchedProc["tkn-scheduler process"]
    KS["kube-scheduler command<br/>(app.NewSchedulerCommand)"]
    PF["WorkspaceGroup<br/>PreFilter"]
    F["WorkspaceGroup<br/>Filter"]
    R["WorkspaceGroup<br/>Reserve / Unreserve"]
    AUX["Auxiliary controller-runtime manager"]
    RS[("ReservationStore")]
    RW["Release worker"]
    TP["Optional terminal-pod resize worker"]
    KS --> PF --> F --> R
    AUX --> RS
    AUX --> RW
    AUX -. "--terminal-pod-resize" .-> TP
  end

  Pod --> KS
  PF <--> RR
  R <--> RR
  RR --> AUX
  PRobj --> AUX
  TR --> AUX
  Pod --> AUX
  RW <--> RR

Co-location holds

Two sequential tasks sharing a workspace, task-a and task-b, were partitioned into one group with a peak of cpu: 50m, memory: 64Mi. Both landed on tekton-scheduler-lab-worker2. The captured final state, after both tasks completed, showed the group’s reservation state as Released.

A reserved peak changes placement

PipelineRun A’s group included a future task-b sized at 3800m CPU. Its active task-a pod requested only 200m. The group nevertheless held its 3800m peak from the moment it landed on a node. PipelineRun B then requested 2000m CPU. Scheduled requests alone would have admitted B to A’s node. The reservation left only 800m available, so the plugin rejected worker2:

"WorkspaceGroup filter rejected node" node="tekton-scheduler-lab-worker2"
pod="claim2-pr-b-v2-task-a-pod" pipelineRun="claim2-pr-b-v2" group="group-0"
availableCPU="800m" availableMemory="6714420Ki" requiredCPU="2" requiredMemory="128Mi"
foreignCPU="3800m" foreignMemory="128Mi"

foreignCPU=3800m is A’s group’s reserved size, not its running pod’s actual 200m request. That’s the plugin’s own accounting doing the rejecting, not the stock scheduler noticing an already-full node. B scheduled onto the other worker immediately after. The current implementation subtracts the full foreign reservation in this path, so the logged 800m availability conservatively counts A’s active 200m pod twice. Correct production accounting would subtract only the foreign reservation headroom, leaving 1000m in this example; B still does not fit.

Release waits for a retry

The co-location run showed that a completed PipelineRun releases its reservation. A separate task with retries: 1 produced two pods sharing the same task label: one Failed, the retry still Running. The reservation remained held. GroupTerminal groups pods by task label rather than trusting a single pod, because Tekton has no separate attempt-number label:

func GroupTerminal(tasks []string, allPods []*corev1.Pod) bool {
podsByTask := map[string][]*corev1.Pod{}
for _, p := range allPods {
task := p.Labels[pipelineTaskLabelKey]
if task == "" {
continue
}
podsByTask[task] = append(podsByTask[task], p)
}
for _, task := range tasks {
pods := podsByTask[task]
if len(pods) == 0 {
return false
}
for _, p := range pods {
if !isPodTerminal(p) {
return false
}
}
}
return true
}

Once the retry itself succeeded, the captured state showed the reservation Released.

Pod resizing does not change the reservation

The resize worker watches for terminal TaskRuns, not terminal pods. A TaskRun can finish while its pod is still draining sidecars, such as Docker-in-Docker. In that interval, the pod can retain the task’s original requests even though its useful work is over. The worker patches those requests down to a configurable floor.

The reservation follows a separate rule. In the captured run, the worker changed task-a’s container requests from 500m/256Mi to 5m/20Mi, confirmed by both the pod field and its log line. Across the captured unresized and resized snapshots, the Reservation Record remained at 500m/256Mi and Reserved. The node’s Allocated resources table had already stopped counting the terminal pod before the resize. That table showed no marginal capacity change. Resizing a teardown pod doesn’t change the group’s reserved peak.

Limits before this becomes a production scheduler

This prototype is not a production-ready reservation system. Its demonstrated group contains sequential tasks. The max(task requests) sizing rule is unsafe when grouped tasks can be runnable concurrently: two 3-CPU tasks require up to 6 CPU, not 3 CPU. A production version needs a concurrency-aware peak calculation, or a deliberately more conservative sum, before it can schedule arbitrary Tekton DAGs and finally tasks.

Reservations affect only Pods scheduled through tkn-scheduler. The default scheduler does not read this CRD, so it can place an unrelated workload into capacity that this plugin considers reserved. A production deployment needs dedicated or tainted CI nodes, or a reservation policy enforced for every workload that can consume these nodes.

The retry run proves that an already-observed non-terminal retry prevents release. It does not cover the gap between a failed attempt and creation or observation of its retry Pod. During that gap, GroupTerminal can see only terminal Pods and release the reservation early. A production release decision needs retry state from the TaskRun or PipelineRun to be authoritative, rather than inferring exhaustion only from the Pods currently visible.

Resource scope is CPU and memory. The Pod-level accounting handles regular containers and init-container semantics, including restartable sidecars, but it is not a complete reproduction of kube-scheduler’s effective request calculation. Pod overhead, ephemeral storage, hugepages, extended resources, and Pod-count limits are outside this prototype’s reservation model.

Reserve and Unreserve cover a failure after this plugin reserves a node but before binding completes. The experiment did not exercise scheduler restart, multiple scheduler replicas, node loss or drain, eviction, preemption, PipelineRun deletion races, conflicting state updates, autoscaler behavior, or a scheduler/Kubernetes version-skew and upgrade policy. Those are design requirements for a production deployment, not outcomes established here.

The current rollback rule also needs stronger ownership semantics before parallel group members are supported. A later scheduling cycle can observe a node claim made by an earlier pod while that earlier pod is still binding. If the earlier bind then fails, its Unreserve clears the shared group claim even if the later pod has progressed using it. The demonstrated sequential workload does not exercise that interleaving. A production design needs a claim generation, per-pod holder set, or equivalent compare-and-swap rule so a rollback can undo only the state transition it owns.

Lazy creation keeps the prototype small, but it puts a reservation lookup and, for the first qualifying pod, a CRD create on PreFilter’s scheduling path. That makes API-server latency and failure part of scheduling latency. A production design could create and size the reservation from PipelineRun state before pods reach the scheduler, leaving the scheduler to read, filter, and claim an existing record. This prototype chose the lazy path and did not measure that trade-off.

References

  1. Tekton TEP-0135: Coscheduling PipelineRun pods
  2. scheduler-plugins: Coscheduling
  3. Kubernetes Scheduling Framework

Similar Articles