# A Kube-Scheduler Plugin for Tekton Capacity Reservation

Tekton exposes four `coschedule` settings: `workspaces`, `pipelineruns`, `isolate-pipelinerun`, and `disabled`. `workspaces` co-locates tasks within one PipelineRun that share a workspace. `pipelineruns` co-locates every task in that PipelineRun, regardless of workspace sharing. Neither decides whether a second PipelineRun may land on the node between one task's pod exiting and the next one starting.

`isolate-pipelinerun` closes that gap by excluding every other PipelineRun from the node entirely. [TEP-0135](https://github.com/tektoncd/community/blob/main/teps/0135-coscheduling-pipelinerun-pods.md) requires operators to configure whether multiple PipelineRuns run concurrently on the same node. `isolate-pipelinerun` makes that policy binary: a node is reserved for one PipelineRun, or isolation is disabled. It can't admit a second PipelineRun only when capacity remains.

I built a kube-scheduler plugin that reserves each workspace-sharing group's peak CPU and memory need while that group is active, on the node it lands on. A second PipelineRun can share that node when there's room and gets pushed elsewhere when there isn't.

The placement problem is small enough to state before the implementation. PipelineRun A has a `200m` task running and a later workspace-sharing task that needs `3800m`. PipelineRun B needs `2000m`. Current Pod requests make B look admissible on A's node; A's future locality requirement does not. The plugin makes that future CPU and memory requirement visible to scheduling.

## What the plugin does

`tkn-scheduler` is a kube-scheduler instance that handles Pods carrying `schedulerName: tkn-scheduler`. It's built on `k8s.io/kubernetes/cmd/kube-scheduler/app`, registering a `WorkspaceGroup` plugin at the PreFilter, Filter, and Reserve extension points via `app.WithPlugin`:

```go
app.WithPlugin(schedplugin.Name, schedplugin.New(store, scheme)),
```

There's no admission webhook or separate controller deployment for reservation creation. PreFilter looks up the pod's Reservation Record (RR) by `<namespace>/<pipelineRun>` and creates it lazily and idempotently on the first qualifying pod. It joins tasks that share a workspace into a group. In the sequential `task-a → task-b` runs exercised here, each group's CPU and memory are the maximum of its member tasks rather than their sum. That is valid because `runAfter` prevents those tasks from overlapping. The prototype does not yet derive a safe peak for arbitrary parallel tasks in the same workspace group.

Filter does the actual admission check. If the group already has a node, every other candidate is rejected outright. For an unassigned group, or for the pod's own request once the group has a node:

```go
// allocatable - scheduledRequests - foreignReservations - required >= 0
//
// scheduledRequests is the node's currently-allocated pod requests MINUS the requests of any
// other pods already on the node that belong to THIS SAME workspace group
func Evaluate(c CapacityCheck, foreignCPU, foreignMemory resource.Quantity) Result {
	scheduledMinusSameGroupCPU := subtractClampToZero(c.NodeScheduledCPU, c.SameGroupPodsCPU)
	scheduledMinusSameGroupMemory := subtractClampToZero(c.NodeScheduledMemory, c.SameGroupPodsMemory)

	availableCPU := subtractClampToZero(c.NodeAllocatableCPU, scheduledMinusSameGroupCPU)
	availableCPU = subtractClampToZero(availableCPU, foreignCPU)

	availableMemory := subtractClampToZero(c.NodeAllocatableMemory, scheduledMinusSameGroupMemory)
	availableMemory = subtractClampToZero(availableMemory, foreignMemory)

	fits := availableCPU.Cmp(c.RequiredCPU) >= 0 && availableMemory.Cmp(c.RequiredMemory) >= 0
	...
}
```

The current `foreignCPU`/`foreignMemory` value is the full reservation total from every other group. That prevents a group's own pending pods from blocking each other, but it also means a foreign group's running pod is counted once in scheduled requests and again in its full reservation. The intended value is each foreign group's unrepresented reservation headroom: its reserved peak minus requests from that group's pods already included in the node total. The same-group subtraction solves a different problem. Without it, a task pod already running on the node would be double-counted against a later task in the same group. A legitimate back-to-back schedule would then look like a capacity shortfall.

Reserve fires on the first pod for a group, atomically setting its node, its reservation state, and a label marking the record active. Later pods in an assigned group can reserve only that node. A scheduling cycle that finds the record already claimed by a *different* node returns a conflict, not a retry. If a bind fails after Reserve succeeds, Unreserve rolls the record back to unclaimed.

The auxiliary controller-runtime manager feeds terminal-pod events and PipelineRun updates to a release worker. It releases a group when all task pods are terminal or when the PipelineRun itself is done. It's retry-aware, so an active retry keeps the reservation held after an earlier attempt of the same task fails. An optional resize worker shrinks a TaskRun's pod to a configurable floor after the TaskRun reaches a terminal state. Its schedule is separate from the reservation's own release.

This diagram separates the scheduler-framework path from the auxiliary release and resize workers:

```mermaid
flowchart TB
  subgraph K8s["Kubernetes API (kind cluster, v1.37.0)"]
    PRobj["PipelineRun"]
    TR["TaskRun"]
    RR["Reservation Record (CRD)"]
    Pod["TaskRun Pod<br/>schedulerName: tkn-scheduler"]
  end

  subgraph SchedProc["tkn-scheduler process"]
    KS["kube-scheduler command<br/>(app.NewSchedulerCommand)"]
    PF["WorkspaceGroup<br/>PreFilter"]
    F["WorkspaceGroup<br/>Filter"]
    R["WorkspaceGroup<br/>Reserve / Unreserve"]
    AUX["Auxiliary controller-runtime manager"]
    RS[("ReservationStore")]
    RW["Release worker"]
    TP["Optional terminal-pod resize worker"]
    KS --> PF --> F --> R
    AUX --> RS
    AUX --> RW
    AUX -. "--terminal-pod-resize" .-> TP
  end

  Pod --> KS
  PF <--> RR
  R <--> RR
  RR --> AUX
  PRobj --> AUX
  TR --> AUX
  Pod --> AUX
  RW <--> RR
```

## Co-location holds

Two sequential tasks sharing a workspace, `task-a` and `task-b`, were partitioned into one group with a peak of `cpu: 50m`, `memory: 64Mi`. Both landed on `tekton-scheduler-lab-worker2`. The captured final state, after both tasks completed, showed the group's reservation state as `Released`.

## A reserved peak changes placement

PipelineRun A's group included a future `task-b` sized at `3800m` CPU. Its active `task-a` pod requested only `200m`. The group nevertheless held its `3800m` peak from the moment it landed on a node. PipelineRun B then requested `2000m` CPU. Scheduled requests alone would have admitted B to A's node. The reservation left only `800m` available, so the plugin rejected worker2:
```
"WorkspaceGroup filter rejected node" node="tekton-scheduler-lab-worker2"
pod="claim2-pr-b-v2-task-a-pod" pipelineRun="claim2-pr-b-v2" group="group-0"
availableCPU="800m" availableMemory="6714420Ki" requiredCPU="2" requiredMemory="128Mi"
foreignCPU="3800m" foreignMemory="128Mi"
```

`foreignCPU=3800m` is A's group's reserved size, not its running pod's actual 200m request. That's the plugin's own accounting doing the rejecting, not the stock scheduler noticing an already-full node. B scheduled onto the other worker immediately after. The current implementation subtracts the full foreign reservation in this path, so the logged `800m` availability conservatively counts A's active `200m` pod twice. Correct production accounting would subtract only the foreign reservation headroom, leaving `1000m` in this example; B still does not fit.

## Release waits for a retry

The co-location run showed that a completed PipelineRun releases its reservation. A separate task with `retries: 1` produced two pods sharing the same task label: one `Failed`, the retry still `Running`. The reservation remained held. `GroupTerminal` groups pods by task label rather than trusting a single pod, because Tekton has no separate attempt-number label:

```go
func GroupTerminal(tasks []string, allPods []*corev1.Pod) bool {
	podsByTask := map[string][]*corev1.Pod{}
	for _, p := range allPods {
		task := p.Labels[pipelineTaskLabelKey]
		if task == "" {
			continue
		}
		podsByTask[task] = append(podsByTask[task], p)
	}

	for _, task := range tasks {
		pods := podsByTask[task]
		if len(pods) == 0 {
			return false
		}
		for _, p := range pods {
			if !isPodTerminal(p) {
				return false
			}
		}
	}
	return true
}
```

Once the retry itself succeeded, the captured state showed the reservation `Released`.

## Pod resizing does not change the reservation

The resize worker watches for terminal TaskRuns, not terminal pods. A TaskRun can finish while its pod is still draining sidecars, such as Docker-in-Docker. In that interval, the pod can retain the task's original requests even though its useful work is over. The worker patches those requests down to a configurable floor.

The reservation follows a separate rule. In the captured run, the worker changed `task-a`'s container requests from `500m`/`256Mi` to `5m`/`20Mi`, confirmed by both the pod field and its log line. Across the captured unresized and resized snapshots, the Reservation Record remained at `500m`/`256Mi` and `Reserved`. The node's `Allocated resources` table had already stopped counting the terminal pod before the resize. That table showed no marginal capacity change. Resizing a teardown pod doesn't change the group's reserved peak.

## Limits before this becomes a production scheduler

This prototype is not a production-ready reservation system. Its demonstrated group contains sequential tasks. The `max(task requests)` sizing rule is unsafe when grouped tasks can be runnable concurrently: two 3-CPU tasks require up to 6 CPU, not 3 CPU. A production version needs a concurrency-aware peak calculation, or a deliberately more conservative sum, before it can schedule arbitrary Tekton DAGs and `finally` tasks.

Reservations affect only Pods scheduled through `tkn-scheduler`. The default scheduler does not read this CRD, so it can place an unrelated workload into capacity that this plugin considers reserved. A production deployment needs dedicated or tainted CI nodes, or a reservation policy enforced for every workload that can consume these nodes.

The retry run proves that an already-observed non-terminal retry prevents release. It does not cover the gap between a failed attempt and creation or observation of its retry Pod. During that gap, `GroupTerminal` can see only terminal Pods and release the reservation early. A production release decision needs retry state from the TaskRun or PipelineRun to be authoritative, rather than inferring exhaustion only from the Pods currently visible.

Resource scope is CPU and memory. The Pod-level accounting handles regular containers and init-container semantics, including restartable sidecars, but it is not a complete reproduction of kube-scheduler's effective request calculation. Pod overhead, ephemeral storage, hugepages, extended resources, and Pod-count limits are outside this prototype's reservation model.

Reserve and Unreserve cover a failure after this plugin reserves a node but before binding completes. The experiment did not exercise scheduler restart, multiple scheduler replicas, node loss or drain, eviction, preemption, PipelineRun deletion races, conflicting state updates, autoscaler behavior, or a scheduler/Kubernetes version-skew and upgrade policy. Those are design requirements for a production deployment, not outcomes established here.

The current rollback rule also needs stronger ownership semantics before parallel group members are supported. A later scheduling cycle can observe a node claim made by an earlier pod while that earlier pod is still binding. If the earlier bind then fails, its Unreserve clears the shared group claim even if the later pod has progressed using it. The demonstrated sequential workload does not exercise that interleaving. A production design needs a claim generation, per-pod holder set, or equivalent compare-and-swap rule so a rollback can undo only the state transition it owns.

Lazy creation keeps the prototype small, but it puts a reservation lookup and, for the first qualifying pod, a CRD create on PreFilter's scheduling path. That makes API-server latency and failure part of scheduling latency. A production design could create and size the reservation from PipelineRun state before pods reach the scheduler, leaving the scheduler to read, filter, and claim an existing record. This prototype chose the lazy path and did not measure that trade-off.
