Pod-level resources for Tekton TaskRuns

Cover image credit: Photo by Sunira Moses on Unsplash

Pod-level resources for Tekton TaskRuns

A mutating webhook gives a Tekton build step and Docker-in-Docker sidecar one Pod-level resource budget, then verifies the limits in node cgroups.

Share
Table of Contents

Tekton has TaskRun-level computeResources, but that setting applies to Steps and explicitly excludes sidecars. Sidecars retain separate resource configuration. Tekton documents the split.

That split matters for Docker-in-Docker (dind) TaskRuns. The step container submits docker build, but the dind sidecar’s Docker daemon performs it. Because Tekton assigns their limits separately, the daemon can’t use CPU or memory left idle by the waiting step.

Kubernetes supports Pod-level resources. A Pod can set spec.resources.requests and spec.resources.limits, allowing containers without their own limits to draw from the Pod’s shared budget. Pod-level resources are beta and enabled by default from Kubernetes 1.34. The missing piece here was a way to put those fields on the Pod Tekton generated.

I used a mutating admission webhook for that step. It receives the Pod before Kubernetes stores it, adds the Pod-level resource values, and clears the individual limits. The generated Pod then reaches the kubelet with one budget instead of two ceilings.

The TaskRun input

The Task referenced by this TaskRun defines one build step and one dind sidecar:

steps:
- name: build-context
computeResources:
requests: { cpu: 100m, memory: 128Mi }
limits: { cpu: 250m, memory: 256Mi }
sidecars:
- name: dind
computeResources:
requests: { cpu: 500m, memory: 512Mi }
limits: { cpu: 2, memory: 2Gi }

The build step writes a Dockerfile, waits for dind to become ready, then sends it docker build. The daemon in the sidecar pulls the image and executes the Dockerfile commands. Once the build starts, the step mostly waits. With separate limits, dind would remain capped at 2 CPUs even while the waiting step used little of its own 250m CPU ceiling.

The desired Pod budget is the sum of the two containers:

requests: 100m + 500m = 600m CPU
128Mi + 512Mi = 640Mi memory
limits: 250m + 2 CPU = 2250m CPU
256Mi + 2Gi = 2304Mi memory

The webhook leaves the container requests unchanged and sets the Pod-level request to their computed total. It pools only the container limits, clearing the individual ceilings while keeping the original per-container requests present. For this Pod, the Pod-level request is the same 600m CPU and 640Mi that the running-container requests summed to. Tekton’s injected prepare and place-scripts init containers had empty resource blocks in this run, so the scheduling footprint did not change.

The webhook

An admission webhook is an HTTP service the API server calls before it stores a proposed Kubernetes object. This mutating webhook returns a JSON Patch that changes the generated Pod before Kubernetes schedules it.

Tekton asks Kubernetes to create the generated Pod. The API server sends that Pod in an AdmissionReview, Kubernetes’s request-and-response envelope for admission webhooks. The webhook returns the resource patch, and Kubernetes applies it before persisting the changed Pod.

sequenceDiagram
    participant T as Tekton
    participant A as Kubernetes API server
    participant W as mutating webhook
    participant S as scheduler
    participant K as kubelet

    T->>A: create generated Pod
    A->>W: AdmissionReview with proposed Pod
    W->>W: compute Pod-level resources
    W-->>A: JSON Patch
    A->>A: apply patch and store Pod
    S->>A: assign a node
    K->>A: observe assigned Pod
    K->>K: start step and dind

The handler sums requests and limits for containers that run together. It keeps each container request, writes the total to spec.resources.requests, and creates one limit-clearing operation for each container. For the two containers in this TaskRun, the handler constructs three patch operations:

add /spec/resources
replace /spec/containers/0/resources/limits {}
replace /spec/containers/1/resources/limits {}

For this TaskRun, the add operation creates this Pod-level value:

resources:
requests:
cpu: 600m
memory: 640Mi
limits:
cpu: 2250m
memory: 2304Mi

The two replace operations clear only resources.limits. They do not clear the whole resource object, so each container keeps its original request.

The generic webhook also recognizes native Kubernetes sidecars: restartable init containers with restartPolicy: Always. It includes those in the running-container total and clears their limits. The dind sidecar in this TaskRun used Tekton’s regular sidecar representation, not that native-sidecar path. Mixed native-sidecar and regular-init layouts need separate validation because a native sidecar can remain running while a later regular init container executes.

The resulting Pod had 600m CPU and 640MiB of requests below its 2250m CPU and 2304MiB limits. Because its Pod-level requests and limits were not equal, it did not meet the Guaranteed criteria and Kubernetes classified it as Burstable.

Cgroup verification

The step’s Dockerfile installed OpenSSL and ran 600 iterations of a 64MB dd and sha256sum pipeline. That kept dind busy long enough to inspect the node before the build completed.

I inspected the Linux cgroup v2 controls, the kernel files that enforce CPU and memory limits, on the kind test node. The Pod has a parent cgroup, with one child, or leaf, cgroup for each running container:

pod slice cpu.max: 225000 100000 memory.max: 2415919104
sidecar-dind leaf cpu.max: 225000 100000 memory.max: 2415919104
step-build-context leaf cpu.max: 225000 100000 memory.max: 2415919104

225000 100000 allows 225ms of CPU time per 100ms period, equivalent to 2.25 CPUs. 2415919104 bytes is 2304MiB. The step leaf no longer showed its original 250m CPU or 256MiB memory limit. The dind leaf no longer showed its original 2-CPU or 2GiB limit. All three cgroups showed the Pod-level limit.

In this kind test cluster, which uses containerd as its runtime, the Pod cgroup and both running-container leaves exposed the same Pod-level values after the webhook cleared the container limits.

The Pod completed with Succeeded, and the build log recorded Successfully tagged demo:local.

Tekton supplied separate container values but could not express one budget shared by the step and sidecar. The webhook inserted that budget before Kubernetes stored the generated Pod. The cgroup reads show the resulting 2.25-CPU and 2304MiB limit on the Pod and both container leaves.

Closing

Tekton supplies separate limits for the build step and dind sidecar. The webhook turns them into one Pod-level budget before scheduling, and the cgroup capture shows that Kubernetes carries that limit through to both running containers. That is the mechanism this experiment establishes.

It does not establish a capacity benefit. The prototype derives its Pod-level limit by summing every concurrently running container’s configured limit, which can be larger than a build needs at any one moment. Demonstrating a benefit requires a baseline workload and a contention test.

The next design step is to make that budget a workload choice rather than a sum of container defaults. A resource-class label on a PipelineRun or TaskRun could select extra-small, small, medium, large, or extra-large; Tekton propagates the label to the generated Pod, and the webhook can apply the corresponding Pod limits and container requests before the Pod is scheduled.

References

  1. Compute Resources in Tekton
  2. Pod-level Resources
  3. Tekton Labels and Annotations
  4. Sidecar Containers

Similar Articles