Chuyển đến nội dung chính

LESSON 18: JOBS AND CRONJOBS

Batch processing with Jobs (single, parallel, indexed, work queue), CronJobs with timezone support (GA K8s 1.27). JobSet (CNCF project) for a group of dependent Jobs — ideal for AI/ML training pipelines.

🔒 DevSecOps — Lesson 18 LESSON 18: JOBS AND CRONJOBS

KUBERNETES: FROM BASIC TO ADVANCED

Module 5: Workload Management__HTMLTAG_60___

xdev.asia

Jobs and CronJobs in Kubernetes__HTMLTAG_66___

In Kubernetes, Deployments and StatefulSets are designed for continuously running workloads — they always try to maintain a certain number of Pods. But many real-world tasks don't need to run forever: process a batch of data, run a database migration, train an ML model, or send mass emails. This is where Jobs and CronJobs come into play.

1. What are jobs? Batch Workloads and Run-to-Completion

A Job in Kubernetes creates one or more Pods with the goal of completing a specific task. Unlike Deployment, Job tracks the number of successful completions — when enough Pods are completed, the Job is considered done.

Important Jobs Features:

  • Run-to-completion: Pod finished running and exited with code 0 meaning success
  • Automatic Retry: If Pod fails, Job automatically creates a new Pod according to backoffLimit
  • Tracking completions: Job knows how many have been completed out of the total needed
  • Parallelism: Multiple Pods can run in parallel to increase throughput

Simple Job example — calculate Pi:

apiVersion: batch/v1
kind: Job
metadata:
  name: pi-calculator
  namespace: default
spec:
  template:
    spec:
      containers:
      - name: pi
        image: perl:5.34
        command: ["perl", "-Mbignum=bpi", "-wle", "print bpi(2000)"]
        resources:
          requests:
            cpu: "250m"
            memory: "64Mi"
          limits:
            cpu: "500m"
            memory: "128Mi"
      restartPolicy: Never
  backoffLimit: 4

Note restartPolicy: Never — for Jobs, you can only use Never or OnFailure, not used Always.

2. Job Completion Modes

Kubernetes supports three completion modes for Jobs, suitable for different use cases.

2.1 NonIndexed (Default)

Job is completed when there are enough successful completions. The Pods are unordered — they all do the same job and the Job needs enough completions Pod to succeed.

apiVersion: batch/v1
kind: Job
metadata:
  name: nonindexed-parallel-job
spec:
  completions: 5        # Cần 5 Pod hoàn thành thành công
  parallelism: 2        # Chạy tối đa 2 Pod cùng lúc
  completionMode: NonIndexed  # Đây là default, có thể bỏ qua
  template:
    spec:
      containers:
      - name: worker
        image: busybox:1.35
        command: ["sh", "-c", "echo Processing task; sleep 10; echo Done"]
      restartPolicy: Never
  backoffLimit: 3

2.2 Indexed Jobs

Indexed Jobs is a very powerful feature — each Pod receives a unique index from 0 to completions-1 via the environment variable JOB_COMPLETION_INDEX. This is ideal for data partitioning: each Pod processes a defined portion of data.

apiVersion: batch/v1
kind: Job
metadata:
  name: indexed-data-processor
spec:
  completions: 10       # 10 partitions
  parallelism: 3        # Xử lý 3 partitions cùng lúc
  completionMode: Indexed
  template:
    spec:
      containers:
      - name: data-processor
        image: python:3.11-slim
        command:
        - python3
        - -c
        - |
          import os
          partition_id = int(os.environ['JOB_COMPLETION_INDEX'])
          total_partitions = 10
          # Xử lý dữ liệu từ partition partition_id
          start = partition_id * 1000
          end = start + 1000
          print(f"Processing records {start} to {end}")
          # ... thực tế sẽ query database hoặc đọc file
        env:
        - name: JOB_COMPLETION_INDEX
          valueFrom:
            fieldRef:
              fieldPath: metadata.annotations['batch.kubernetes.io/job-completion-index']
      restartPolicy: Never
  backoffLimit: 6

Kubernetes automatically injects the variable JOB_COMPLETION_INDEX into each Pod. Pod 0 processes partition 0, Pod 1 processes partition 1, and so on. — never duplicates even if Pod restarts.

2.3 Work Queue

With the work queue pattern, multiple Pods take tasks from one queue (Redis, RabbitMQ, SQS). The job is completed when the queue is empty and there are no more Pods in process.

apiVersion: batch/v1
kind: Job
metadata:
  name: queue-worker
spec:
  parallelism: 4    # 4 workers đồng thời
  # completions không set = work queue mode (hoàn thành khi 1 Pod exit 0)
  template:
    spec:
      containers:
      - name: worker
        image: my-queue-worker:v1.2
        env:
        - name: QUEUE_URL
          value: "redis://redis-service:6379/queue:tasks"
        - name: MAX_TASKS
          value: "100"
        resources:
          requests:
            cpu: "500m"
            memory: "256Mi"
      restartPolicy: OnFailure

3. Job Parameters Details

Understanding Job parameters helps you optimize for each use case:

  • completions: Total number of Pods that need to complete successfully. Default is 1.
  • parallelism: Maximum number of Pods running simultaneously. Default is 1.
  • backoffLimit: Number of retries before the Job is marked failed. Default is 6.
  • activeDeadlineSeconds: Maximum time (seconds) the Job is allowed to run. Exceeded → Job terminated.
  • ttlSecondsAfterFinished: Delete Job (and Pods) N seconds after completion.
apiVersion: batch/v1
kind: Job
metadata:
  name: time-limited-job
spec:
  completions: 3
  parallelism: 3
  backoffLimit: 2
  activeDeadlineSeconds: 600    # Job phải xong trong 10 phút
  ttlSecondsAfterFinished: 3600 # Xóa sau 1 giờ
  template:
    spec:
      containers:
      - name: worker
        image: busybox:1.35
        command: ["sh", "-c", "sleep 30 && echo completed"]
      restartPolicy: Never

4. Pod Failure Policies (K8s 1.31+)

From Kubernetes 1.31, Pod Failure Policy allows you to define granular behavior when a Pod fails — retry is not always recommended.

apiVersion: batch/v1
kind: Job
metadata:
  name: job-with-failure-policy
spec:
  completions: 5
  parallelism: 2
  backoffLimit: 6
  podFailurePolicy:
    rules:
    # Nếu Pod exit với code 42 (business error), đừng retry — fail ngay
    - action: FailJob
      onExitCodes:
        containerName: main
        operator: In
        values: [42]
    # Nếu node bị preempt (OOM, spot interruption), ignore và retry
    - action: Ignore
      onPodConditions:
      - type: DisruptionTarget
    # Các lỗi khác: retry như bình thường
    - action: Count
      onExitCodes:
        operator: NotIn
        values: [0, 42]
  template:
    spec:
      containers:
      - name: main
        image: my-batch-processor:v2
        command: ["./process"]
      restartPolicy: Never

The action can be used:

  • FailJob: Stop the entire Job immediately, mark failed
  • Ignore: Not counted in backoffLimit, create new Pod
  • Count: Counts into backoffLimit as usual (default behavior)

5. Job TTL — Automatic Cleanup

Jobs and their Pods will persist forever after completion without a cleanup mechanism. Use ttlSecondsAfterFinished to automatically delete:

apiVersion: batch/v1
kind: Job
metadata:
  name: cleanup-demo
spec:
  ttlSecondsAfterFinished: 300  # Xóa 5 phút sau khi xong (kể cả failed)
  template:
    spec:
      containers:
      - name: task
        image: busybox:1.35
        command: ["echo", "Hello from Job"]
      restartPolicy: Never

You can also patch existing Jobs: kubectl patch job old-job -p '{"spec":{"ttlSecondsAfterFinished":0}}' — this removes the Job immediately.

6. CronJobs — Schedule Task

CronJob automatically creates Jobs on a regular schedule, using familiar cron syntax.

apiVersion: batch/v1
kind: CronJob
metadata:
  name: daily-report
  namespace: production
spec:
  schedule: "0 2 * * *"    # 2 giờ sáng mỗi ngày
  timeZone: "Asia/Ho_Chi_Minh"   # GA từ K8s 1.27
  concurrencyPolicy: Forbid        # Không chạy job mới nếu job cũ đang chạy
  startingDeadlineSeconds: 300     # Nếu trễ quá 5 phút, bỏ qua
  successfulJobsHistoryLimit: 3    # Giữ 3 successful jobs gần nhất
  failedJobsHistoryLimit: 1        # Giữ 1 failed job gần nhất
  jobTemplate:
    spec:
      ttlSecondsAfterFinished: 86400  # Xóa sau 24 giờ
      template:
        spec:
          containers:
          - name: report-generator
            image: my-report-app:v1.5
            command: ["python", "generate_report.py", "--date", "yesterday"]
            env:
            - name: DB_HOST
              valueFrom:
                secretKeyRef:
                  name: db-credentials
                  key: host
          restartPolicy: OnFailure

6.1 CronJob Timezone Support (GA K8s 1.27)

Before K8s 1.27, all CronJobs used the UTC of the controller. From K8s 1.27, timeZone field is GA — you can specify any timezone according to IANA timezone database:

spec:
  schedule: "0 9 * * 1-5"          # 9 giờ sáng thứ 2-6
  timeZone: "Asia/Ho_Chi_Minh"     # Vietnam timezone (UTC+7)

Popular timezones:

  • Asia/Ho_Chi_Minh — Vietnam (UTC+7)
  • Asia/Singapore — Singapore (UTC+8)
  • America/New_York — Eastern US
  • Europe/London — UK
  • UTC — Coordinated Universal Time

6.2 ConcurrencyPolicy

Important thing to decide: what to do if an old Job is not finished when the new Job is scheduled to run?

  • Allow (default): Create a new Job even if the old Job is running — be careful with race conditions
  • Forbid: Skip the new Job, the old Job still continues
  • Replace: Delete old Job, create new Job to replace

7. JobSet — CNCF Project for Distributed Jobs

JobSet is a CNCF project (currently in the Sandbox stage) designed to coordinate multiple dependent Jobs. This is the ideal tool for distributed ML training pipelines__HTMLTAG_267___.

JobSet Settings:

kubectl apply --server-side -f \
  https://github.com/kubernetes-sigs/jobset/releases/download/v0.7.0/manifests.yaml

7.1 JobSet for Distributed ML Training

Scenario: Train a model with Parameter Server architecture — one group of pods as parameter servers (store gradients), another group as workers (calculation).

apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
  name: ml-training-pytorch
  namespace: ml-training
  annotations:
    jobset.sigs.k8s.io/exclusive-topology: kubernetes.io/hostname
spec:
  failurePolicy:
    maxRestarts: 3        # Restart toàn bộ JobSet nếu có failure
  replicatedJobs:
  # Parameter Server: lưu model state, nhận gradients từ workers
  - name: parameter-server
    replicas: 1
    template:
      spec:
        completions: 2
        parallelism: 2
        completionMode: Indexed
        template:
          spec:
            containers:
            - name: ps
              image: pytorch/pytorch:2.2-cuda12.1-cudnn8-runtime
              command: ["python", "train.py", "--role", "ps"]
              env:
              - name: ROLE
                value: "parameter-server"
              - name: JOB_INDEX
                valueFrom:
                  fieldRef:
                    fieldPath: metadata.annotations['batch.kubernetes.io/job-completion-index']
              resources:
                requests:
                  cpu: "4"
                  memory: "16Gi"
            restartPolicy: Never

  # Workers: tính toán gradients, gửi lên PS
  - name: worker
    replicas: 1
    template:
      spec:
        completions: 8      # 8 workers
        parallelism: 8
        completionMode: Indexed
        template:
          spec:
            containers:
            - name: worker
              image: pytorch/pytorch:2.2-cuda12.1-cudnn8-runtime
              command: ["python", "train.py", "--role", "worker"]
              env:
              - name: ROLE
                value: "worker"
              - name: PS_HOSTS
                value: "ml-training-pytorch-parameter-server-0-0.ml-training-pytorch:8080,ml-training-pytorch-parameter-server-0-1.ml-training-pytorch:8080"
              resources:
                requests:
                  cpu: "4"
                  memory: "16Gi"
                  nvidia.com/gpu: "1"
                limits:
                  nvidia.com/gpu: "1"
            restartPolicy: Never

7.2 Outstanding Features of JobSet__HTMLTAG_276___
  • Failure policy propagation: If a Job in the set fails, the entire JobSet can restart or fail together — no "orphan" Jobs__HTMLTAG_281___
  • DNS-based communication: Jobs in JobSet automatically have DNS records to communicate with each other ({jobset-name}-{job-name}-{job-index}-{pod-index}.{jobset-name})
  • Exclusive topology: Ensure Pods of the same Job are scheduled on the same rack/node (reduces network latency)
  • Startup sequencing__HTMLTAG_294___: Only start worker after PS is ready

7.3 JobSet Tracking

# Xem trạng thái JobSet
kubectl get jobset -n ml-training

# Xem chi tiết
kubectl describe jobset ml-training-pytorch -n ml-training

# Xem logs của parameter server
kubectl logs -l jobset.sigs.k8s.io/job-name=parameter-server -n ml-training

# Xem logs của tất cả workers
kubectl logs -l jobset.sigs.k8s.io/job-name=worker -n ml-training --prefix

8. Summary: When to Use What?

  • SimpleJob: One task, run once — use basic Job
  • Parallel processing without order: NonIndexed Job with completions + parallelism
  • Data partitioning: Indexed Job — each Pod processes the specified partition
  • Queue-based processing: Work queue Job + Redis/RabbitMQ
  • Scheduled tasks: CronJob with timezone support
  • Distributed training/HPC: JobSet for multi-job coordination

Jobs and CronJobs are the foundation of every batch processing system on Kubernetes. Understanding completion modes and failure policies helps you build reliable pipelines, especially with AI/ML workloads becoming more popular.