Runtime Provisioner 0.8.0 Readme

This is a REST application used to provision KSML applications in Kubernetes.

When a deployment is requested, this provisioner fetches the KSML Helm Charts from the configured Helm registry, generates values.yaml based on the requested configuration and deploys it in the same Kubernetes cluster where this provisioner is deployed.

The Provisioner requires certain privileges scoped to the namespace in Kubernetes to deploy KSML pod including a ServiceMonitor resource for Prometheus monitoring.

Limitations

The Provisioner has the following limitations:

  1. The Provisioner can only deploy KSML applications in the same Kubernetes cluster where it is deployed.

  2. The Provisioner can pull charts from OCI-compatible registries only. Nexus registry is not supported.

  3. KSML deployment names are generated as {tenant}-{instance}-{environment}-{application}-ksml. Kubernetes limits label values to 63 characters, so if the combined name exceeds the allowed length, it is replaced by {tenant}-{instance}-{hash16}-ksml where the hash is a 16-character SHA256-based prefix of the full name (e.g. mytenant-myinstance-c319f3cd9b7c02cc-ksml). The original identity is always preserved in the pod labels and can be used to locate the pod:

    kubectl get pods -n <namespace> \
      -l axual.io/tenant=<tenant>,axual.io/instance=<instance>,axual.io/environment=<environment>,axual.io/application=<application>

Required Configuration

Set these environment variables before you start the Provisioner.

Environment Variables Description

NAMESPACE

The Kubernetes namespace where KSML application will be deployed.

REGISTRY_URL

The Helm Chart Registry URL where KSML Helm Charts are pulled from.

CHART_NAME

The name of the KSML Helm Chart.

CHART_VERSION

The version of the KSML Helm Chart.

Optional Configuration

These environment variables change the default behaviour of the Provisioner.

Environment Variables Description

REGISTRY_AUTH_ENABLED

Set to true if Helm Chart registry requires authentication. Default false.

REGISTRY_USERNAME

Username of the Helm Chart Registry defined in REGISTRY_URL.

REGISTRY_PASSWORD

Password of the Helm Chart Registry defined in REGISTRY_URL.

CUSTOM_VALUES_FILE

Path to a custom values file containing extra configurations for KSML deployment. This is useful when configs like resources, securityContext, topologySpreadConstraints etc need to be passed.

CLIENT_CA_FILE

Path to a base64-encoded PEM file containing one or more CA certificates. It is used to validate the Helm Registry’s server certificate.

KSML_ENABLED

Set to false to serve Kafka Connect logs only. The provisioner then starts without any of the Helm or registry settings above. Default true.

CONNECT_ENABLED

Set to false to switch Kafka Connect log reading off. The endpoint then answers 501. Default true.

CONNECT_NAMESPACES

Namespaces to search for Kafka Connect worker pods, separated by commas. Defaults to the provisioner’s own namespace. Connect clusters usually run in another one.

CONNECT_CONTAINER_NAME

Names the Kafka Connect worker container. Only needed for a pod where the container cannot be found automatically.

CONNECT_FILTER_WINDOW_PERCENT

Raw lines read per worker pod when a request filters by connector, as a percentage of the lines asked for. Default 2000, meaning 20x. See “Filtering and the read window” below.

CONNECT_FILTER_WINDOW_MAX_LINES

Ceiling on that window, in lines. Default 50000, which only binds above a percentage of 5000, because tailLines is capped at 1000. See “Filtering and the read window” below.

CONNECT_DOWNLOAD_CAP_BYTES

Most a whole-log download may copy in total, across every worker. Default 67108864, 64 MiB. A safety net, not a feature limit: the container runtime’s rotation size caps a real download first. Zero or less falls back to the default.

MAX_CONCURRENT_LOG_STREAMS

How many log reads may run at the same time, for KSML and Kafka Connect together. It counts pod streams, not requests, because one Connect request reads every worker of the cluster. Default 200.

INSECURE_SKIP_TLS_VERIFY

If true, the validation of Helm Registry’s server certificate. This can lead to man in the middle attacks and should not be used in Production.

DISTRIBUTED_TRACING_ENABLED

If true, enables distributed tracing with OpenTelemetry. To configure the exporter, see section “Enable Distributed Tracing.” Default false.

BASIC_AUTH_ENABLED

Set to true if HTTP server should have basic authentication. Default false.

BASIC_AUTH_USERNAME

Username of the basic authentication of HTTP server.

BASIC_AUTH_PASSWORD

Password of the basic authentication of HTTP server.

SERVER_TLS_CERTIFICATE

PEM encoded x509 certificate for HTTP server.

SERVER_TLS_PRIVATE_KEY

PEM encoded private key for HTTP server.

Filtering and the read window

Kubernetes cannot filter log content. tailLines is applied to raw lines by the kubelet, and the connector filter only runs afterwards, inside the provisioner. So reading only as many lines as were asked for returns almost nothing on a cluster where the wanted connector is a small share of the traffic.

A filtered request is therefore read over a wider raw window:

raw lines per pod = min(tailLines * CONNECT_FILTER_WINDOW_PERCENT / 100, CONNECT_FILTER_WINDOW_MAX_LINES)

with the requested count as a floor, so the window is never narrower than the answer needs. A request without a filter reads exactly tailLines per pod, because every line read is a line returned.

The endpoint caps tailLines at 1000, so at the default percentage of 2000 the widest window is 1000 * 2000 / 100 = 20000 raw lines per pod. The ceiling of 50000 sits above every value the endpoint can produce at that percentage, and only binds once CONNECT_FILTER_WINDOW_PERCENT is raised above 5000. Raising the ceiling alone changes nothing.

How far back this reaches is the reason the ceiling is high:

how far back = raw lines read / lines the worker writes per second

Measured on a real worker writing 0.39 lines per second: 100 raw lines reached 1.6 minutes back, 2000 reached 33 minutes, 5000 reached 76 minutes, and the default maximum of 20000 reached about 14 hours. On a cluster at normal log levels, around 0.07 lines per second, 5000 reached about 20 hours. Raise the percentage for a quiet connector on a chatty cluster; lower it if reads feel slow.

Reading wide is cheap. Non-matching lines are dropped as they are read, so memory follows the answer rather than the window: the lines one request holds are bounded at 4 MiB across its worker pods. That allowance is split into a reserved floor per pod plus a shared pool, so the answer does not depend on which kubelet replied first, and one busy worker can still use everything the others did not reserve. The answer is serialised on top of that, so peak memory per request is a few times the 4 MiB rather than exactly it. Measured read times for one worker were 0.17 s for 5000 lines and 0.37 s for a whole 25 000 line log.

What the window cannot do is outlive the pod. Kubernetes keeps roughly containerLogMaxSize * containerLogMaxFiles of log per container, commonly 50 MiB, and deletes it with the pod. Nothing is stored elsewhere.

Required Kubernetes Permissions

Runtime Provisioner requires a certain set of permissions on the service account to deploy KSML applications. These permissions are described below.

The required service account, role and role binding are automatically created by the Provisioner Helm Chart. No special configurations are required.

Kubernetes Resource API Group Permissions

configmaps

GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE

pods

GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE

pods

metrics.k8s.io

GET, LIST

pods/log

GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE

secrets

GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE

serviceaccounts

GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE

services

GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE

statefulsets

apps

GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE

jobs

batch

GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE

prometheusrules

monitoring.coreos.com

GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE

servicemonitors

monitoring.coreos.com

GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE

ingresses

networking.k8s.io

GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE

networkpolicies

networking.k8s.io

GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE

persistentvolumeclaims

storage.k8s.io

GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE

Provisioning KSML applications in different namespace

By default, KSML applications will be provisioned in the same namespace where the provisioner is deployed. If the applications need to be deployed in a different namespace (eg ksml), pass below values.

env:
  - name: NAMESPACE
    value: "ksml"

rbac:
  namespace: "ksml"

The NAMESPACE environment variable informs the provisioner about where to deploy KSML applications. The rbac.namespace ensures that the provisioner has the necessary permissions to deploy KSML applications in the target namespace.

Bring your own ServiceAccount and RBAC

By default, the helm chart deploys a service account and associated RBAC resources with correct permissions. But for some reason, if you want to use your own service account and RBAC resources, set below configuration:

serviceAccount:
  create: false
  name: my-custom-sa
rbac:
  create: false

A custom service account can be passed via serviceAccount.name. It is your responsibility to ensure this service account has correct permissions to deploy KSML applications (see previous section for exact permissions). The easiest way to do this is to create a Role with relevant permissions and a RoleBinding to associate the role with the service account.

In most cases though, you don’t need this and can keep life simple by having rbac.create: true and let the Helm chart take care of the complexity.

Who can read which logs

The service does not check who is calling. Basic auth, the NetworkPolicy and Platform Manager’s group check are the controls around it.

route.enabled=true and ingress.enabled=true publish the service outside the cluster, where anything that reaches it can deploy and undeploy KSML applications and read these logs. Turn BASIC_AUTH_ENABLED on before you enable either, because it defaults to false.

CONNECT_NAMESPACES is the boundary. The chart grants read permission only for those namespaces, so the service cannot read one that is not listed. Tenants that must not see each other’s logs need separate namespaces, or their own provisioner. The address is stored per Connect cluster, so a second provisioner needs no extra work.

Prometheus metrics

Served on port 2112 and scraped by the ServiceMonitor in the chart. Reading Kafka Connect logs adds four metrics, each labelled tenant, instance and connect_cluster:

Metric Type Meaning

kafkaconnect_log_lines_read_total

counter

Lines taken from the worker pods, before filtering

kafkaconnect_log_lines_filtered_total

counter

Lines removed by the connector filter

kafkaconnect_log_read_errors_total

counter

Reads that ended with an error, with a mode label of batch or stream

kafkaconnect_log_streams_active

gauge

Streams open right now

read minus filtered is what reached the caller. These are not the GET /metrics endpoint below, which reports CPU and memory of one KSML deployment.

Enable Distributed Tracing

Runtime Provisioner supports distributed tracing with OpenTelemetry. To enable, set environment variable DISTRIBUTED_TRACING_ENABLED to true.

To export traces to an OTel backend, configure below environment variables.

Environment Variable Description

OTEL_EXPORTER_OTLP_ENDPOINT

Endpoint of OTel collector. Must start with http/https and include port number.

OTEL_EXPORTER_OTLP_HEADERS

Optional headers.

OTEL_SERVICE_NAME

A unique name to recognize the service.

Enable Basic Authentication on the HTTP server

Basic authentication can be enabled by supplying the below configuration:

Environment Variable Description

BASIC_AUTH_ENABLED

Set to “true” to enable Basic authentication. Default is “false”.

BASIC_AUTH_USERNAME

Username for Basic authentication. Required when BASIC_AUTH_ENABLED is “true”.

BASIC_AUTH_PASSWORD

Password for Basic authentication. Required when BASIC_AUTH_ENABLED is “true”.

Below is an example configuration:

env:
  - name: BASIC_AUTH_ENABLED
    value: "true"
  - name: BASIC_AUTH_USERNAME
    value: "username"
  - name: BASIC_AUTH_PASSWORD
    value: "password"

Enable TLS on the HTTP server

TLS can be enabled by supplying the below configuration:

Environment Variable Description

SERVER_TLS_CERTIFICATE

Certificate in PEM format.

SERVER_TLS_PRIVATE_KEY

Private key in PEM format.

Below is an example configuration:

env:
  - name: SERVER_TLS_CERTIFICATE
    valueFrom:
      secretKeyRef:
        key: tls.crt
        name: runtime-provisioner
  - name: SERVER_TLS_PRIVATE_KEY
    valueFrom:
      secretKeyRef:
        key: tls.key
        name: runtime-provisioner

Note that in the above example, the secret runtime-provisioner must exist with certificate in tls.crt and private key in tls.key.

An OpenShift Route created with the default route.tls.termination: passthrough hands the TLS session straight to the pod, so it needs both variables above. Without them the pod serves plain HTTP and the handshake fails with no error on the Route itself. Set route.tls.termination: edge instead if you want the router to terminate TLS.

Setting fsGroup for KSML PersistentVolume Access

When using a PersistentVolume (PV) for KSML Stateful feature, the volume is mounted at /ksml-store path which is owned by the root user. Without additional configuration, the KSML process running in the Pod cannot create the required subfolder /ksml-store/data.

To ensure the KSML process can write to the volume, set podSecurityContext.fsGroup so that the files in the volume are owned by the group of the user running the KSML process. The default user is part of group 0 so set fsGroup to 0 unless you are using a different user. For example:

runtime-provisioner:
  customValues:
    podSecurityContext:
      fsGroup: 0

This ensures that the KSML Pod can create and manage the /ksml-store/data folder inside the PV automatically.

REST API endpoints

The Provisioner exposes the following HTTP endpoints.

POST /deploy

Request Body:

{
  "tenant": "dizzl",
  "instance": "sandbox",
  "environment": "dev",
  "application": "sensor-inspector",
  "replicas": 1,
  "definition": "{streams: {sensor_source_avro: {topic: ksml_sensordata_avro, keyType: string, valueType: 'avro:SensorData'}}, functions: {log_message: {type: forEach, parameters: [{name: format, type: string}], code: 'log.info(\"Consumed {} message - key={}, value={}\", format, key, value)'}}, pipelines: {consume_avro: {from: sensor_source_avro, forEach: {code: 'log_message(key, value, format=\"AVRO\")'}}}}",
  "kafkaProperties": {
    "bootstrap.servers": "broker:9093",
    "schema.registry.url": "http://schemaregistry:8081",
    "application.id": "io.axual.sensor-inspector"
  },
  "schemas": [
    {
      "schemaFileName": "SensorData.avsc",
      "schemaBody": "..."
    }
  ]
}

Returns 200 OK if deployment is successful.

GET /status

Returns the status of KSML deployment(s). This endpoint supports two modes:

Detailed Status (single deployment)

Query Parameters: - tenant (required): The tenant name - instance (required): The instance name - environment (required): The environment name - application (required): The application name - deploymentMode (optional): Either StatefulSet or Job

Example: GET /status?tenant=dizzl&instance=sandbox&environment=dev&application=sensor-inspector

Returns the status of a specific KSML deployment including Helm and Kubernetes details.

The status can be one of Running, Starting, Failing, or Completed (for Jobs).

{
  "helm": {
    "name": "dizzl-sandbox-dev-sensor-inspector-ksml",
    "chart": "ksml",
    "version": "0.0.0-snapshot",
    "appVersion": "snapshot",
    "namespace": "default",
    "revision": 1,
    "status": "deployed",
    "deployedAt": "2024-08-26T13:04:20Z"
  },
  "kubernetes": {
    "name": "dizzl-sandbox-dev-sensor-inspector-ksml",
    "status": "Running",
    "replicas": 1,
    "readyReplicas": 1,
    "restartsCount": 3,
    "lastRestart": "2024-08-26T14:01:05Z"
  }
}

Bulk Status (all deployments for tenant/instance)

Query Parameters: - tenant (required): The tenant name - instance (required): The instance name

Example: GET /status?tenant=dizzl&instance=sandbox

Returns the status of all KSML deployments for the given tenant and instance combination. This searches for pods with matching axual.io/tenant and axual.io/instance labels. Each status includes resource limits (CPU in cores, memory in MB) from the KSML container.

{
  "statuses": [
    {
      "name": "dizzl-sandbox-dev-sensor-inspector-ksml",
      "status": "Starting",
      "replicas": 1,
      "readyReplicas": 0,
      "resourceLimits": {
        "cpuCores": "1",
        "memoryMB": "4096"
      }
    },
    {
      "name": "dizzl-sandbox-prod-sensor-inspector-ksml",
      "status": "Running",
      "replicas": 1,
      "readyReplicas": 1,
      "resourceLimits": {
        "cpuCores": "0.5",
        "memoryMB": "2048"
      }
    }
  ],
  "tenant": "dizzl",
  "instance": "sandbox"
}

GET /logs?tenant=dizzl&instance=sandbox&environment=dev&application=sensor-inspector&tailLines=10

Returns or streams logs for the KSML deployment. The response format is determined by the Accept request header.

Query Parameters: - tenant (required): The tenant name - instance (required): The instance name - environment (required): The environment name - application (required): The application name - deploymentMode (optional): Either StatefulSet or Job - tailLines (optional): Number of log lines to return or replay at stream start (1–1000, default 100)

Batch mode (default — no Accept header, or Accept: application/json):

Returns the last N lines of logs for the KSML deployment as a single JSON response.

Sample response:

{
  "tailLines": 1,
  "logs": [
    {
      "replicaName": "0",
      "lines": "2024-07-16T13:06:39,394Z [io.axual.sensor-inspector-6a395bc1-f3bc-4eb2-9331-c3d1b4e45f8b-StreamThread-1] INFO  ksml.function.log_message - Consumed AVRO message - key=sensor27, value={'city': 'Xanten'}\n"
    },
    {
      "replicaName": "1",
      "lines": "2024-07-16T13:06:39,841Z [io.axual.sensor-inspector-1393ae07-9321-424d-9453-b86383c816b3-StreamThread-1] INFO  ksml.function.log_message - Consumed AVRO message - key=sensor28, value={'city': 'Amsterdam'}\n"
    }
  ]
}

Streaming mode (Accept: application/x-ndjson):

Streams real-time logs as NDJSON over a chunked HTTP response. The connection stays open and each log line is forwarded to the client as it arrives, making it suitable for live log tailing. A keepalive message is emitted every 15 seconds on idle connections to prevent intermediate proxies from closing the connection.

Each line in the response body is a self-contained JSON object:

Object type Field Description

Log line

line

A single log line from the pod

Keepalive

keepalive

Empty string; sent every 15 s to keep the connection alive

Error

error

Emitted when the log stream terminates with an error

Sample response (each line is a separate JSON object):

{"line":"2024-07-16T13:06:39,394Z [StreamThread-1] INFO ksml.function.log_message - Consumed AVRO message - key=sensor27"}
{"line":"2024-07-16T13:06:39,841Z [StreamThread-1] INFO ksml.function.log_message - Consumed AVRO message - key=sensor28"}
{"keepalive":""}
{"line":"2024-07-16T13:06:54,102Z [StreamThread-1] INFO ksml.function.log_message - Consumed AVRO message - key=sensor29"}
{"error":"Pod CrashLoopBackOff: exit code 1"}

GET /metrics?tenant=dizzl&instance=sandbox&environment=dev&application=sensor-inspector&deploymentMode=StatefulSet

Returns the current CPU and memory usage metrics for the KSML deployment. This endpoint requires the Kubernetes Metrics Server to be installed in the cluster.

Query Parameters: - tenant (required): The tenant name - instance (required): The instance name - environment (required): The environment name - application (required): The application name - deploymentMode (optional): Either StatefulSet or Job

Sample response:

{
  "cpu": {
    "usedMillicores": "250"
  },
  "memory": {
    "usedMb": "512"
  }
}

Note: If the Metrics Server is not installed or metrics are not yet available for a newly started pod, the endpoint will return a 404 error with an appropriate message.

POST /undeploy

Request Body:

{
  "tenant": "dizzl",
  "instance": "sandbox",
  "environment": "dev",
  "application": "sensor-inspector"
}

Returns 200 OK if removal of deployment is successful

GET /kafkaconnect/logs?tenant=dizzl&instance=sandbox&connectCluster=payments&tailLines=10&connector=payments-sink

Reads the worker pods of one Kafka Connect cluster and merges their logs into one answer. Pods are found by the axual.io/tenant, axual.io/instance and axual.io/connect-cluster labels, so connectCluster must match the label, not the Helm release name.

tailLines is how many lines of history the answer holds in total, across the workers, from 0 to 1000. 0 means live lines only, which only a stream can serve. It bounds the answer and not the read: every worker is read over the whole total, and the merge then keeps the newest tailLines across the cluster, so one busy worker can supply the whole answer. It means the same in both modes, so a stream shows the same window a batch request would answer with, within the one limit described under streaming below.

connector keeps only the lines naming that connector. A JSON entry is matched on its connector.context field, which is exact. Any other line falls back to looking for the name in the text, which is still needed: Kafka’s own producer and consumer threads log outside the task’s logging context and name the connector in their client id instead, and 7 lines in 618 on a live worker were of that shape. Either way it runs after the read, so the read window is widened first, see “Filtering and the read window”.

There is no way to read the log of a container that has already stopped. That was tried and taken out: the node keeps the old log on its own schedule, so it was usually gone, and a worker killed outright writes no reason into its log anyway. To look at a crashed worker use kubectl logs --previous and kubectl describe pod, which also shows the exit code.

minLevel keeps only the lines at that severity or above, from TRACE, DEBUG, INFO, WARN, ERROR and FATAL. Any other value is a 400. It widens the read window for the same reason connector does, including when no connector is given: warnings and errors are 1.4% of a worker’s log, measured over 5995 lines, so a window sized to the request comes back nearly empty.

A kept line keeps the lines that belong to it. Two thirds of a worker’s lines are indented frames under a Java error and carry no level of their own, so a filter judging each line alone would keep Error while starting connector and drop the ninety lines explaining why.

A line whose own severity cannot be read follows the line above it, exactly as an indented frame does, because that is what such a line nearly always is. Keeping it regardless would leak everything a worker prints before its logging starts, and that is not a small amount: the first few hundred lines of a Kafka Connect worker are shell tracing and a dumped configuration file, none of which carry a level. A reader who asked for errors would get those instead of an empty window saying nothing matched. A blank line follows the line above it for the same reason.

So a log format this reader cannot classify at all disappears under a threshold. That is the honest outcome and it explains itself, because the empty window says how far it looked and which filter emptied it, and both layouts the chart writes always carry a level.

The severity is read from either log format: the log.level field of a JSON entry, or a level word near the start of a plain-text line. Both are what the chart’s two layouts produce.

Both stay supported even though kafka-connect-helm 0.5.0 writes JSON and no longer offers a choice. A cluster upgraded to that chart keeps hours of plain text in the same file until it rotates out, so one answer routinely mixes the two, and a cluster on an older chart, or one this chart did not create, can send plain text indefinitely. Note that the KSML endpoint is a different reader entirely: it streams lines without parsing them, so KSML has no severity filter and is unaffected by any of this.

Answers with one JSON document by default:

{
  "tailLines": 10,
  "entries": [
    { "line": "...", "origin": "kc-alpha-kafka-connect-connect-0", "timestamp": "2026-09-01T08:54:10.092028739Z" },
    { "line": "...", "origin": "kc-alpha-kafka-connect-connect-1", "timestamp": "2026-09-01T08:54:10.301114000Z" }
  ],
  "linesRead": 4000,
  "linesReturned": 10,
  "truncated": false,
  "degraded": ["kc-alpha-kafka-connect-connect-2"]
}

entries is the answer, in time order, oldest first. linesRead counts every raw line taken off the workers and linesReturned how many the answer holds, which is what tells “the cluster has nothing more” apart from “plenty was read and none of it matched”. truncated means the byte budget cut the answer short. degraded names the worker pods that could not be read at all, so the answer covers only the rest of the cluster.

Send Accept: application/x-ndjson to stream instead. Each line is then one JSON object, and every log line carries origin, the pod it came from, and timestamp, when the node stamped it:

{"line":"2026-08-11 13:21:42 INFO ...","origin":"kc-alpha-kafka-connect-connect-0","timestamp":"2026-08-11T13:21:42.104512Z"}
{"event":"streamEnded","origin":"kc-alpha-kafka-connect-connect-1","line":"--- worker ... ended ---"}
{"keepalive":""}

A stream sends the history first and then follows the log. Both come from one read per worker, opened with the window and follow together, so the kubelet seeks back over the window and then keeps following the same file. So a client needs one request, not two.

Request What arrives

Accept: application/json, tailLines=100

history only, one document

Accept: application/x-ndjson, tailLines=100

the same window, line by line, then new output as it is written

Accept: application/x-ndjson, tailLines=0

new output only, no history

Do not fetch the history with a second batch request. Two requests read two different windows that overlap at the recent end, and the client then receives the same lines from both with no way to tell them apart. A client that already holds the history opens the stream with tailLines=0, which asks for new output only. Platform Manager sends a fixed tailLines on its own stream, so through Platform Manager a stream always carries history: ask it once and do not fetch the history yourself.

The kubelet’s join is exact. Where the history ends is our judgement. One read cannot lose a line between the window and the new output, and cannot send the same line twice. But the stream carries no marker between the two, so the provisioner decides where the history stops: when the window has been read, when a line stamped later than the request arrives, when the read falls silent for 400 ms, or after 5 seconds, whichever comes first. Lines that miss the merged block are still sent, straight after it, in their own worker’s order. So a busy cluster can show a few lines out of order at the seam, and a stream forced to end its history early can send more lines than tailLines asked for. No line is lost, and none arrives twice.

That last point is the one difference between the two modes: tailLines bounds a batch answer exactly, while on a stream it bounds the merged block and the rest of the window still follows it.

The history lines arrive in time order across the workers, exactly as entries does, and each carries origin and timestamp, so a stream describes a line the same way the batch answer does. timestamp is absent on a line the node never stamped, which is a stack trace frame or another continuation of the line above it. New output cannot be ordered across workers at all, because holding lines back to sort them would delay a live tail, so each worker’s new lines are sent as they arrive.

A stream stays open while at least one worker still sends lines, and closes after 30 minutes.

When a filtered stream reads a wide window and matches nothing, the silent-stream notice says so, for example No lines for this connector in the last 4000 lines read from its workers. Waiting for output from the application…​. A reader can then tell a quiet connector from a window that did not look far enough back.

The notice names the filter that emptied the window, because a reader told the wrong one lifts the wrong control and still sees nothing. With minLevel alone it reads No lines at WARN or above in the last 4000 lines read from the workers. …​, and with both filters No lines for this connector at WARN or above in the last 4000 lines read from its workers. …​. A minLevel of TRACE keeps every level, so it is not named, exactly as it does not widen the read window.

When every worker returns fewer raw lines than the window asked for, the read reached the start of the log the node still has, and the notice says that instead: No lines for this connector in the whole log the workers still have, 340 lines. …​, or without a filter This is the whole log the workers still have, 340 lines. …​. That is the difference between “ask for more” and “there is no more”, which a line count alone does not show. The node’s retention is the wall no request can pass: the kubelet keeps containerLogMaxSize per file and containerLogMaxFiles of them, 10Mi and 5 by default.

When the plain-text level reader can be removed

The reader that finds a level in a plain-text line exists for three reasons, and all three end:

  • A cluster upgraded to kafka-connect-helm 0.5.0 keeps plain-text history in the same log file until it rotates out. One filtered answer during testing came back with 3 JSON lines and 95 plain ones.

  • A cluster still on an older chart has not switched yet.

  • A Kafka Connect cluster this chart did not create can send anything.

So it becomes safe to delete once every cluster runs 0.5.0 or later and their log files have rotated. Written down because otherwise it is either deleted too early or kept for ever by default.

The KSML endpoint never reaches this reader: it streams lines without parsing them. Platform Manager’s and the console’s ability to show a plain-text line is a separate matter and must stay, because KSML sends plain text permanently.

A very long connector label

Most logging contexts are short, 21 to 31 bytes, such as [payments-sink|task-0]. A connector that copies between two clusters produced a 98-byte one, because it carries a whole client id:

[AdminClient clientId=src-payments-sink->dst-payments-sink|payments-sink|heartbeats-target-admin]

It appeared on 4 lines out of thousands, so it costs nothing today. It could matter if a customer runs many cross-cluster connectors and one logs heavily, since the field is on every line and counts against the read window as well as disk. The fix, if it is ever needed, is a shorter field such as the connector name alone. Do not remove the field: it is what the connector filter matches on.

The three numbers that decide when the history stops

The split between history and new output is the provisioner’s own judgement, so three constants decide it. They are constants on purpose: no cluster has needed a different value, and a setting is a contract to keep.

Constant Value What it does

historyIdleGap

400 ms

per worker: how long a worker’s read may say nothing before its history counts as complete. The replay arrives far faster than a worker writes, so the pause between them is wide

historyDeadline

5 s

a clock bound on the phase, per worker and again across the workers, because a worker that keeps writing never falls silent and a log shorter than the window never reaches its edge. The second use is what stops one silent worker holding the merged block back from all of them

clockSkewAllowance

1 s

per worker: how far a worker node’s clock may run ahead of ours before a replayed line looks like new output. The kubelet stamps lines on the node, so that comparison crosses two clocks

Reading any of them too early costs the merge for the rest of the replay, or a little latency. None of them can lose or repeat a line. If a cluster ever needs different values, they follow the CONNECT_FILTER_WINDOW_PERCENT pattern: an environment variable, a chart value and a row in the settings table.

Answers 404 when no worker pod matches the labels, 400 for a bad parameter, 503 when the stream limit is full, and 501 in Docker mode, where there is no Kubernetes API to read.

GET /kafkaconnect/logs/download?tenant=dizzl&instance=sandbox&connectCluster=payments

Sends every worker’s current log file as one plain-text attachment, in labelled sections (===== worker <pod> =====). It shares nothing with the endpoint above: no filtering, no merging, no line budget. Bytes go from the Kubernetes stream to the socket in 64 KiB chunks, so the provisioner’s memory does not grow with the log.

connector, minLevel and tailLines are ignored, because reading everything is the point.

What it can return is bounded by the container runtime, not by this service. Kubernetes serves the container’s current log file only. Rotated files stay on the node and no API reaches them. Measured on the test cluster: a container that had written 59 MiB held 83 MB across five files on disk, docker logs on the node returned all 62 MB of it, and the Kubernetes API returned 2.4 MB, the current file. So this endpoint returns one rotation window, and how much that is depends on when the last rotation happened.

For scale, the runtime’s rotation size is the ceiling: 10Mi with the kubelet’s containerLogMaxSize default, 20MB with the Docker json-file driver’s max-size. Reading a whole file measured about 19 MiB/s through the Kubernetes API, so a few workers take a few seconds.

CONNECT_DOWNLOAD_CAP_BYTES caps the total across every worker, 64 MiB by default. It is a safety net rather than a feature limit: a real cluster never comes near it, and it exists for the case where somebody configures a much larger rotation size. Hitting it still produces a valid file, with a final line saying it was truncated and how many workers were copied.

Kubelet timestamps are deliberately off. The worker writes one JSON object per line and a timestamp prefix would stop the file being readable as NDJSON; every entry already carries its own @timestamp.

A worker that cannot be read is named in the file and the others are still copied. If no worker can be read the answer is a status code, not a file whose every line is an error message.

GET /healthz and GET /readyz

/healthz reports only that the process runs, and makes no Kubernetes call. /readyz lists one pod in each configured namespace, so an ok also proves the service account may read pods. Both answer {"status":"ok"} and need no authentication.

The /ksml/ prefix

Every KSML endpoint above is also served under /ksml/, for example /ksml/logs. Both spellings behave identically. The flat paths are deprecated and will be removed in a later major version.


Docker Runtime Mode (Docker Compose only)

Not for production use. The Docker runtime is intended exclusively for local development and testing via Docker Compose. It has no high-availability, no persistent storage management, and no Kubernetes RBAC isolation. Do not run it in any shared or production environment.

The Runtime Provisioner can run in Docker mode by setting DEPLOYMENT_TARGET=docker. In this mode it manages KSML applications as plain Docker containers on the local Docker daemon instead of deploying Helm charts into Kubernetes.

How it works

When a /deploy request arrives the provisioner: 1. Writes a ksml-runner.yaml configuration file (with inline Kafka settings, schema registry config, and KSML definition) into a per-container directory on the host. 2. Starts a Docker container from KSML_IMAGE, mounting that directory so the KSML runner can read its config. 3. Attaches the container to any Docker networks listed in DOCKER_CONFIG_NETWORKS so it can reach Kafka brokers and schema registries running in the same Compose project.

On /undeploy the container is stopped and removed. Config files written for that container are also cleaned up.

The /logs and /status endpoints work the same way as in Kubernetes mode. The /metrics endpoint is not supported in Docker mode and always returns 404.

Required configuration

Set these environment variables to run the Provisioner in Docker mode.

Environment Variable Description

DEPLOYMENT_TARGET

Must be set to docker.

KSML_IMAGE

Docker image reference for the KSML runner (e.g. registry.example.com/ksml-runner:1.2.0).

DOCKER_CONFIG_DIR

Absolute path on the host where per-container config files are written and bind-mounted into KSML containers.

Optional configuration

These environment variables change the default behaviour in Docker mode.

Environment Variable Description

DOCKER_CONFIG_VOLUME

Name of a Docker named volume to mount instead of a bind mount. Required on Docker Desktop for Mac, where host bind-mounts from inside a container are not supported. See note below.

DOCKER_CONFIG_NETWORKS

Comma-separated list of Docker network names to attach KSML containers to (e.g. my-compose-project_default). Allows KSML containers to reach Kafka and schema registry services in the same Compose project.

Docker Compose setup

Below is a minimal example showing the provisioner alongside Kafka. The config directory is shared as a named volume so it works on both Linux and Docker Desktop for Mac.

volumes:
  ksml-configs:
    name: ksml-configs   # fixed name — must match DOCKER_CONFIG_VOLUME

services:
  runtime-provisioner:
    image: registry.example.com/runtime-provisioner:0.8.0
    environment:
      DEPLOYMENT_TARGET: docker
      KSML_IMAGE: registry.example.com/ksml-runner:1.2.0
      DOCKER_CONFIG_DIR: /ksml-configs          # path inside the provisioner container
      DOCKER_CONFIG_VOLUME: ksml-configs         # named volume shared with KSML containers
      DOCKER_CONFIG_NETWORKS: my-project_default # network where Kafka is reachable
    volumes:
      - ksml-configs:/ksml-configs
      - /var/run/docker.sock:/var/run/docker.sock  # Docker socket — provisioner manages sibling containers
    ports:
      - "8000:8000"

  kafka:
    image: apache/kafka:4.1.0
    # ...

Docker socket mount: The provisioner needs access to the Docker daemon socket (/var/run/docker.sock) to start and stop KSML containers. On Linux this is a bind mount from the host; on Docker Desktop for Mac the socket is available at the same path inside containers by default.

Named volume on Docker Desktop for Mac: Docker Desktop runs containers inside a Linux VM, so a plain host path bind-mount written by the provisioner container cannot be read by KSML sibling containers. Use a named volume (DOCKER_CONFIG_VOLUME) instead. The volume name: in the Compose definition must be a fixed string matching DOCKER_CONFIG_VOLUME — omit name: and Compose will prefix it with the project name, causing a mismatch.