Runtime Provisioner 0.8.0 Readme
This is a REST application used to provision KSML applications in Kubernetes.
When a deployment is requested, this provisioner fetches the KSML Helm Charts from the configured Helm registry, generates values.yaml based on the requested configuration and deploys it in the same Kubernetes cluster where this provisioner is deployed.
The Provisioner requires certain privileges scoped to the namespace in Kubernetes to deploy KSML pod including a ServiceMonitor resource for Prometheus monitoring.
Limitations
The Provisioner has the following limitations:
-
The Provisioner can only deploy KSML applications in the same Kubernetes cluster where it is deployed.
-
The Provisioner can pull charts from OCI-compatible registries only. Nexus registry is not supported.
-
KSML deployment names are generated as
{tenant}-{instance}-{environment}-{application}-ksml. Kubernetes limits label values to 63 characters, so if the combined name exceeds the allowed length, it is replaced by{tenant}-{instance}-{hash16}-ksmlwhere the hash is a 16-character SHA256-based prefix of the full name (e.g.mytenant-myinstance-c319f3cd9b7c02cc-ksml). The original identity is always preserved in the pod labels and can be used to locate the pod:kubectl get pods -n <namespace> \ -l axual.io/tenant=<tenant>,axual.io/instance=<instance>,axual.io/environment=<environment>,axual.io/application=<application>
Required Configuration
Set these environment variables before you start the Provisioner.
| Environment Variables | Description |
|---|---|
NAMESPACE |
The Kubernetes namespace where KSML application will be deployed. |
REGISTRY_URL |
The Helm Chart Registry URL where KSML Helm Charts are pulled from. |
CHART_NAME |
The name of the KSML Helm Chart. |
CHART_VERSION |
The version of the KSML Helm Chart. |
Optional Configuration
These environment variables change the default behaviour of the Provisioner.
| Environment Variables | Description |
|---|---|
REGISTRY_AUTH_ENABLED |
Set to |
REGISTRY_USERNAME |
Username of the Helm Chart Registry defined in
|
REGISTRY_PASSWORD |
Password of the Helm Chart Registry defined in
|
CUSTOM_VALUES_FILE |
Path to a custom values file containing
extra configurations for KSML deployment. This is useful when configs
like |
CLIENT_CA_FILE |
Path to a base64-encoded PEM file containing one or more CA certificates. It is used to validate the Helm Registry’s server certificate. |
KSML_ENABLED |
Set to |
CONNECT_ENABLED |
Set to |
CONNECT_NAMESPACES |
Namespaces to search for Kafka Connect worker pods, separated by commas. Defaults to the provisioner’s own namespace. Connect clusters usually run in another one. |
CONNECT_CONTAINER_NAME |
Names the Kafka Connect worker container. Only needed for a pod where the container cannot be found automatically. |
CONNECT_FILTER_WINDOW_PERCENT |
Raw lines read per worker
pod when a request filters by connector, as a percentage of the lines
asked for. Default |
CONNECT_FILTER_WINDOW_MAX_LINES |
Ceiling on that
window, in lines. Default |
CONNECT_DOWNLOAD_CAP_BYTES |
Most a whole-log download may
copy in total, across every worker. Default |
MAX_CONCURRENT_LOG_STREAMS |
How many log reads may run at
the same time, for KSML and Kafka Connect together. It counts pod
streams, not requests, because one Connect request reads every worker of
the cluster. Default |
INSECURE_SKIP_TLS_VERIFY |
If true, the validation of Helm Registry’s server certificate. This can lead to man in the middle attacks and should not be used in Production. |
DISTRIBUTED_TRACING_ENABLED |
If true, enables distributed
tracing with OpenTelemetry. To configure the exporter, see section
“Enable Distributed Tracing.” Default |
BASIC_AUTH_ENABLED |
Set to |
BASIC_AUTH_USERNAME |
Username of the basic authentication of HTTP server. |
BASIC_AUTH_PASSWORD |
Password of the basic authentication of HTTP server. |
SERVER_TLS_CERTIFICATE |
PEM encoded x509 certificate for HTTP server. |
SERVER_TLS_PRIVATE_KEY |
PEM encoded private key for HTTP server. |
Filtering and the read window
Kubernetes cannot filter log content. tailLines is applied to raw
lines by the kubelet, and the connector filter only runs afterwards,
inside the provisioner. So reading only as many lines as were asked for
returns almost nothing on a cluster where the wanted connector is a
small share of the traffic.
A filtered request is therefore read over a wider raw window:
raw lines per pod = min(tailLines * CONNECT_FILTER_WINDOW_PERCENT / 100, CONNECT_FILTER_WINDOW_MAX_LINES)
with the requested count as a floor, so the window is never narrower
than the answer needs. A request without a filter reads exactly
tailLines per pod, because every line read is a line returned.
The endpoint caps tailLines at 1000, so at the default percentage of
2000 the widest window is 1000 * 2000 / 100 = 20000 raw lines per
pod. The ceiling of 50000 sits above every value the endpoint can
produce at that percentage, and only binds once
CONNECT_FILTER_WINDOW_PERCENT is raised above 5000.
Raising the ceiling alone changes nothing.
How far back this reaches is the reason the ceiling is high:
how far back = raw lines read / lines the worker writes per second
Measured on a real worker writing 0.39 lines per second: 100 raw lines reached 1.6 minutes back, 2000 reached 33 minutes, 5000 reached 76 minutes, and the default maximum of 20000 reached about 14 hours. On a cluster at normal log levels, around 0.07 lines per second, 5000 reached about 20 hours. Raise the percentage for a quiet connector on a chatty cluster; lower it if reads feel slow.
Reading wide is cheap. Non-matching lines are dropped as they are read, so memory follows the answer rather than the window: the lines one request holds are bounded at 4 MiB across its worker pods. That allowance is split into a reserved floor per pod plus a shared pool, so the answer does not depend on which kubelet replied first, and one busy worker can still use everything the others did not reserve. The answer is serialised on top of that, so peak memory per request is a few times the 4 MiB rather than exactly it. Measured read times for one worker were 0.17 s for 5000 lines and 0.37 s for a whole 25 000 line log.
What the window cannot do is outlive the pod. Kubernetes keeps roughly
containerLogMaxSize * containerLogMaxFiles of log per container,
commonly 50 MiB, and deletes it with the pod. Nothing is stored
elsewhere.
Required Kubernetes Permissions
Runtime Provisioner requires a certain set of permissions on the service account to deploy KSML applications. These permissions are described below.
The required service account, role and role binding are automatically created by the Provisioner Helm Chart. No special configurations are required.
| Kubernetes Resource | API Group | Permissions |
|---|---|---|
configmaps |
GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE |
|
pods |
GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE |
|
pods |
metrics.k8s.io |
GET, LIST |
pods/log |
GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE |
|
secrets |
GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE |
|
serviceaccounts |
GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE |
|
services |
GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE |
|
statefulsets |
apps |
GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE |
jobs |
batch |
GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE |
prometheusrules |
monitoring.coreos.com |
GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE |
servicemonitors |
monitoring.coreos.com |
GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE |
ingresses |
networking.k8s.io |
GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE |
networkpolicies |
networking.k8s.io |
GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE |
persistentvolumeclaims |
storage.k8s.io |
GET, LIST, WATCH, CREATE, PATCH, UPDATE, DELETE |
Provisioning KSML applications in different namespace
By default, KSML applications will be provisioned in the same namespace
where the provisioner is deployed. If the applications need to be
deployed in a different namespace (eg ksml), pass below values.
env:
- name: NAMESPACE
value: "ksml"
rbac:
namespace: "ksml"
The NAMESPACE environment variable informs the provisioner about where
to deploy KSML applications. The rbac.namespace ensures that the
provisioner has the necessary permissions to deploy KSML applications in
the target namespace.
Bring your own ServiceAccount and RBAC
By default, the helm chart deploys a service account and associated RBAC resources with correct permissions. But for some reason, if you want to use your own service account and RBAC resources, set below configuration:
serviceAccount:
create: false
name: my-custom-sa
rbac:
create: false
A custom service account can be passed via serviceAccount.name. It is
your responsibility to ensure this service account has correct
permissions to deploy KSML applications (see previous section for exact
permissions). The easiest way to do this is to create a Role with
relevant permissions and a RoleBinding to associate the role with the
service account.
In most cases though, you don’t need this and can keep life simple by
having rbac.create: true and let the Helm chart take care of the
complexity.
Who can read which logs
The service does not check who is calling. Basic auth, the NetworkPolicy and Platform Manager’s group check are the controls around it.
route.enabled=true and ingress.enabled=true publish the service
outside the cluster, where anything that reaches it can deploy and
undeploy KSML applications and read these logs. Turn
BASIC_AUTH_ENABLED on before you enable either, because it
defaults to false.
CONNECT_NAMESPACES is the boundary. The chart grants read
permission only for those namespaces, so the service cannot read one
that is not listed. Tenants that must not see each other’s logs need
separate namespaces, or their own provisioner. The address is stored per
Connect cluster, so a second provisioner needs no extra work.
Prometheus metrics
Served on port 2112 and scraped by the ServiceMonitor in the chart.
Reading Kafka Connect logs adds four metrics, each labelled tenant,
instance and connect_cluster:
| Metric | Type | Meaning |
|---|---|---|
|
counter |
Lines taken from the worker pods, before filtering |
|
counter |
Lines removed by the connector filter |
|
counter |
Reads
that ended with an error, with a |
|
gauge |
Streams open right now |
read minus filtered is what reached the caller. These are not the
GET /metrics endpoint below, which reports CPU and memory of one KSML
deployment.
Enable Distributed Tracing
Runtime Provisioner supports distributed tracing with OpenTelemetry. To
enable, set environment variable DISTRIBUTED_TRACING_ENABLED
to true.
To export traces to an OTel backend, configure below environment variables.
| Environment Variable | Description |
|---|---|
OTEL_EXPORTER_OTLP_ENDPOINT |
Endpoint of OTel collector. Must start with http/https and include port number. |
OTEL_EXPORTER_OTLP_HEADERS |
Optional headers. |
OTEL_SERVICE_NAME |
A unique name to recognize the service. |
Enable Basic Authentication on the HTTP server
Basic authentication can be enabled by supplying the below configuration:
| Environment Variable | Description |
|---|---|
BASIC_AUTH_ENABLED |
Set to “true” to enable Basic authentication. Default is “false”. |
BASIC_AUTH_USERNAME |
Username for Basic authentication. Required when BASIC_AUTH_ENABLED is “true”. |
BASIC_AUTH_PASSWORD |
Password for Basic authentication. Required when BASIC_AUTH_ENABLED is “true”. |
Below is an example configuration:
env:
- name: BASIC_AUTH_ENABLED
value: "true"
- name: BASIC_AUTH_USERNAME
value: "username"
- name: BASIC_AUTH_PASSWORD
value: "password"
Enable TLS on the HTTP server
TLS can be enabled by supplying the below configuration:
| Environment Variable | Description |
|---|---|
SERVER_TLS_CERTIFICATE |
Certificate in PEM format. |
SERVER_TLS_PRIVATE_KEY |
Private key in PEM format. |
Below is an example configuration:
env:
- name: SERVER_TLS_CERTIFICATE
valueFrom:
secretKeyRef:
key: tls.crt
name: runtime-provisioner
- name: SERVER_TLS_PRIVATE_KEY
valueFrom:
secretKeyRef:
key: tls.key
name: runtime-provisioner
Note that in the above example, the secret runtime-provisioner must
exist with certificate in tls.crt and private key in tls.key.
An OpenShift Route created with the default
route.tls.termination: passthrough hands the TLS session straight to
the pod, so it needs both variables above. Without them the pod serves
plain HTTP and the handshake fails with no error on the Route itself.
Set route.tls.termination: edge instead if you want the router to
terminate TLS.
Setting fsGroup for KSML PersistentVolume Access
When using a PersistentVolume (PV) for KSML Stateful feature, the
volume is mounted at /ksml-store path which is owned by the root user.
Without additional configuration, the KSML process running in the Pod
cannot create the required subfolder /ksml-store/data.
To ensure the KSML process can write to the volume, set
podSecurityContext.fsGroup so that the files in the volume are owned
by the group of the user running the KSML process. The default user is
part of group 0 so set fsGroup to 0 unless you are using a different
user. For example:
runtime-provisioner:
customValues:
podSecurityContext:
fsGroup: 0
This ensures that the KSML Pod can create and manage the
/ksml-store/data folder inside the PV automatically.
REST API endpoints
The Provisioner exposes the following HTTP endpoints.
POST /deploy
Request Body:
{
"tenant": "dizzl",
"instance": "sandbox",
"environment": "dev",
"application": "sensor-inspector",
"replicas": 1,
"definition": "{streams: {sensor_source_avro: {topic: ksml_sensordata_avro, keyType: string, valueType: 'avro:SensorData'}}, functions: {log_message: {type: forEach, parameters: [{name: format, type: string}], code: 'log.info(\"Consumed {} message - key={}, value={}\", format, key, value)'}}, pipelines: {consume_avro: {from: sensor_source_avro, forEach: {code: 'log_message(key, value, format=\"AVRO\")'}}}}",
"kafkaProperties": {
"bootstrap.servers": "broker:9093",
"schema.registry.url": "http://schemaregistry:8081",
"application.id": "io.axual.sensor-inspector"
},
"schemas": [
{
"schemaFileName": "SensorData.avsc",
"schemaBody": "..."
}
]
}
Returns 200 OK if deployment is successful.
GET /status
Returns the status of KSML deployment(s). This endpoint supports two modes:
Detailed Status (single deployment)
Query Parameters: - tenant (required): The tenant name - instance
(required): The instance name - environment (required): The
environment name - application (required): The application name -
deploymentMode (optional): Either StatefulSet or Job
Example:
GET /status?tenant=dizzl&instance=sandbox&environment=dev&application=sensor-inspector
Returns the status of a specific KSML deployment including Helm and Kubernetes details.
The status can be one of Running, Starting, Failing, or
Completed (for Jobs).
{
"helm": {
"name": "dizzl-sandbox-dev-sensor-inspector-ksml",
"chart": "ksml",
"version": "0.0.0-snapshot",
"appVersion": "snapshot",
"namespace": "default",
"revision": 1,
"status": "deployed",
"deployedAt": "2024-08-26T13:04:20Z"
},
"kubernetes": {
"name": "dizzl-sandbox-dev-sensor-inspector-ksml",
"status": "Running",
"replicas": 1,
"readyReplicas": 1,
"restartsCount": 3,
"lastRestart": "2024-08-26T14:01:05Z"
}
}
Bulk Status (all deployments for tenant/instance)
Query Parameters: - tenant (required): The tenant name - instance
(required): The instance name
Example: GET /status?tenant=dizzl&instance=sandbox
Returns the status of all KSML deployments for the given tenant and
instance combination. This searches for pods with matching
axual.io/tenant and axual.io/instance labels. Each status includes
resource limits (CPU in cores, memory in MB) from the KSML container.
{
"statuses": [
{
"name": "dizzl-sandbox-dev-sensor-inspector-ksml",
"status": "Starting",
"replicas": 1,
"readyReplicas": 0,
"resourceLimits": {
"cpuCores": "1",
"memoryMB": "4096"
}
},
{
"name": "dizzl-sandbox-prod-sensor-inspector-ksml",
"status": "Running",
"replicas": 1,
"readyReplicas": 1,
"resourceLimits": {
"cpuCores": "0.5",
"memoryMB": "2048"
}
}
],
"tenant": "dizzl",
"instance": "sandbox"
}
GET /logs?tenant=dizzl&instance=sandbox&environment=dev&application=sensor-inspector&tailLines=10
Returns or streams logs for the KSML deployment. The response format is
determined by the Accept request header.
Query Parameters: - tenant (required): The tenant name - instance
(required): The instance name - environment (required): The
environment name - application (required): The application name -
deploymentMode (optional): Either StatefulSet or Job - tailLines
(optional): Number of log lines to return or replay at stream start
(1–1000, default 100)
Batch mode (default — no Accept header, or
Accept: application/json):
Returns the last N lines of logs for the KSML deployment as a single JSON response.
Sample response:
{
"tailLines": 1,
"logs": [
{
"replicaName": "0",
"lines": "2024-07-16T13:06:39,394Z [io.axual.sensor-inspector-6a395bc1-f3bc-4eb2-9331-c3d1b4e45f8b-StreamThread-1] INFO ksml.function.log_message - Consumed AVRO message - key=sensor27, value={'city': 'Xanten'}\n"
},
{
"replicaName": "1",
"lines": "2024-07-16T13:06:39,841Z [io.axual.sensor-inspector-1393ae07-9321-424d-9453-b86383c816b3-StreamThread-1] INFO ksml.function.log_message - Consumed AVRO message - key=sensor28, value={'city': 'Amsterdam'}\n"
}
]
}
Streaming mode (Accept: application/x-ndjson):
Streams real-time logs as NDJSON over a chunked HTTP response. The connection stays open and each log line is forwarded to the client as it arrives, making it suitable for live log tailing. A keepalive message is emitted every 15 seconds on idle connections to prevent intermediate proxies from closing the connection.
Each line in the response body is a self-contained JSON object:
| Object type | Field | Description |
|---|---|---|
Log line |
|
A single log line from the pod |
Keepalive |
|
Empty string; sent every 15 s to keep the connection alive |
Error |
|
Emitted when the log stream terminates with an error |
Sample response (each line is a separate JSON object):
{"line":"2024-07-16T13:06:39,394Z [StreamThread-1] INFO ksml.function.log_message - Consumed AVRO message - key=sensor27"}
{"line":"2024-07-16T13:06:39,841Z [StreamThread-1] INFO ksml.function.log_message - Consumed AVRO message - key=sensor28"}
{"keepalive":""}
{"line":"2024-07-16T13:06:54,102Z [StreamThread-1] INFO ksml.function.log_message - Consumed AVRO message - key=sensor29"}
{"error":"Pod CrashLoopBackOff: exit code 1"}
GET /metrics?tenant=dizzl&instance=sandbox&environment=dev&application=sensor-inspector&deploymentMode=StatefulSet
Returns the current CPU and memory usage metrics for the KSML deployment. This endpoint requires the Kubernetes Metrics Server to be installed in the cluster.
Query Parameters: - tenant (required): The tenant name - instance
(required): The instance name - environment (required): The
environment name - application (required): The application name -
deploymentMode (optional): Either StatefulSet or Job
Sample response:
{
"cpu": {
"usedMillicores": "250"
},
"memory": {
"usedMb": "512"
}
}
Note: If the Metrics Server is not installed or metrics are not yet available for a newly started pod, the endpoint will return a 404 error with an appropriate message.
POST /undeploy
Request Body:
{
"tenant": "dizzl",
"instance": "sandbox",
"environment": "dev",
"application": "sensor-inspector"
}
Returns 200 OK if removal of deployment is successful
GET /kafkaconnect/logs?tenant=dizzl&instance=sandbox&connectCluster=payments&tailLines=10&connector=payments-sink
Reads the worker pods of one Kafka Connect cluster and merges their logs
into one answer. Pods are found by the axual.io/tenant,
axual.io/instance and axual.io/connect-cluster labels, so
connectCluster must match the label, not the Helm release name.
tailLines is how many lines of history the answer holds in total,
across the workers, from 0 to 1000. 0 means live lines only, which
only a stream can serve. It bounds the answer and not the read: every
worker is read over the whole total, and the merge then keeps the newest
tailLines across the cluster, so one busy worker can supply the whole
answer. It means the same in both modes, so a stream shows the same
window a batch request would answer with, within the one limit described
under streaming below.
connector keeps only the lines naming that connector. A JSON entry is
matched on its connector.context field, which is exact. Any other line
falls back to looking for the name in the text, which is still needed:
Kafka’s own producer and consumer threads log outside the task’s logging
context and name the connector in their client id instead, and 7 lines
in 618 on a live worker were of that shape. Either way it runs after the
read, so the read window is widened first, see “Filtering and the read
window”.
There is no way to read the log of a container that has already stopped.
That was tried and taken out: the node keeps the old log on its own
schedule, so it was usually gone, and a worker killed outright writes no
reason into its log anyway. To look at a crashed worker use
kubectl logs --previous and kubectl describe pod, which also shows
the exit code.
minLevel keeps only the lines at that severity or above, from TRACE,
DEBUG, INFO, WARN, ERROR and FATAL. Any other value is a 400.
It widens the read window for the same reason connector does,
including when no connector is given: warnings and errors are 1.4% of a
worker’s log, measured over 5995 lines, so a window sized to the request
comes back nearly empty.
A kept line keeps the lines that belong to it. Two thirds of a worker’s
lines are indented frames under a Java error and carry no level of their
own, so a filter judging each line alone would keep
Error while starting connector and drop the ninety lines explaining
why.
A line whose own severity cannot be read follows the line above it, exactly as an indented frame does, because that is what such a line nearly always is. Keeping it regardless would leak everything a worker prints before its logging starts, and that is not a small amount: the first few hundred lines of a Kafka Connect worker are shell tracing and a dumped configuration file, none of which carry a level. A reader who asked for errors would get those instead of an empty window saying nothing matched. A blank line follows the line above it for the same reason.
So a log format this reader cannot classify at all disappears under a threshold. That is the honest outcome and it explains itself, because the empty window says how far it looked and which filter emptied it, and both layouts the chart writes always carry a level.
The severity is read from either log format: the log.level field of a
JSON entry, or a level word near the start of a plain-text line. Both
are what the chart’s two layouts produce.
Both stay supported even though kafka-connect-helm 0.5.0 writes JSON
and no longer offers a choice. A cluster upgraded to that chart keeps
hours of plain text in the same file until it rotates out, so one answer
routinely mixes the two, and a cluster on an older chart, or one this
chart did not create, can send plain text indefinitely. Note that the
KSML endpoint is a different reader entirely: it streams lines without
parsing them, so KSML has no severity filter and is unaffected by any of
this.
Answers with one JSON document by default:
{
"tailLines": 10,
"entries": [
{ "line": "...", "origin": "kc-alpha-kafka-connect-connect-0", "timestamp": "2026-09-01T08:54:10.092028739Z" },
{ "line": "...", "origin": "kc-alpha-kafka-connect-connect-1", "timestamp": "2026-09-01T08:54:10.301114000Z" }
],
"linesRead": 4000,
"linesReturned": 10,
"truncated": false,
"degraded": ["kc-alpha-kafka-connect-connect-2"]
}
entries is the answer, in time order, oldest first. linesRead counts
every raw line taken off the workers and linesReturned how many the
answer holds, which is what tells “the cluster has nothing more” apart
from “plenty was read and none of it matched”. truncated means the
byte budget cut the answer short. degraded names the worker pods that
could not be read at all, so the answer covers only the rest of the
cluster.
Send Accept: application/x-ndjson to stream instead. Each line is then
one JSON object, and every log line carries origin, the pod it came
from, and timestamp, when the node stamped it:
{"line":"2026-08-11 13:21:42 INFO ...","origin":"kc-alpha-kafka-connect-connect-0","timestamp":"2026-08-11T13:21:42.104512Z"}
{"event":"streamEnded","origin":"kc-alpha-kafka-connect-connect-1","line":"--- worker ... ended ---"}
{"keepalive":""}
A stream sends the history first and then follows the log. Both come
from one read per worker, opened with the window and follow
together, so the kubelet seeks back over the window and then keeps
following the same file. So a client needs one request, not two.
| Request | What arrives |
|---|---|
|
history only, one document |
|
the same window, line by line, then new output as it is written |
|
new output only, no history |
Do not fetch the history with a second batch request. Two requests read
two different windows that overlap at the recent end, and the client
then receives the same lines from both with no way to tell them apart. A
client that already holds the history opens the stream with
tailLines=0, which asks for new output only. Platform Manager sends a
fixed tailLines on its own stream, so through Platform Manager a
stream always carries history: ask it once and do not fetch the history
yourself.
The kubelet’s join is exact. Where the history ends is our judgement.
One read cannot lose a line between the window and the new output, and
cannot send the same line twice. But the stream carries no marker
between the two, so the provisioner decides where the history stops:
when the window has been read, when a line stamped later than the
request arrives, when the read falls silent for 400 ms, or after 5
seconds, whichever comes first. Lines that miss the merged block are
still sent, straight after it, in their own worker’s order. So a busy
cluster can show a few lines out of order at the seam, and a stream
forced to end its history early can send more lines than tailLines
asked for. No line is lost, and none arrives twice.
That last point is the one difference between the two modes: tailLines
bounds a batch answer exactly, while on a stream it bounds the merged
block and the rest of the window still follows it.
The history lines arrive in time order across the workers, exactly as
entries does, and each carries origin and timestamp, so a stream
describes a line the same way the batch answer does. timestamp is
absent on a line the node never stamped, which is a stack trace frame or
another continuation of the line above it. New output cannot be ordered
across workers at all, because holding lines back to sort them would
delay a live tail, so each worker’s new lines are sent as they arrive.
A stream stays open while at least one worker still sends lines, and closes after 30 minutes.
When a filtered stream reads a wide window and matches nothing, the
silent-stream notice says so, for example
No lines for this connector in the last 4000 lines read from its workers. Waiting for output from the application….
A reader can then tell a quiet connector from a window that did not look
far enough back.
The notice names the filter that emptied the window, because a reader
told the wrong one lifts the wrong control and still sees nothing. With
minLevel alone it reads
No lines at WARN or above in the last 4000 lines read from the workers. …,
and with both filters
No lines for this connector at WARN or above in the last 4000 lines read from its workers. ….
A minLevel of TRACE keeps every level, so it is not named, exactly
as it does not widen the read window.
When every worker returns fewer raw lines than the window asked for, the
read reached the start of the log the node still has, and the notice
says that instead:
No lines for this connector in the whole log the workers still have, 340 lines. …,
or without a filter
This is the whole log the workers still have, 340 lines. …. That is
the difference between “ask for more” and “there is no more”, which
a line count alone does not show. The node’s retention is the wall no
request can pass: the kubelet keeps containerLogMaxSize per file and
containerLogMaxFiles of them, 10Mi and 5 by default.
When the plain-text level reader can be removed
The reader that finds a level in a plain-text line exists for three reasons, and all three end:
-
A cluster upgraded to
kafka-connect-helm0.5.0 keeps plain-text history in the same log file until it rotates out. One filtered answer during testing came back with 3 JSON lines and 95 plain ones. -
A cluster still on an older chart has not switched yet.
-
A Kafka Connect cluster this chart did not create can send anything.
So it becomes safe to delete once every cluster runs 0.5.0 or later and their log files have rotated. Written down because otherwise it is either deleted too early or kept for ever by default.
The KSML endpoint never reaches this reader: it streams lines without parsing them. Platform Manager’s and the console’s ability to show a plain-text line is a separate matter and must stay, because KSML sends plain text permanently.
A very long connector label
Most logging contexts are short, 21 to 31 bytes, such as
[payments-sink|task-0]. A connector that copies between
two clusters produced a 98-byte one, because it carries a whole client
id:
[AdminClient clientId=src-payments-sink->dst-payments-sink|payments-sink|heartbeats-target-admin]
It appeared on 4 lines out of thousands, so it costs nothing today. It could matter if a customer runs many cross-cluster connectors and one logs heavily, since the field is on every line and counts against the read window as well as disk. The fix, if it is ever needed, is a shorter field such as the connector name alone. Do not remove the field: it is what the connector filter matches on.
The three numbers that decide when the history stops
The split between history and new output is the provisioner’s own judgement, so three constants decide it. They are constants on purpose: no cluster has needed a different value, and a setting is a contract to keep.
| Constant | Value | What it does |
|---|---|---|
|
400 ms |
per worker: how long a worker’s read may say nothing before its history counts as complete. The replay arrives far faster than a worker writes, so the pause between them is wide |
|
5 s |
a clock bound on the phase, per worker and again across the workers, because a worker that keeps writing never falls silent and a log shorter than the window never reaches its edge. The second use is what stops one silent worker holding the merged block back from all of them |
|
1 s |
per worker: how far a worker node’s clock may run ahead of ours before a replayed line looks like new output. The kubelet stamps lines on the node, so that comparison crosses two clocks |
Reading any of them too early costs the merge for the rest of the
replay, or a little latency. None of them can lose or repeat a line. If
a cluster ever needs different values, they follow the
CONNECT_FILTER_WINDOW_PERCENT pattern: an environment
variable, a chart value and a row in the settings table.
Answers 404 when no worker pod matches the labels, 400 for a bad parameter, 503 when the stream limit is full, and 501 in Docker mode, where there is no Kubernetes API to read.
GET /kafkaconnect/logs/download?tenant=dizzl&instance=sandbox&connectCluster=payments
Sends every worker’s current log file as one plain-text attachment, in
labelled sections (===== worker <pod> =====). It shares
nothing with the endpoint above: no filtering, no merging, no line
budget. Bytes go from the Kubernetes stream to the socket in 64 KiB
chunks, so the provisioner’s memory does not grow with the log.
connector, minLevel and tailLines are ignored, because reading
everything is the point.
What it can return is bounded by the container runtime, not by this
service. Kubernetes serves the container’s current log file only.
Rotated files stay on the node and no API reaches them. Measured on the
test cluster: a container that had written 59 MiB held 83 MB across five
files on disk, docker logs on the node returned all 62 MB of it, and
the Kubernetes API returned 2.4 MB, the current file. So this endpoint
returns one rotation window, and how much that is depends on when the
last rotation happened.
For scale, the runtime’s rotation size is the ceiling: 10Mi with the
kubelet’s containerLogMaxSize default, 20MB with the Docker json-file
driver’s max-size. Reading a whole file measured about 19 MiB/s
through the Kubernetes API, so a few workers take a few seconds.
CONNECT_DOWNLOAD_CAP_BYTES caps the total across every
worker, 64 MiB by default. It is a safety net rather than a feature
limit: a real cluster never comes near it, and it exists for the case
where somebody configures a much larger rotation size. Hitting it still
produces a valid file, with a final line saying it was truncated and how
many workers were copied.
Kubelet timestamps are deliberately off. The worker writes one JSON
object per line and a timestamp prefix would stop the file being
readable as NDJSON; every entry already carries its own @timestamp.
A worker that cannot be read is named in the file and the others are still copied. If no worker can be read the answer is a status code, not a file whose every line is an error message.
Docker Runtime Mode (Docker Compose only)
Not for production use. The Docker runtime is intended exclusively for local development and testing via Docker Compose. It has no high-availability, no persistent storage management, and no Kubernetes RBAC isolation. Do not run it in any shared or production environment.
The Runtime Provisioner can run in Docker mode by setting
DEPLOYMENT_TARGET=docker. In this mode it manages KSML
applications as plain Docker containers on the local Docker daemon
instead of deploying Helm charts into Kubernetes.
How it works
When a /deploy request arrives the provisioner: 1. Writes a
ksml-runner.yaml configuration file (with inline Kafka settings,
schema registry config, and KSML definition) into a per-container
directory on the host. 2. Starts a Docker container from
KSML_IMAGE, mounting that directory so the KSML runner can read
its config. 3. Attaches the container to any Docker networks listed in
DOCKER_CONFIG_NETWORKS so it can reach Kafka brokers and
schema registries running in the same Compose project.
On /undeploy the container is stopped and removed. Config files
written for that container are also cleaned up.
The /logs and /status endpoints work the same way as in Kubernetes
mode. The /metrics endpoint is not supported in Docker mode and always
returns 404.
Required configuration
Set these environment variables to run the Provisioner in Docker mode.
| Environment Variable | Description |
|---|---|
|
Must be set to |
|
Docker image reference for the KSML runner
(e.g. |
|
Absolute path on the host where per-container config files are written and bind-mounted into KSML containers. |
Optional configuration
These environment variables change the default behaviour in Docker mode.
| Environment Variable | Description |
|---|---|
|
Name of a Docker named volume to mount instead of a bind mount. Required on Docker Desktop for Mac, where host bind-mounts from inside a container are not supported. See note below. |
|
Comma-separated list of Docker
network names to attach KSML containers to
(e.g. |
Docker Compose setup
Below is a minimal example showing the provisioner alongside Kafka. The config directory is shared as a named volume so it works on both Linux and Docker Desktop for Mac.
volumes:
ksml-configs:
name: ksml-configs # fixed name — must match DOCKER_CONFIG_VOLUME
services:
runtime-provisioner:
image: registry.example.com/runtime-provisioner:0.8.0
environment:
DEPLOYMENT_TARGET: docker
KSML_IMAGE: registry.example.com/ksml-runner:1.2.0
DOCKER_CONFIG_DIR: /ksml-configs # path inside the provisioner container
DOCKER_CONFIG_VOLUME: ksml-configs # named volume shared with KSML containers
DOCKER_CONFIG_NETWORKS: my-project_default # network where Kafka is reachable
volumes:
- ksml-configs:/ksml-configs
- /var/run/docker.sock:/var/run/docker.sock # Docker socket — provisioner manages sibling containers
ports:
- "8000:8000"
kafka:
image: apache/kafka:4.1.0
# ...
Docker socket mount: The provisioner needs access to the Docker daemon socket (
/var/run/docker.sock) to start and stop KSML containers. On Linux this is a bind mount from the host; on Docker Desktop for Mac the socket is available at the same path inside containers by default.
Named volume on Docker Desktop for Mac: Docker Desktop runs containers inside a Linux VM, so a plain host path bind-mount written by the provisioner container cannot be read by KSML sibling containers. Use a named volume (
DOCKER_CONFIG_VOLUME) instead. The volumename:in the Compose definition must be a fixed string matchingDOCKER_CONFIG_VOLUME— omitname:and Compose will prefix it with the project name, causing a mismatch.