Schema Distribution

This guide explains what Schema Distribution copies and why, why it applies only to the legacy registry, and what to do instead when the platform uses Apicurio Registry.

Type

Explanation

Goal

Understand whether your deployment needs Schema Distribution at all.

Audience

Platform Operators and architects planning schemas across a multi-cluster instance.

When to use

Read before enabling the Schema Distributor, and when moving from the legacy registry to Apicurio Registry.

Schema Distribution is a legacy feature. It only works with the Confluent-based Legacy Schema Registry, because it copies that registry’s Kafka schemas topic from one cluster to the others.

It does not work with Apicurio Registry, which is the registry Axual now ships. If you use Apicurio Registry, do not enable the Schema Distributor. See Schema Distribution with Apicurio.

Contents

About Schema Distribution

The Confluent-based Legacy Schema Registry holds the schema definitions for topics on the Kafka cluster. A producer can then write a record carrying the schema identifier from the registry rather than the full schema. Schema Distribution keeps the same identifier for a schema on every cluster, so a consumer finds the correct schema even when the producer wrote the record on a different Kafka cluster.

The Schema Distributor does this by reading the schemas topic the legacy registry uses, and distributing the contents to the same topic on all other Kafka clusters. This introduces two requirements:

  • Schemas are registered on a single cluster only.

  • The schema registries on all other clusters run as read-only registries.

The Kafka cluster with registration enabled is the primary cluster.

Schema Distribution is started on Axual Distributor connected to the primary cluster, and distributes the contents of the schemas topic to the remote clusters.

Schema Distribution with Cluster 1 as primary cluster
Figure 1. Schema Distribution with Cluster 1 as primary cluster

Internally the registry stores each schema under a monotonically increasing id, keyed by subject name. Registering a schema that is already present under another subject maps the new subject to the existing id rather than creating a second one. Axual adds a layer on top of that, because the topic name a tenant sees in Self-Service is mapped to a technical name, so mytopic is stored as tenant-instance-environment-mytopic.

Only the primary cluster accepts registrations, because it is the cluster the instance API runs on. Letting two clusters assign ids independently would produce the same schema under different ids.

The Schema Distributor is configured with its own subject format and the subject format of every target cluster, and translates between them as it distributes, so the same schema keeps the same id everywhere.

The subjects in the legacy registry are bound to the topic name, often in the form of <topic-name>-key and <topic-name>-value for the key and value schemas. When different clusters use different topic naming patterns, the Schema Distributor transforms the subject names to match the naming pattern of the target cluster.

Schema Distribution with Apicurio

Apicurio Registry does not keep its schemas in a plain Kafka topic that can be copied between clusters, so Kafka cannot replicate them and the Schema Distributor does not work with it.

For a multi-cluster setup with Apicurio Registry, run one shared Apicurio Registry instance that every Kafka cluster can reach. All clusters then use the same schemas with the same schema ids, so no schema distribution is needed. This is the only supported way to share schemas across clusters with Apicurio Registry today.

What you must do: give every cluster that uses this shared Apicurio Registry the same topic naming pattern. This is the topic.pattern setting on clients and the topicPattern setting in the Axual Distributor config. Use one value everywhere, for example {tenant}-{instance}-{environment}-{topic}.

Why: Apicurio Registry saves each schema under a name called the subject. Axual builds the subject from the real (technical) Kafka topic name, so the same topic must get the same technical name on every cluster. No Schema Distributor changes names between clusters here, so different patterns give one topic a different subject on each cluster, and reading or registering schemas fails.

This rule is only about sharing schemas. Message distribution still works when clusters use different patterns, because each record carries its schema id and the consumer finds the schema by that id.

Do not enable schemaDistributor in the Axual Distributor Helm chart when you use Apicurio Registry. It would only create an unused schemas topic and access control entries, and would distribute no schemas.

Enabling the Schema Distributor

The schemaDistributor values are listed in schemaDistributor, and Start the connectors in the right order places it first among the connectors, on the primary cluster only.