fix-1828: started work - #2280
Conversation
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
|
@ryanjbaxter can you trigger copilot in here please? |
| } | ||
|
|
||
| private boolean isDeploymentReady(String deploymentName, String namespace) throws ApiException { | ||
| private boolean isDeploymentReady(String deploymentName, String namespace, int expectedReplicas) |
There was a problem hiding this comment.
in the new IT that I added, there is a need for two replicas, to really test the HA set-up
Signed-off-by: wind57 <eugen.rabii@gmail.com>
| envVars.add(new V1EnvVar().name("SPRING_CLOUD_KUBERNETES_SECRETS_ENABLED").value("TRUE")); | ||
|
|
||
| if (enableHa) { | ||
| envVars.add(new V1EnvVar().name("SPRING_CLOUD_KUBERNETES_LEADER_ELECTION_ENABLED").value("true")); |
There was a problem hiding this comment.
two properties are needed to enable HA
| } | ||
|
|
||
| /** | ||
| * <pre> |
There was a problem hiding this comment.
I've added a single IT, that goes through a cycle of leader / no leader / leader
| * | ||
| * @author wind57 | ||
| */ | ||
| sealed interface ConfigurationWatcherStateStore permits LeaseConfigurationWatcherStateStore { |
There was a problem hiding this comment.
this is the definition of the resource version store. Methods in this one are executed only by the leader
Signed-off-by: wind57 <eugen.rabii@gmail.com>
Signed-off-by: wind57 <eugen.rabii@gmail.com>
| } | ||
|
|
||
| @Override | ||
| public ConfigurationWatcherState readOrCreate() { |
There was a problem hiding this comment.
if HA is enabled and this is the first call, this will create an empty lease. Otherwise, it will read whatever is stored there.
We store in the spring.cloud.kubernetes.configuration.watcher/configmap-resource-version annotation, something like : "default=123,prod=456", so each namespace tracks its own resource version checkpoint.
We need such a store because during a downtime when there is no leader at all, resource version can progress ( meaning configmap is updated ), but since there is no leader, events can get lost. As such, we always increment and "store" ( via this implementation ) the most recent resource version we have observed. This happens in the handlers onAdd / onDelete / onUpdate.
So for example:
- we are now the leader and the resourceVersion is at
1. - we lose leadership, so our store stays at
1. - configmap progresses to
resourceVersion=2 - another leader is elected, it reads the store, sees that it holds
1 - starts the informer at
resourceVersion=1, so it can replay events it has not seen
| } | ||
|
|
||
| @Override | ||
| public void onApplicationEvent(ApplicationEvent event) { |
There was a problem hiding this comment.
these events are triggered only when the instance became the leader
| LOG.debug(() -> "Secret " + secret.getMetadata().getName() + " was deleted in namespace " | ||
| + secret.getMetadata().getNamespace()); | ||
| onEvent.accept(secret); | ||
| writeResourceVersion(secret); |
There was a problem hiding this comment.
after every update that we receive in the handler, write the resource version to the store
Signed-off-by: wind57 <eugen.rabii@gmail.com>
| } | ||
| } | ||
|
|
||
| public final void start(Map<String, String> storedResourceVersions, |
There was a problem hiding this comment.
With HA enabled:
-
@PostConstruct does nothing, so the informers are not started.
-
The native leader-election callback emits
StartLeadingEvent. -
ConfigurationWatcherHACoordinatorreceives that event. -
It calls
stateStore.readOrCreate()- If the HA state Lease does not exist, it creates it and returns an empty state.
- Otherwise, it reads the stored
ConfigMapandSecretresource versions.
-
The coordinator calls
start(...)on the available detectors, passing:- the stored resource versions
- a writer that persists future resource versions
-
Each detector creates its
InformerResourceVersionResolver. -
For each watched namespace, the informer sends its initial
LISTrequest:- if a stored checkpoint exists, the request starts from that resource version
- otherwise, it starts normally
-
The informer then starts its
WATCH. -
Resource events are handled normally:
- ConfigMap or Secret change is processed
- refresh or bus notification is triggered
- the processed resource version is written to the HA state Lease
-
Later
LIST/WATCHrequests use the resource version supplied by the informer, not the stored checkpoint. -
If the leader loses leadership, StopLeadingEvent is emitted.
-
The coordinator stops both detectors.
-
The next leader reads the persisted state and repeats the process from the stored checkpoints.
TL;DRThis PR adds high-availability support to the Kubernetes client-based Configuration Watcher. The Configuration Watcher can now run with multiple replicas while ensuring that only one replica actively watches ConfigMaps and Secrets at a time. If the leader fails, another replica acquires leadership, restores the persisted informer checkpoints, and continues processing changes. How It WorksHA requires both properties to be enabled: The two properties have separate responsibilities:
When HA is disabled, the existing behavior is unchanged. ConfigMap and Secret informers start during normal bean initialization. When HA is enabled:
The HA coordinator is initialized before the leader-election callbacks so that it is ready to receive leadership events. Persistent State The Configuration Watcher stores its state in a Kubernetes Lease. The leader-election lock and the Configuration Watcher state lease are separate resources. The state lease stores the last processed resource version independently for:
The values are stored in annotations on the lease. For example: The default configuration is: The lease is created automatically when the first leader starts, if it does not already exist. Resource-Version Replay
Repeated refresh notifications must therefore be tolerated by refresh targets.
|
No description provided.