diff --git a/etcd/README.md b/etcd/README.md index b92d553df7e..871c190ecc0 100644 --- a/etcd/README.md +++ b/etcd/README.md @@ -40,10 +40,12 @@ The brief window where status is empty is acceptable since the healthcheck contr A **pacemaker resource** is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. -For Two Node OpenShift with Fencing, we manage three resource types: +For Two Node OpenShift with Fencing, the node's `resources` array tracks: - **Kubelet**: The Kubernetes node agent and a prerequisite for etcd - **Etcd**: The distributed key-value store -- **FencingAgent**: Used to isolate failed nodes during a quorum loss event (tracked separately) +- **TaintAlertAgent** / **UntaintAlertAgent**: Alert agents whose CIB configuration and per-node script presence are tracked as optional entries + +Fencing agents are tracked separately in each node's `fencingAgents` array, not in `resources`. ### Status Structure @@ -69,7 +71,7 @@ status: # Optional on creation, populated via status subresou - type: Member - type: FencingAvailable - type: FencingHealthy - resources: # Required: Pacemaker resources on this node (min 2) + resources: # Required: Pacemaker resources on this node (min 2: Kubelet + Etcd) - name: Kubelet # Both Kubelet and Etcd must be present conditions: # Required: Resource-level conditions (min 8 items) - type: Healthy @@ -82,6 +84,13 @@ status: # Optional on creation, populated via status subresou - type: Schedulable - name: Etcd conditions: [...] # Same 8 conditions as Kubelet (abbreviated) + - name: TaintAlertAgent # Optional: alert-agent entry + conditions: # Required: 3 conditions for alert-agent entries + - type: Healthy + - type: Enabled # reason: ScriptConfigured (agent registered in CIB) + - type: Operational # reason: ScriptPresent (script present on this node) + - name: UntaintAlertAgent # Optional + conditions: [...] # Same 3 conditions as TaintAlertAgent (abbreviated) fencingAgents: # Required: Fencing agents for THIS node (1-8) - name: # e.g., "master-0_redfish" (unique, max 300 chars) method: # Fencing method: "Redfish" or "IPMI" @@ -104,6 +113,27 @@ Unlike regular pacemaker resources (Kubelet, Etcd), fencing agents are tracked s - **FencingAvailable**: True if at least one agent is healthy (fencing works), False if all agents unhealthy (degrades operator) - **FencingHealthy**: True if all agents are healthy (ideal state), False if any agent is unhealthy (emits warning events) +### Alert Agents + +Alert agents are pacemaker alert handlers used to automatically taint a node after it is fenced, and remove +that taint once the node rejoins the cluster. There are two known alert agents: `TaintAlertAgent` and +`UntaintAlertAgent`. An alert agent is registered in the CIB (a single object shared by all nodes via +Pacemaker's CIB replication), and the script it invokes must exist locally on whichever node the triggering +event occurs on, since Pacemaker executes it there. Script delivery is handled independently per node by MCO. + +Both aspects are tracked as optional entries in each node's `resources` array, reusing the `PacemakerClusterResourceStatus` type. An alert agent is not a +pacemaker-managed resource, so only a subset of the resource conditions applies to it: +- `Enabled` (reason `ScriptConfigured`) - the alert agent is registered in the CIB as expected +- `Operational` (reason `ScriptPresent`) - the alert agent's script is present and executable on this node +- `Healthy` - aggregate of the two conditions above + +The name-gated validation on `PacemakerClusterResourceStatus` requires all eight conditions only for the `Kubelet` and `Etcd` resources, so alert-agent entries carry just these three. + +Alert-agent health is deliberately **not** folded into the node or cluster `Healthy` aggregates, since a missing or misconfigured alert agent doesn't reflect etcd/kubelet state. It is surfaced by the cluster-etcd-operator as an error that degrades the operator: a broken alert agent means post-fencing taint/untaint automation is broken, which is actionable on its own. + +The alert-agent resource entries are optional. They are omitted, or reported `Unknown`/`Pending`, when this +status has not yet been collected by the status collector, including by a collector version that predates them. + ### Cluster-Level Conditions | Condition | True | False | @@ -128,7 +158,8 @@ Unlike regular pacemaker resources (Kubelet, Etcd), fencing agents are tracked s ### Resource-Level Conditions -Each resource in the `resources` array and each fencing agent in the `fencingAgents` array has its own conditions. +Each pacemaker-managed resource (`Kubelet`, `Etcd`) in the `resources` array and each fencing agent in the +`fencingAgents` array has all eight of the following conditions. | Condition | True | False | |-----------|------|-------| @@ -141,6 +172,15 @@ Each resource in the `resources` array and each fencing agent in the `fencingAge | `Started` | Resource is started (`Started`) | Resource is stopped (`Stopped`) | | `Schedulable` | Resource is schedulable (`Schedulable`) | Resource is not schedulable (`Unschedulable`) | +Alert-agent entries (`TaintAlertAgent`, `UntaintAlertAgent`) reuse this type but populate only `Healthy`, +`Enabled`, and `Operational`, with alert-agent-specific reasons: + +| Condition | True | False | Unknown | +|-----------|------|-------|---------| +| `Healthy` | Alert agent entry is healthy (`ResourceHealthy`) | Alert agent entry has issues (`ResourceUnhealthy`) | Not yet observed | +| `Enabled` | Alert agent registered in the CIB as expected (`ScriptConfigured`) | Not registered or misconfigured (descriptive reason) | Not yet observed (`Pending`) | +| `Operational` | Alert agent script present on this node (`ScriptPresent`) | Script missing from this node (descriptive reason) | Not yet observed (`Pending`) | + ### Validation Rules **Resource naming:** @@ -170,8 +210,8 @@ Each resource in the `resources` array and each fencing agent in the `fencingAge **Status fields:** - `status` - Optional on creation (pointer type), populated via status subresource -- When status is present, all fields within are required: - - `conditions` - Required array of cluster conditions (min 3 items) +- When status is present, `conditions`, `lastUpdated`, and `nodes` are required: + - `conditions` - Required array of cluster conditions (min 3 items: Healthy, InService, NodeCountAsExpected) - `lastUpdated` - Required timestamp for staleness detection - `nodes` - Required array of control-plane node statuses (min 0, max 5; empty allowed for catastrophic failures) @@ -179,20 +219,22 @@ Each resource in the `resources` array and each fencing agent in the `fencingAge - `nodeName` - Required, RFC 1123 subdomain - `addresses` - Required (min 1, max 8 items) - `conditions` - Required (min 9 items with specific types enforced via XValidation) -- `resources` - Required (min 2 items: Kubelet and Etcd) +- `resources` - Required (min 2 items: Kubelet and Etcd; may also contain optional TaintAlertAgent / UntaintAlertAgent entries) - `fencingAgents` - Required (min 1, max 8 items) **Conditions validation:** - Cluster-level: MinItems=3 (Healthy, InService, NodeCountAsExpected) - Node-level: MinItems=9 (Healthy, Online, InService, Active, Ready, Clean, Member, FencingAvailable, FencingHealthy) -- Resource-level: MinItems=8 (Healthy, InService, Managed, Enabled, Operational, Active, Started, Schedulable) -- Fencing agent-level: MinItems=8 (same conditions as resources) +- Resource-level: MinItems=3, MaxItems=16. Healthy, Enabled, and Operational are always required. For Kubelet and Etcd, name-gated XValidation additionally requires InService, Managed, Active, Started, and Schedulable (8 total). Alert-agent entries (TaintAlertAgent, UntaintAlertAgent) require only the three always-required conditions. +- Fencing agent-level: MinItems=8 (Healthy, InService, Managed, Enabled, Operational, Active, Started, Schedulable) All condition arrays have XValidation rules to ensure specific condition types are present. **Resource names:** -- Valid values are: `Kubelet`, `Etcd` -- Both resources must be present in each node's `resources` array +- Valid values are: `Kubelet`, `Etcd`, `TaintAlertAgent`, `UntaintAlertAgent` +- `Kubelet` and `Etcd` must be present in each node's `resources` array +- `TaintAlertAgent` and `UntaintAlertAgent` are optional; when present, neither is required to appear +- Names must be unique within the `resources` array (enforced via the `name` list-map key) **Fencing agent fields:** - `name`: Unique identifier for the fencing agent (e.g., "master-0_redfish") diff --git a/etcd/v1/tests/pacemakerclusters.etcd.openshift.io/DualReplica.yaml b/etcd/v1/tests/pacemakerclusters.etcd.openshift.io/DualReplica.yaml index dde9ce6e9f1..25fb1b128b6 100644 --- a/etcd/v1/tests/pacemakerclusters.etcd.openshift.io/DualReplica.yaml +++ b/etcd/v1/tests/pacemakerclusters.etcd.openshift.io/DualReplica.yaml @@ -1933,3 +1933,896 @@ tests: reason: Schedulable message: "Schedulable" expectedStatusError: "must be a valid global unicast IPv4 or IPv6 address in canonical form" + + - name: Should accept alert-agent resources with only the conditions required for alert agents + initial: | + apiVersion: etcd.openshift.io/v1 + kind: PacemakerCluster + metadata: + name: cluster + updated: | + apiVersion: etcd.openshift.io/v1 + kind: PacemakerCluster + metadata: + name: cluster + status: + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ClusterHealthy + message: "Cluster is healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: NodeCountAsExpected + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: AsExpected + message: "Expected nodes present" + lastUpdated: "2024-01-01T00:00:01Z" + nodes: + - nodeName: master-0.example.com + addresses: + - type: InternalIP + address: "192.168.1.1" + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: NodeHealthy + message: "Node healthy" + - type: Online + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Online + message: "Online" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Ready + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Ready + message: "Ready" + - type: Clean + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Clean + message: "Clean" + - type: Member + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Member + message: "Member" + - type: FencingAvailable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: FencingAvailable + message: "Fencing available" + - type: FencingHealthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: FencingHealthy + message: "All fencing agents healthy" + resources: + - name: Kubelet + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + - name: Etcd + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + - name: TaintAlertAgent + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptConfigured + message: "Alert agent registered in the CIB" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptPresent + message: "Script present" + - name: UntaintAlertAgent + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptConfigured + message: "Alert agent registered in the CIB" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptPresent + message: "Script present" + fencingAgents: + - name: master-0.example.com_redfish + method: Redfish + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + expected: | + apiVersion: etcd.openshift.io/v1 + kind: PacemakerCluster + metadata: + name: cluster + status: + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ClusterHealthy + message: "Cluster is healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: NodeCountAsExpected + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: AsExpected + message: "Expected nodes present" + lastUpdated: "2024-01-01T00:00:01Z" + nodes: + - nodeName: master-0.example.com + addresses: + - type: InternalIP + address: "192.168.1.1" + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: NodeHealthy + message: "Node healthy" + - type: Online + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Online + message: "Online" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Ready + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Ready + message: "Ready" + - type: Clean + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Clean + message: "Clean" + - type: Member + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Member + message: "Member" + - type: FencingAvailable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: FencingAvailable + message: "Fencing available" + - type: FencingHealthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: FencingHealthy + message: "All fencing agents healthy" + resources: + - name: Kubelet + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + - name: Etcd + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + - name: TaintAlertAgent + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptConfigured + message: "Alert agent registered in the CIB" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptPresent + message: "Script present" + - name: UntaintAlertAgent + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptConfigured + message: "Alert agent registered in the CIB" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptPresent + message: "Script present" + fencingAgents: + - name: master-0.example.com_redfish + method: Redfish + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + - name: Should reject an alert-agent resource missing the Operational condition + initial: | + apiVersion: etcd.openshift.io/v1 + kind: PacemakerCluster + metadata: + name: cluster + updated: | + apiVersion: etcd.openshift.io/v1 + kind: PacemakerCluster + metadata: + name: cluster + status: + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ClusterHealthy + message: "Cluster is healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: NodeCountAsExpected + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: AsExpected + message: "Expected nodes present" + lastUpdated: "2024-01-01T00:00:01Z" + nodes: + - nodeName: master-0.example.com + addresses: + - type: InternalIP + address: "192.168.1.1" + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: NodeHealthy + message: "Node healthy" + - type: Online + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Online + message: "Online" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Ready + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Ready + message: "Ready" + - type: Clean + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Clean + message: "Clean" + - type: Member + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Member + message: "Member" + - type: FencingAvailable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: FencingAvailable + message: "Fencing available" + - type: FencingHealthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: FencingHealthy + message: "All fencing agents healthy" + resources: + - name: Kubelet + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + - name: Etcd + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + - name: TaintAlertAgent + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptConfigured + message: "Alert agent registered in the CIB" + - type: Extra + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Extra + message: "Placeholder" + fencingAgents: + - name: master-0.example.com_redfish + method: Redfish + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + expectedStatusError: "conditions must contain Healthy, Enabled, and Operational for TaintAlertAgent and UntaintAlertAgent resources" + - name: Should reject a Kubelet resource missing the conditions required for pacemaker-managed resources + initial: | + apiVersion: etcd.openshift.io/v1 + kind: PacemakerCluster + metadata: + name: cluster + updated: | + apiVersion: etcd.openshift.io/v1 + kind: PacemakerCluster + metadata: + name: cluster + status: + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ClusterHealthy + message: "Cluster is healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: NodeCountAsExpected + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: AsExpected + message: "Expected nodes present" + lastUpdated: "2024-01-01T00:00:01Z" + nodes: + - nodeName: master-0.example.com + addresses: + - type: InternalIP + address: "192.168.1.1" + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: NodeHealthy + message: "Node healthy" + - type: Online + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Online + message: "Online" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Ready + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Ready + message: "Ready" + - type: Clean + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Clean + message: "Clean" + - type: Member + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Member + message: "Member" + - type: FencingAvailable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: FencingAvailable + message: "Fencing available" + - type: FencingHealthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: FencingHealthy + message: "All fencing agents healthy" + resources: + - name: Kubelet + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - name: Etcd + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + fencingAgents: + - name: master-0.example.com_redfish + method: Redfish + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + expectedStatusError: "conditions must contain Healthy, InService, Managed, Enabled, Operational, Active, Started, and Schedulable for Kubelet and Etcd resources" diff --git a/etcd/v1/types_pacemakercluster.go b/etcd/v1/types_pacemakercluster.go index a481f5e1bd4..fa5ab2f0ca0 100644 --- a/etcd/v1/types_pacemakercluster.go +++ b/etcd/v1/types_pacemakercluster.go @@ -277,12 +277,24 @@ const ( // In Two Node OpenShift with Fencing, we do not expect any resources to be disabled. // When True, the resource is enabled with reason "Enabled". This is the normal operating state. // When False, the resource is disabled with reason "Disabled". This is an unexpected state. + // For alert-agent script resources (TaintAlertAgent, UntaintAlertAgent), an alert agent cannot be + // disabled, so this condition instead tracks whether the alert agent is registered in the CIB with + // the expected script path and event filter. When True, the agent is configured with reason + // "ScriptConfigured". When False, the agent is not registered with reason "AlertAgentNotRegistered", or is + // registered but misconfigured with reason "ScriptMisconfigured". When Unknown, + // registration has not yet been observed this run with reason "Pending"; this is expected to be temporary. ResourceEnabledConditionType = "Enabled" // ResourceOperationalConditionType tracks whether a resource is operational (not failed). // A failed resource is one that is not able to start or is in an error state. // When True, the resource is operational with reason "Operational". This is the normal operating state. // When False, the resource has failed with reason "Failed". This is an unexpected state. + // For alert-agent script resources (TaintAlertAgent, UntaintAlertAgent), this condition instead + // tracks whether the alert agent's script file is present and executable on this node, since the + // script is delivered independently to each node by MCO. When True, the script is present with + // reason "ScriptPresent". When False, the script is missing from this node with reason + // "ScriptMissing". When Unknown, presence has not yet been observed this run with reason + // "Pending"; this is expected to be temporary. ResourceOperationalConditionType = "Operational" // ResourceActiveConditionType tracks whether a resource is active. @@ -350,6 +362,24 @@ const ( // Resources that are disabled are stopped and not automatically managed or started by the cluster. // This is an unexpected state. ResourceEnabledReasonDisabled = "Disabled" + + // ResourceEnabledReasonScriptConfigured means an alert-agent script resource is registered in the + // CIB with the expected script path and event filter. This is the normal operating state for an + // alert-agent resource and is used in place of "Enabled". + ResourceEnabledReasonScriptConfigured = "ScriptConfigured" + + // ResourceEnabledReasonAlertAgentNotRegistered means an alert-agent script resource is not registered in + // the CIB at all. This is an unexpected state. + ResourceEnabledReasonAlertAgentNotRegistered = "AlertAgentNotRegistered" + + // ResourceEnabledReasonScriptMisconfigured means an alert-agent script resource is registered in + // the CIB but with an unexpected script path or event filter. This is an unexpected state. + ResourceEnabledReasonScriptMisconfigured = "ScriptMisconfigured" + + // ResourceEnabledReasonPending means an alert-agent script resource's CIB registration has not yet + // been observed this run by the status collector. Used only with status "Unknown". This is expected + // to be temporary, e.g. immediately after upgrade or before the first successful CIB collection. + ResourceEnabledReasonPending = "Pending" ) // ResourceOperational condition reasons @@ -361,6 +391,20 @@ const ( // ResourceOperationalReasonFailed means the resource has failed. // A failed resource is one that is not able to start or is in an error state. This is an unexpected state. ResourceOperationalReasonFailed = "Failed" + + // ResourceOperationalReasonScriptPresent means an alert-agent script resource's script file is + // present and executable on this node. This is the normal operating state for an alert-agent + // resource and is used in place of "Operational". + ResourceOperationalReasonScriptPresent = "ScriptPresent" + + // ResourceOperationalReasonScriptMissing means an alert-agent script resource's script file is + // missing or not executable on this node. This is an unexpected state. + ResourceOperationalReasonScriptMissing = "ScriptMissing" + + // ResourceOperationalReasonPending means an alert-agent script resource's presence on this node has + // not yet been observed this run by the status collector. Used only with status "Unknown". This is + // expected to be temporary, e.g. before MCO has delivered the script or before the first collection. + ResourceOperationalReasonPending = "Pending" ) // ResourceActive condition reasons @@ -431,8 +475,10 @@ type PacemakerNodeAddress struct { } // PacemakerClusterResourceName represents the name of a pacemaker resource. +// This includes both pacemaker-managed resources (Kubelet, Etcd) and pacemaker +// alert agents whose scripts are tracked per node (e.g. TaintAlertAgent, UntaintAlertAgent). // Fencing agents are tracked separately in the fencingAgents field. -// +kubebuilder:validation:Enum=Kubelet;Etcd +// +kubebuilder:validation:Enum=Kubelet;Etcd;TaintAlertAgent;UntaintAlertAgent // +enum type PacemakerClusterResourceName string @@ -445,6 +491,16 @@ const ( // PacemakerClusterResourceNameEtcd is the etcd pacemaker resource. // The etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. PacemakerClusterResourceNameEtcd PacemakerClusterResourceName = "Etcd" + + // PacemakerClusterResourceNameTaintAlertAgent is the alert agent that taints a node after it is fenced. + // Its entry tracks the alert agent's configuration in the CIB + // and the presence of its script on this node. + PacemakerClusterResourceNameTaintAlertAgent PacemakerClusterResourceName = "TaintAlertAgent" + + // PacemakerClusterResourceNameUntaintAlertAgent is the alert agent that removes a node's taint once + // it rejoins the cluster. Its entry tracks the alert agent's + // configuration in the CIB and the presence of its script on this node. + PacemakerClusterResourceNameUntaintAlertAgent PacemakerClusterResourceName = "UntaintAlertAgent" ) // FencingMethod represents the method used by a fencing agent to isolate failed nodes. @@ -506,7 +562,7 @@ type PacemakerClusterStatus struct { // The "Healthy" condition is an aggregate that tracks the overall health of the cluster. // The "InService" condition tracks whether the cluster is in service (not in maintenance mode). // The "NodeCountAsExpected" condition tracks whether the expected number of nodes are present. - // Each of these conditions is required, so the array must contain at least 3 items. + // Each of these three conditions is required, so the array must contain at least 3 items. // +listType=map // +listMapKey=type // +kubebuilder:validation:MinItems=3 @@ -589,11 +645,15 @@ type PacemakerClusterNodeStatus struct { // +required Addresses []PacemakerNodeAddress `json:"addresses,omitempty"` - // resources contains the status of pacemaker resources scheduled on this node. + // resources contains the status of pacemaker resources tracked on this node. // Each resource entry includes the resource name and its health conditions. - // For Two Node OpenShift with Fencing, we track Kubelet and Etcd resources per node. - // Both resources are required to be present, so the array must contain at least 2 items. - // Valid resource names are "Kubelet" and "Etcd". + // For Two Node OpenShift with Fencing, we track the Kubelet and Etcd pacemaker-managed resources + // per node. Both are required to be present, so the array must contain at least 2 items. + // The array may also contain optional alert-agent script entries named "TaintAlertAgent" and + // "UntaintAlertAgent". Their entries track whether the alert agent is configured in the CIB + // and whether its script is present on this node. Alert-agent entries are + // optional and are omitted when a status collector version that predates them has not reported them. + // Valid resource names are "Kubelet", "Etcd", "TaintAlertAgent", and "UntaintAlertAgent". // Fencing agents are tracked separately in the fencingAgents field. // +listType=map // +listMapKey=name @@ -672,15 +732,21 @@ type PacemakerClusterFencingAgentStatus struct { Method FencingMethod `json:"method,omitempty"` } -// PacemakerClusterResourceStatus represents the status of a pacemaker resource scheduled on a node. +// PacemakerClusterResourceStatus represents the status of a resource tracked on a node. // A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or // applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. -// For Two Node OpenShift with Fencing, we track two resources per node: +// For Two Node OpenShift with Fencing, we track the following pacemaker-managed resources per node: // - Kubelet (the Kubernetes node agent and a prerequisite for etcd) // - Etcd (the distributed key-value store) // +// The same type is reused to track alert-agent scripts (TaintAlertAgent, UntaintAlertAgent). An alert +// agent is not a pacemaker-managed resource, so only a subset of the pacemaker condition types applies +// to it: the required conditions are enforced conditionally based on the resource name. +// // Fencing agents are tracked separately in the fencingAgents field because they are mapped to // their target node (the node they can fence), not the node where monitoring operations are scheduled. +// +kubebuilder:validation:XValidation:rule="!(self.name == 'Kubelet' || self.name == 'Etcd') || ['Healthy','InService','Managed','Enabled','Operational','Active','Started','Schedulable'].all(t, self.conditions.exists(c, c.type == t))",message="conditions must contain Healthy, InService, Managed, Enabled, Operational, Active, Started, and Schedulable for Kubelet and Etcd resources" +// +kubebuilder:validation:XValidation:rule="!(self.name == 'TaintAlertAgent' || self.name == 'UntaintAlertAgent') || ['Healthy','Enabled','Operational'].all(t, self.conditions.exists(c, c.type == t))",message="conditions must contain Healthy, Enabled, and Operational for TaintAlertAgent and UntaintAlertAgent resources" type PacemakerClusterResourceStatus struct { // conditions represent the observations of the resource's current state. // Known condition types are: "Healthy", "InService", "Managed", "Enabled", "Operational", @@ -693,26 +759,28 @@ type PacemakerClusterResourceStatus struct { // The "Active" condition tracks whether the resource is active (available to be used). // The "Started" condition tracks whether the resource is started. // The "Schedulable" condition tracks whether the resource is schedulable (not blocked). - // Each of these conditions is required, so the array must contain at least 8 items. + // Which conditions are required depends on the resource name: + // - For the pacemaker-managed resources "Kubelet" and "Etcd", all eight condition types listed + // above are required, so the array must contain at least 8 items. + // - For the alert-agent resources "TaintAlertAgent" and "UntaintAlertAgent", only "Healthy", + // "Enabled" (reason "ScriptConfigured"), and "Operational" (reason "ScriptPresent") are + // required, so the array must contain at least 3 items. The remaining condition types do not + // apply to an alert agent and may be omitted. + // The array must contain at least 3 items in all cases; the exact set of required conditions for + // each resource name is enforced by name-gated validation rules on this type. // +listType=map // +listMapKey=type - // +kubebuilder:validation:MinItems=8 + // +kubebuilder:validation:MinItems=3 // +kubebuilder:validation:MaxItems=16 - // +kubebuilder:validation:XValidation:rule="self.exists(c, c.type == 'Healthy')",message="conditions must contain a condition of type Healthy" - // +kubebuilder:validation:XValidation:rule="self.exists(c, c.type == 'InService')",message="conditions must contain a condition of type InService" - // +kubebuilder:validation:XValidation:rule="self.exists(c, c.type == 'Managed')",message="conditions must contain a condition of type Managed" - // +kubebuilder:validation:XValidation:rule="self.exists(c, c.type == 'Enabled')",message="conditions must contain a condition of type Enabled" - // +kubebuilder:validation:XValidation:rule="self.exists(c, c.type == 'Operational')",message="conditions must contain a condition of type Operational" - // +kubebuilder:validation:XValidation:rule="self.exists(c, c.type == 'Active')",message="conditions must contain a condition of type Active" - // +kubebuilder:validation:XValidation:rule="self.exists(c, c.type == 'Started')",message="conditions must contain a condition of type Started" - // +kubebuilder:validation:XValidation:rule="self.exists(c, c.type == 'Schedulable')",message="conditions must contain a condition of type Schedulable" // +required Conditions []metav1.Condition `json:"conditions,omitempty"` - // name is the name of the pacemaker resource. - // Valid values are "Kubelet" and "Etcd". + // name is the name of the resource. + // Valid values are "Kubelet", "Etcd", "TaintAlertAgent", and "UntaintAlertAgent". // The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. // The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. + // The TaintAlertAgent and UntaintAlertAgent entries track alert-agent configuration and per-node + // script presence rather than a pacemaker-managed resource. // Fencing agents are tracked separately in the node's fencingAgents field. // +required Name PacemakerClusterResourceName `json:"name,omitempty"` diff --git a/etcd/v1/zz_generated.crd-manifests/0000_25_etcd_01_pacemakerclusters.crd.yaml b/etcd/v1/zz_generated.crd-manifests/0000_25_etcd_01_pacemakerclusters.crd.yaml index 0cad7e278b2..c1cca32951d 100644 --- a/etcd/v1/zz_generated.crd-manifests/0000_25_etcd_01_pacemakerclusters.crd.yaml +++ b/etcd/v1/zz_generated.crd-manifests/0000_25_etcd_01_pacemakerclusters.crd.yaml @@ -58,7 +58,7 @@ spec: The "Healthy" condition is an aggregate that tracks the overall health of the cluster. The "InService" condition tracks whether the cluster is in service (not in maintenance mode). The "NodeCountAsExpected" condition tracks whether the expected number of nodes are present. - Each of these conditions is required, so the array must contain at least 3 items. + Each of these three conditions is required, so the array must contain at least 3 items. items: description: Condition contains details for one aspect of the current state of this API Resource. @@ -452,21 +452,29 @@ spec: rule: '!format.dns1123Subdomain().validate(self).hasValue()' resources: description: |- - resources contains the status of pacemaker resources scheduled on this node. + resources contains the status of pacemaker resources tracked on this node. Each resource entry includes the resource name and its health conditions. - For Two Node OpenShift with Fencing, we track Kubelet and Etcd resources per node. - Both resources are required to be present, so the array must contain at least 2 items. - Valid resource names are "Kubelet" and "Etcd". + For Two Node OpenShift with Fencing, we track the Kubelet and Etcd pacemaker-managed resources + per node. Both are required to be present, so the array must contain at least 2 items. + The array may also contain optional alert-agent script entries named "TaintAlertAgent" and + "UntaintAlertAgent". Their entries track whether the alert agent is configured in the CIB + and whether its script is present on this node. Alert-agent entries are + optional and are omitted when a status collector version that predates them has not reported them. + Valid resource names are "Kubelet", "Etcd", "TaintAlertAgent", and "UntaintAlertAgent". Fencing agents are tracked separately in the fencingAgents field. items: description: |- - PacemakerClusterResourceStatus represents the status of a pacemaker resource scheduled on a node. + PacemakerClusterResourceStatus represents the status of a resource tracked on a node. A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. - For Two Node OpenShift with Fencing, we track two resources per node: + For Two Node OpenShift with Fencing, we track the following pacemaker-managed resources per node: - Kubelet (the Kubernetes node agent and a prerequisite for etcd) - Etcd (the distributed key-value store) + The same type is reused to track alert-agent scripts (TaintAlertAgent, UntaintAlertAgent). An alert + agent is not a pacemaker-managed resource, so only a subset of the pacemaker condition types applies + to it: the required conditions are enforced conditionally based on the resource name. + Fencing agents are tracked separately in the fencingAgents field because they are mapped to their target node (the node they can fence), not the node where monitoring operations are scheduled. properties: @@ -483,7 +491,15 @@ spec: The "Active" condition tracks whether the resource is active (available to be used). The "Started" condition tracks whether the resource is started. The "Schedulable" condition tracks whether the resource is schedulable (not blocked). - Each of these conditions is required, so the array must contain at least 8 items. + Which conditions are required depends on the resource name: + - For the pacemaker-managed resources "Kubelet" and "Etcd", all eight condition types listed + above are required, so the array must contain at least 8 items. + - For the alert-agent resources "TaintAlertAgent" and "UntaintAlertAgent", only "Healthy", + "Enabled" (reason "ScriptConfigured"), and "Operational" (reason "ScriptPresent") are + required, so the array must contain at least 3 items. The remaining condition types do not + apply to an alert agent and may be omitted. + The array must contain at least 3 items in all cases; the exact set of required conditions for + each resource name is enforced by name-gated validation rules on this type. items: description: Condition contains details for one aspect of the current state of this API Resource. @@ -541,51 +557,42 @@ spec: - type type: object maxItems: 16 - minItems: 8 + minItems: 3 type: array x-kubernetes-list-map-keys: - type x-kubernetes-list-type: map - x-kubernetes-validations: - - message: conditions must contain a condition of type - Healthy - rule: self.exists(c, c.type == 'Healthy') - - message: conditions must contain a condition of type - InService - rule: self.exists(c, c.type == 'InService') - - message: conditions must contain a condition of type - Managed - rule: self.exists(c, c.type == 'Managed') - - message: conditions must contain a condition of type - Enabled - rule: self.exists(c, c.type == 'Enabled') - - message: conditions must contain a condition of type - Operational - rule: self.exists(c, c.type == 'Operational') - - message: conditions must contain a condition of type - Active - rule: self.exists(c, c.type == 'Active') - - message: conditions must contain a condition of type - Started - rule: self.exists(c, c.type == 'Started') - - message: conditions must contain a condition of type - Schedulable - rule: self.exists(c, c.type == 'Schedulable') name: description: |- - name is the name of the pacemaker resource. - Valid values are "Kubelet" and "Etcd". + name is the name of the resource. + Valid values are "Kubelet", "Etcd", "TaintAlertAgent", and "UntaintAlertAgent". The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. + The TaintAlertAgent and UntaintAlertAgent entries track alert-agent configuration and per-node + script presence rather than a pacemaker-managed resource. Fencing agents are tracked separately in the node's fencingAgents field. enum: - Kubelet - Etcd + - TaintAlertAgent + - UntaintAlertAgent type: string required: - conditions - name type: object + x-kubernetes-validations: + - message: conditions must contain Healthy, InService, Managed, + Enabled, Operational, Active, Started, and Schedulable + for Kubelet and Etcd resources + rule: '!(self.name == ''Kubelet'' || self.name == ''Etcd'') + || [''Healthy'',''InService'',''Managed'',''Enabled'',''Operational'',''Active'',''Started'',''Schedulable''].all(t, + self.conditions.exists(c, c.type == t))' + - message: conditions must contain Healthy, Enabled, and Operational + for TaintAlertAgent and UntaintAlertAgent resources + rule: '!(self.name == ''TaintAlertAgent'' || self.name == + ''UntaintAlertAgent'') || [''Healthy'',''Enabled'',''Operational''].all(t, + self.conditions.exists(c, c.type == t))' maxItems: 8 minItems: 2 type: array diff --git a/etcd/v1/zz_generated.featuregated-crd-manifests/pacemakerclusters.etcd.openshift.io/DualReplica.yaml b/etcd/v1/zz_generated.featuregated-crd-manifests/pacemakerclusters.etcd.openshift.io/DualReplica.yaml index c9ec52f3829..ba6f442bc97 100644 --- a/etcd/v1/zz_generated.featuregated-crd-manifests/pacemakerclusters.etcd.openshift.io/DualReplica.yaml +++ b/etcd/v1/zz_generated.featuregated-crd-manifests/pacemakerclusters.etcd.openshift.io/DualReplica.yaml @@ -59,7 +59,7 @@ spec: The "Healthy" condition is an aggregate that tracks the overall health of the cluster. The "InService" condition tracks whether the cluster is in service (not in maintenance mode). The "NodeCountAsExpected" condition tracks whether the expected number of nodes are present. - Each of these conditions is required, so the array must contain at least 3 items. + Each of these three conditions is required, so the array must contain at least 3 items. items: description: Condition contains details for one aspect of the current state of this API Resource. @@ -453,21 +453,29 @@ spec: rule: '!format.dns1123Subdomain().validate(self).hasValue()' resources: description: |- - resources contains the status of pacemaker resources scheduled on this node. + resources contains the status of pacemaker resources tracked on this node. Each resource entry includes the resource name and its health conditions. - For Two Node OpenShift with Fencing, we track Kubelet and Etcd resources per node. - Both resources are required to be present, so the array must contain at least 2 items. - Valid resource names are "Kubelet" and "Etcd". + For Two Node OpenShift with Fencing, we track the Kubelet and Etcd pacemaker-managed resources + per node. Both are required to be present, so the array must contain at least 2 items. + The array may also contain optional alert-agent script entries named "TaintAlertAgent" and + "UntaintAlertAgent". Their entries track whether the alert agent is configured in the CIB + and whether its script is present on this node. Alert-agent entries are + optional and are omitted when a status collector version that predates them has not reported them. + Valid resource names are "Kubelet", "Etcd", "TaintAlertAgent", and "UntaintAlertAgent". Fencing agents are tracked separately in the fencingAgents field. items: description: |- - PacemakerClusterResourceStatus represents the status of a pacemaker resource scheduled on a node. + PacemakerClusterResourceStatus represents the status of a resource tracked on a node. A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. - For Two Node OpenShift with Fencing, we track two resources per node: + For Two Node OpenShift with Fencing, we track the following pacemaker-managed resources per node: - Kubelet (the Kubernetes node agent and a prerequisite for etcd) - Etcd (the distributed key-value store) + The same type is reused to track alert-agent scripts (TaintAlertAgent, UntaintAlertAgent). An alert + agent is not a pacemaker-managed resource, so only a subset of the pacemaker condition types applies + to it: the required conditions are enforced conditionally based on the resource name. + Fencing agents are tracked separately in the fencingAgents field because they are mapped to their target node (the node they can fence), not the node where monitoring operations are scheduled. properties: @@ -484,7 +492,15 @@ spec: The "Active" condition tracks whether the resource is active (available to be used). The "Started" condition tracks whether the resource is started. The "Schedulable" condition tracks whether the resource is schedulable (not blocked). - Each of these conditions is required, so the array must contain at least 8 items. + Which conditions are required depends on the resource name: + - For the pacemaker-managed resources "Kubelet" and "Etcd", all eight condition types listed + above are required, so the array must contain at least 8 items. + - For the alert-agent resources "TaintAlertAgent" and "UntaintAlertAgent", only "Healthy", + "Enabled" (reason "ScriptConfigured"), and "Operational" (reason "ScriptPresent") are + required, so the array must contain at least 3 items. The remaining condition types do not + apply to an alert agent and may be omitted. + The array must contain at least 3 items in all cases; the exact set of required conditions for + each resource name is enforced by name-gated validation rules on this type. items: description: Condition contains details for one aspect of the current state of this API Resource. @@ -542,51 +558,42 @@ spec: - type type: object maxItems: 16 - minItems: 8 + minItems: 3 type: array x-kubernetes-list-map-keys: - type x-kubernetes-list-type: map - x-kubernetes-validations: - - message: conditions must contain a condition of type - Healthy - rule: self.exists(c, c.type == 'Healthy') - - message: conditions must contain a condition of type - InService - rule: self.exists(c, c.type == 'InService') - - message: conditions must contain a condition of type - Managed - rule: self.exists(c, c.type == 'Managed') - - message: conditions must contain a condition of type - Enabled - rule: self.exists(c, c.type == 'Enabled') - - message: conditions must contain a condition of type - Operational - rule: self.exists(c, c.type == 'Operational') - - message: conditions must contain a condition of type - Active - rule: self.exists(c, c.type == 'Active') - - message: conditions must contain a condition of type - Started - rule: self.exists(c, c.type == 'Started') - - message: conditions must contain a condition of type - Schedulable - rule: self.exists(c, c.type == 'Schedulable') name: description: |- - name is the name of the pacemaker resource. - Valid values are "Kubelet" and "Etcd". + name is the name of the resource. + Valid values are "Kubelet", "Etcd", "TaintAlertAgent", and "UntaintAlertAgent". The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. + The TaintAlertAgent and UntaintAlertAgent entries track alert-agent configuration and per-node + script presence rather than a pacemaker-managed resource. Fencing agents are tracked separately in the node's fencingAgents field. enum: - Kubelet - Etcd + - TaintAlertAgent + - UntaintAlertAgent type: string required: - conditions - name type: object + x-kubernetes-validations: + - message: conditions must contain Healthy, InService, Managed, + Enabled, Operational, Active, Started, and Schedulable + for Kubelet and Etcd resources + rule: '!(self.name == ''Kubelet'' || self.name == ''Etcd'') + || [''Healthy'',''InService'',''Managed'',''Enabled'',''Operational'',''Active'',''Started'',''Schedulable''].all(t, + self.conditions.exists(c, c.type == t))' + - message: conditions must contain Healthy, Enabled, and Operational + for TaintAlertAgent and UntaintAlertAgent resources + rule: '!(self.name == ''TaintAlertAgent'' || self.name == + ''UntaintAlertAgent'') || [''Healthy'',''Enabled'',''Operational''].all(t, + self.conditions.exists(c, c.type == t))' maxItems: 8 minItems: 2 type: array diff --git a/etcd/v1/zz_generated.swagger_doc_generated.go b/etcd/v1/zz_generated.swagger_doc_generated.go index e9e47b47cfd..0f512ec3ffa 100644 --- a/etcd/v1/zz_generated.swagger_doc_generated.go +++ b/etcd/v1/zz_generated.swagger_doc_generated.go @@ -47,7 +47,7 @@ var map_PacemakerClusterNodeStatus = map[string]string{ "conditions": "conditions represent the observations of the node's current state. Known condition types are: \"Healthy\", \"Online\", \"InService\", \"Active\", \"Ready\", \"Clean\", \"Member\", \"FencingAvailable\", \"FencingHealthy\". The \"Healthy\" condition is an aggregate that tracks the overall health of the node. The \"Online\" condition tracks whether the node is online. The \"InService\" condition tracks whether the node is in service (not in maintenance mode). The \"Active\" condition tracks whether the node is active (not in standby mode). The \"Ready\" condition tracks whether the node is ready (not in a pending state). The \"Clean\" condition tracks whether the node is in a clean (status known) state. The \"Member\" condition tracks whether the node is a member of the cluster. The \"FencingAvailable\" condition tracks whether this node can be fenced by at least one healthy agent. The \"FencingHealthy\" condition tracks whether all fencing agents for this node are healthy. Each of these conditions is required, so the array must contain at least 9 items.", "nodeName": "nodeName is the name of the node. This is expected to match the Kubernetes node's name, which must be a lowercase RFC 1123 subdomain consisting of lowercase alphanumeric characters, '-' or '.', starting and ending with an alphanumeric character, and be at most 253 characters in length.", "addresses": "addresses is a list of IP addresses for the node. Pacemaker allows multiple IP addresses for Corosync communication between nodes. The first address in this list is used for IP-based peer URLs for etcd membership. Each address must be a valid global unicast IPv4 or IPv6 address in canonical form (e.g., \"192.168.1.1\" not \"192.168.001.001\", or \"2001:db8::1\" not \"2001:0db8::1\"). This excludes loopback, link-local, and multicast addresses.", - "resources": "resources contains the status of pacemaker resources scheduled on this node. Each resource entry includes the resource name and its health conditions. For Two Node OpenShift with Fencing, we track Kubelet and Etcd resources per node. Both resources are required to be present, so the array must contain at least 2 items. Valid resource names are \"Kubelet\" and \"Etcd\". Fencing agents are tracked separately in the fencingAgents field.", + "resources": "resources contains the status of pacemaker resources tracked on this node. Each resource entry includes the resource name and its health conditions. For Two Node OpenShift with Fencing, we track the Kubelet and Etcd pacemaker-managed resources per node. Both are required to be present, so the array must contain at least 2 items. The array may also contain optional alert-agent script entries named \"TaintAlertAgent\" and \"UntaintAlertAgent\". Their entries track whether the alert agent is configured in the CIB and whether its script is present on this node. Alert-agent entries are optional and are omitted when a status collector version that predates them has not reported them. Valid resource names are \"Kubelet\", \"Etcd\", \"TaintAlertAgent\", and \"UntaintAlertAgent\". Fencing agents are tracked separately in the fencingAgents field.", "fencingAgents": "fencingAgents contains the status of fencing agents that can fence this node. Unlike resources (which are scheduled to run on this node), fencing agents are mapped to the node they can fence (their target), not the node where monitoring operations run. Each fencing agent entry includes a unique name, fencing type, target node, and health conditions. A node is considered fence-capable if at least one fencing agent is healthy. A healthy node is expected to have at least 1 fencing agent, but the list may be empty when fencing agent discovery fails. Names must be unique within this array.", } @@ -56,9 +56,9 @@ func (PacemakerClusterNodeStatus) SwaggerDoc() map[string]string { } var map_PacemakerClusterResourceStatus = map[string]string{ - "": "PacemakerClusterResourceStatus represents the status of a pacemaker resource scheduled on a node. A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. For Two Node OpenShift with Fencing, we track two resources per node:\n - Kubelet (the Kubernetes node agent and a prerequisite for etcd)\n - Etcd (the distributed key-value store)\n\nFencing agents are tracked separately in the fencingAgents field because they are mapped to their target node (the node they can fence), not the node where monitoring operations are scheduled.", - "conditions": "conditions represent the observations of the resource's current state. Known condition types are: \"Healthy\", \"InService\", \"Managed\", \"Enabled\", \"Operational\", \"Active\", \"Started\", \"Schedulable\". The \"Healthy\" condition is an aggregate that tracks the overall health of the resource. The \"InService\" condition tracks whether the resource is in service (not in maintenance mode). The \"Managed\" condition tracks whether the resource is managed by pacemaker. The \"Enabled\" condition tracks whether the resource is enabled. The \"Operational\" condition tracks whether the resource is operational (not failed). The \"Active\" condition tracks whether the resource is active (available to be used). The \"Started\" condition tracks whether the resource is started. The \"Schedulable\" condition tracks whether the resource is schedulable (not blocked). Each of these conditions is required, so the array must contain at least 8 items.", - "name": "name is the name of the pacemaker resource. Valid values are \"Kubelet\" and \"Etcd\". The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. Fencing agents are tracked separately in the node's fencingAgents field.", + "": "PacemakerClusterResourceStatus represents the status of a resource tracked on a node. A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. For Two Node OpenShift with Fencing, we track the following pacemaker-managed resources per node:\n - Kubelet (the Kubernetes node agent and a prerequisite for etcd)\n - Etcd (the distributed key-value store)\n\nThe same type is reused to track alert-agent scripts (TaintAlertAgent, UntaintAlertAgent). An alert agent is not a pacemaker-managed resource, so only a subset of the pacemaker condition types applies to it: the required conditions are enforced conditionally based on the resource name.\n\nFencing agents are tracked separately in the fencingAgents field because they are mapped to their target node (the node they can fence), not the node where monitoring operations are scheduled.", + "conditions": "conditions represent the observations of the resource's current state. Known condition types are: \"Healthy\", \"InService\", \"Managed\", \"Enabled\", \"Operational\", \"Active\", \"Started\", \"Schedulable\". The \"Healthy\" condition is an aggregate that tracks the overall health of the resource. The \"InService\" condition tracks whether the resource is in service (not in maintenance mode). The \"Managed\" condition tracks whether the resource is managed by pacemaker. The \"Enabled\" condition tracks whether the resource is enabled. The \"Operational\" condition tracks whether the resource is operational (not failed). The \"Active\" condition tracks whether the resource is active (available to be used). The \"Started\" condition tracks whether the resource is started. The \"Schedulable\" condition tracks whether the resource is schedulable (not blocked). Which conditions are required depends on the resource name:\n - For the pacemaker-managed resources \"Kubelet\" and \"Etcd\", all eight condition types listed\n above are required, so the array must contain at least 8 items.\n - For the alert-agent resources \"TaintAlertAgent\" and \"UntaintAlertAgent\", only \"Healthy\",\n \"Enabled\" (reason \"ScriptConfigured\"), and \"Operational\" (reason \"ScriptPresent\") are\n required, so the array must contain at least 3 items. The remaining condition types do not\n apply to an alert agent and may be omitted.\nThe array must contain at least 3 items in all cases; the exact set of required conditions for each resource name is enforced by name-gated validation rules on this type.", + "name": "name is the name of the resource. Valid values are \"Kubelet\", \"Etcd\", \"TaintAlertAgent\", and \"UntaintAlertAgent\". The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. The TaintAlertAgent and UntaintAlertAgent entries track alert-agent configuration and per-node script presence rather than a pacemaker-managed resource. Fencing agents are tracked separately in the node's fencingAgents field.", } func (PacemakerClusterResourceStatus) SwaggerDoc() map[string]string { @@ -67,7 +67,7 @@ func (PacemakerClusterResourceStatus) SwaggerDoc() map[string]string { var map_PacemakerClusterStatus = map[string]string{ "": "PacemakerClusterStatus contains the actual pacemaker cluster status information. As part of validating the status object, we need to ensure that the lastUpdated timestamp may not be set to an earlier timestamp than the current value. The validation rule checks if oldSelf has lastUpdated before comparing, to handle the initial status creation case.", - "conditions": "conditions represent the observations of the pacemaker cluster's current state. Known condition types are: \"Healthy\", \"InService\", \"NodeCountAsExpected\". The \"Healthy\" condition is an aggregate that tracks the overall health of the cluster. The \"InService\" condition tracks whether the cluster is in service (not in maintenance mode). The \"NodeCountAsExpected\" condition tracks whether the expected number of nodes are present. Each of these conditions is required, so the array must contain at least 3 items.", + "conditions": "conditions represent the observations of the pacemaker cluster's current state. Known condition types are: \"Healthy\", \"InService\", \"NodeCountAsExpected\". The \"Healthy\" condition is an aggregate that tracks the overall health of the cluster. The \"InService\" condition tracks whether the cluster is in service (not in maintenance mode). The \"NodeCountAsExpected\" condition tracks whether the expected number of nodes are present. Each of these three conditions is required, so the array must contain at least 3 items.", "lastUpdated": "lastUpdated is the timestamp when this status was last updated. This is useful for identifying stale status reports. It must be a valid timestamp in RFC3339 format. Once set, this field cannot be removed and cannot be set to an earlier timestamp than the current value.", "nodes": "nodes provides detailed status for each control-plane node in the Pacemaker cluster. While Pacemaker supports up to 32 nodes, the limit is set to 5 (max OpenShift control-plane nodes). For Two Node OpenShift with Fencing, exactly 2 nodes are expected in a healthy cluster. An empty list indicates a catastrophic failure where Pacemaker reports no nodes.", } diff --git a/etcd/v1alpha1/tests/pacemakerclusters.etcd.openshift.io/DualReplica.yaml b/etcd/v1alpha1/tests/pacemakerclusters.etcd.openshift.io/DualReplica.yaml index 6c919c01946..be30c164fd3 100644 --- a/etcd/v1alpha1/tests/pacemakerclusters.etcd.openshift.io/DualReplica.yaml +++ b/etcd/v1alpha1/tests/pacemakerclusters.etcd.openshift.io/DualReplica.yaml @@ -1933,3 +1933,896 @@ tests: reason: Schedulable message: "Schedulable" expectedStatusError: "must be a valid global unicast IPv4 or IPv6 address in canonical form" + + - name: Should accept alert-agent resources with only the conditions required for alert agents + initial: | + apiVersion: etcd.openshift.io/v1alpha1 + kind: PacemakerCluster + metadata: + name: cluster + updated: | + apiVersion: etcd.openshift.io/v1alpha1 + kind: PacemakerCluster + metadata: + name: cluster + status: + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ClusterHealthy + message: "Cluster is healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: NodeCountAsExpected + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: AsExpected + message: "Expected nodes present" + lastUpdated: "2024-01-01T00:00:01Z" + nodes: + - nodeName: master-0.example.com + addresses: + - type: InternalIP + address: "192.168.1.1" + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: NodeHealthy + message: "Node healthy" + - type: Online + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Online + message: "Online" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Ready + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Ready + message: "Ready" + - type: Clean + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Clean + message: "Clean" + - type: Member + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Member + message: "Member" + - type: FencingAvailable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: FencingAvailable + message: "Fencing available" + - type: FencingHealthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: FencingHealthy + message: "All fencing agents healthy" + resources: + - name: Kubelet + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + - name: Etcd + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + - name: TaintAlertAgent + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptConfigured + message: "Alert agent registered in the CIB" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptPresent + message: "Script present" + - name: UntaintAlertAgent + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptConfigured + message: "Alert agent registered in the CIB" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptPresent + message: "Script present" + fencingAgents: + - name: master-0.example.com_redfish + method: Redfish + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + expected: | + apiVersion: etcd.openshift.io/v1alpha1 + kind: PacemakerCluster + metadata: + name: cluster + status: + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ClusterHealthy + message: "Cluster is healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: NodeCountAsExpected + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: AsExpected + message: "Expected nodes present" + lastUpdated: "2024-01-01T00:00:01Z" + nodes: + - nodeName: master-0.example.com + addresses: + - type: InternalIP + address: "192.168.1.1" + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: NodeHealthy + message: "Node healthy" + - type: Online + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Online + message: "Online" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Ready + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Ready + message: "Ready" + - type: Clean + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Clean + message: "Clean" + - type: Member + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Member + message: "Member" + - type: FencingAvailable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: FencingAvailable + message: "Fencing available" + - type: FencingHealthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: FencingHealthy + message: "All fencing agents healthy" + resources: + - name: Kubelet + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + - name: Etcd + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + - name: TaintAlertAgent + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptConfigured + message: "Alert agent registered in the CIB" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptPresent + message: "Script present" + - name: UntaintAlertAgent + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptConfigured + message: "Alert agent registered in the CIB" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptPresent + message: "Script present" + fencingAgents: + - name: master-0.example.com_redfish + method: Redfish + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + - name: Should reject an alert-agent resource missing the Operational condition + initial: | + apiVersion: etcd.openshift.io/v1alpha1 + kind: PacemakerCluster + metadata: + name: cluster + updated: | + apiVersion: etcd.openshift.io/v1alpha1 + kind: PacemakerCluster + metadata: + name: cluster + status: + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ClusterHealthy + message: "Cluster is healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: NodeCountAsExpected + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: AsExpected + message: "Expected nodes present" + lastUpdated: "2024-01-01T00:00:01Z" + nodes: + - nodeName: master-0.example.com + addresses: + - type: InternalIP + address: "192.168.1.1" + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: NodeHealthy + message: "Node healthy" + - type: Online + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Online + message: "Online" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Ready + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Ready + message: "Ready" + - type: Clean + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Clean + message: "Clean" + - type: Member + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Member + message: "Member" + - type: FencingAvailable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: FencingAvailable + message: "Fencing available" + - type: FencingHealthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: FencingHealthy + message: "All fencing agents healthy" + resources: + - name: Kubelet + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + - name: Etcd + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + - name: TaintAlertAgent + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ScriptConfigured + message: "Alert agent registered in the CIB" + - type: Extra + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Extra + message: "Placeholder" + fencingAgents: + - name: master-0.example.com_redfish + method: Redfish + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + expectedStatusError: "conditions must contain Healthy, Enabled, and Operational for TaintAlertAgent and UntaintAlertAgent resources" + - name: Should reject a Kubelet resource missing the conditions required for pacemaker-managed resources + initial: | + apiVersion: etcd.openshift.io/v1alpha1 + kind: PacemakerCluster + metadata: + name: cluster + updated: | + apiVersion: etcd.openshift.io/v1alpha1 + kind: PacemakerCluster + metadata: + name: cluster + status: + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ClusterHealthy + message: "Cluster is healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: NodeCountAsExpected + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: AsExpected + message: "Expected nodes present" + lastUpdated: "2024-01-01T00:00:01Z" + nodes: + - nodeName: master-0.example.com + addresses: + - type: InternalIP + address: "192.168.1.1" + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: NodeHealthy + message: "Node healthy" + - type: Online + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Online + message: "Online" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Ready + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Ready + message: "Ready" + - type: Clean + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Clean + message: "Clean" + - type: Member + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Member + message: "Member" + - type: FencingAvailable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: FencingAvailable + message: "Fencing available" + - type: FencingHealthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: FencingHealthy + message: "All fencing agents healthy" + resources: + - name: Kubelet + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - name: Etcd + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + fencingAgents: + - name: master-0.example.com_redfish + method: Redfish + conditions: + - type: Healthy + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: ResourceHealthy + message: "Healthy" + - type: InService + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: InService + message: "In service" + - type: Managed + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Managed + message: "Managed" + - type: Enabled + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Enabled + message: "Enabled" + - type: Operational + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Operational + message: "Operational" + - type: Active + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Active + message: "Active" + - type: Started + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Started + message: "Started" + - type: Schedulable + status: "True" + lastTransitionTime: "2024-01-01T00:00:00Z" + reason: Schedulable + message: "Schedulable" + expectedStatusError: "conditions must contain Healthy, InService, Managed, Enabled, Operational, Active, Started, and Schedulable for Kubelet and Etcd resources" diff --git a/etcd/v1alpha1/types_pacemakercluster.go b/etcd/v1alpha1/types_pacemakercluster.go index b627413474b..053a1db9e71 100644 --- a/etcd/v1alpha1/types_pacemakercluster.go +++ b/etcd/v1alpha1/types_pacemakercluster.go @@ -277,12 +277,24 @@ const ( // In Two Node OpenShift with Fencing, we do not expect any resources to be disabled. // When True, the resource is enabled with reason "Enabled". This is the normal operating state. // When False, the resource is disabled with reason "Disabled". This is an unexpected state. + // For alert-agent script resources (TaintAlertAgent, UntaintAlertAgent), an alert agent cannot be + // disabled, so this condition instead tracks whether the alert agent is registered in the CIB with + // the expected script path and event filter. When True, the agent is configured with reason + // "ScriptConfigured". When False, the agent is not registered with reason "AlertAgentNotRegistered", or is + // registered but misconfigured with reason "ScriptMisconfigured". When Unknown, + // registration has not yet been observed this run with reason "Pending"; this is expected to be temporary. ResourceEnabledConditionType = "Enabled" // ResourceOperationalConditionType tracks whether a resource is operational (not failed). // A failed resource is one that is not able to start or is in an error state. // When True, the resource is operational with reason "Operational". This is the normal operating state. // When False, the resource has failed with reason "Failed". This is an unexpected state. + // For alert-agent script resources (TaintAlertAgent, UntaintAlertAgent), this condition instead + // tracks whether the alert agent's script file is present and executable on this node, since the + // script is delivered independently to each node by MCO. When True, the script is present with + // reason "ScriptPresent". When False, the script is missing from this node with reason + // "ScriptMissing". When Unknown, presence has not yet been observed this run with reason + // "Pending"; this is expected to be temporary. ResourceOperationalConditionType = "Operational" // ResourceActiveConditionType tracks whether a resource is active. @@ -350,6 +362,24 @@ const ( // Resources that are disabled are stopped and not automatically managed or started by the cluster. // This is an unexpected state. ResourceEnabledReasonDisabled = "Disabled" + + // ResourceEnabledReasonScriptConfigured means an alert-agent script resource is registered in the + // CIB with the expected script path and event filter. This is the normal operating state for an + // alert-agent resource and is used in place of "Enabled". + ResourceEnabledReasonScriptConfigured = "ScriptConfigured" + + // ResourceEnabledReasonAlertAgentNotRegistered means an alert-agent script resource is not registered in + // the CIB at all. This is an unexpected state. + ResourceEnabledReasonAlertAgentNotRegistered = "AlertAgentNotRegistered" + + // ResourceEnabledReasonScriptMisconfigured means an alert-agent script resource is registered in + // the CIB but with an unexpected script path or event filter. This is an unexpected state. + ResourceEnabledReasonScriptMisconfigured = "ScriptMisconfigured" + + // ResourceEnabledReasonPending means an alert-agent script resource's CIB registration has not yet + // been observed this run by the status collector. Used only with status "Unknown". This is expected + // to be temporary, e.g. immediately after upgrade or before the first successful CIB collection. + ResourceEnabledReasonPending = "Pending" ) // ResourceOperational condition reasons @@ -361,6 +391,20 @@ const ( // ResourceOperationalReasonFailed means the resource has failed. // A failed resource is one that is not able to start or is in an error state. This is an unexpected state. ResourceOperationalReasonFailed = "Failed" + + // ResourceOperationalReasonScriptPresent means an alert-agent script resource's script file is + // present and executable on this node. This is the normal operating state for an alert-agent + // resource and is used in place of "Operational". + ResourceOperationalReasonScriptPresent = "ScriptPresent" + + // ResourceOperationalReasonScriptMissing means an alert-agent script resource's script file is + // missing or not executable on this node. This is an unexpected state. + ResourceOperationalReasonScriptMissing = "ScriptMissing" + + // ResourceOperationalReasonPending means an alert-agent script resource's presence on this node has + // not yet been observed this run by the status collector. Used only with status "Unknown". This is + // expected to be temporary, e.g. before MCO has delivered the script or before the first collection. + ResourceOperationalReasonPending = "Pending" ) // ResourceActive condition reasons @@ -431,8 +475,10 @@ type PacemakerNodeAddress struct { } // PacemakerClusterResourceName represents the name of a pacemaker resource. +// This includes both pacemaker-managed resources (Kubelet, Etcd) and pacemaker +// alert agents whose scripts are tracked per node (e.g. TaintAlertAgent, UntaintAlertAgent). // Fencing agents are tracked separately in the fencingAgents field. -// +kubebuilder:validation:Enum=Kubelet;Etcd +// +kubebuilder:validation:Enum=Kubelet;Etcd;TaintAlertAgent;UntaintAlertAgent // +enum type PacemakerClusterResourceName string @@ -445,6 +491,16 @@ const ( // PacemakerClusterResourceNameEtcd is the etcd pacemaker resource. // The etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. PacemakerClusterResourceNameEtcd PacemakerClusterResourceName = "Etcd" + + // PacemakerClusterResourceNameTaintAlertAgent is the alert agent that taints a node after it is fenced. + // Its entry tracks the alert agent's configuration in the CIB + // and the presence of its script on this node. + PacemakerClusterResourceNameTaintAlertAgent PacemakerClusterResourceName = "TaintAlertAgent" + + // PacemakerClusterResourceNameUntaintAlertAgent is the alert agent that removes a node's taint once + // it rejoins the cluster. Its entry tracks the alert agent's + // configuration in the CIB and the presence of its script on this node. + PacemakerClusterResourceNameUntaintAlertAgent PacemakerClusterResourceName = "UntaintAlertAgent" ) // FencingMethod represents the method used by a fencing agent to isolate failed nodes. @@ -506,7 +562,7 @@ type PacemakerClusterStatus struct { // The "Healthy" condition is an aggregate that tracks the overall health of the cluster. // The "InService" condition tracks whether the cluster is in service (not in maintenance mode). // The "NodeCountAsExpected" condition tracks whether the expected number of nodes are present. - // Each of these conditions is required, so the array must contain at least 3 items. + // Each of these three conditions is required, so the array must contain at least 3 items. // +listType=map // +listMapKey=type // +kubebuilder:validation:MinItems=3 @@ -589,11 +645,15 @@ type PacemakerClusterNodeStatus struct { // +required Addresses []PacemakerNodeAddress `json:"addresses,omitempty"` - // resources contains the status of pacemaker resources scheduled on this node. + // resources contains the status of pacemaker resources tracked on this node. // Each resource entry includes the resource name and its health conditions. - // For Two Node OpenShift with Fencing, we track Kubelet and Etcd resources per node. - // Both resources are required to be present, so the array must contain at least 2 items. - // Valid resource names are "Kubelet" and "Etcd". + // For Two Node OpenShift with Fencing, we track the Kubelet and Etcd pacemaker-managed resources + // per node. Both are required to be present, so the array must contain at least 2 items. + // The array may also contain optional alert-agent script entries named "TaintAlertAgent" and + // "UntaintAlertAgent". Their entries track whether the alert agent is configured in the CIB + // and whether its script is present on this node. Alert-agent entries are + // optional and are omitted when a status collector version that predates them has not reported them. + // Valid resource names are "Kubelet", "Etcd", "TaintAlertAgent", and "UntaintAlertAgent". // Fencing agents are tracked separately in the fencingAgents field. // +listType=map // +listMapKey=name @@ -672,15 +732,21 @@ type PacemakerClusterFencingAgentStatus struct { Method FencingMethod `json:"method,omitempty"` } -// PacemakerClusterResourceStatus represents the status of a pacemaker resource scheduled on a node. +// PacemakerClusterResourceStatus represents the status of a resource tracked on a node. // A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or // applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. -// For Two Node OpenShift with Fencing, we track two resources per node: +// For Two Node OpenShift with Fencing, we track the following pacemaker-managed resources per node: // - Kubelet (the Kubernetes node agent and a prerequisite for etcd) // - Etcd (the distributed key-value store) // +// The same type is reused to track alert-agent scripts (TaintAlertAgent, UntaintAlertAgent). An alert +// agent is not a pacemaker-managed resource, so only a subset of the pacemaker condition types applies +// to it: the required conditions are enforced conditionally based on the resource name. +// // Fencing agents are tracked separately in the fencingAgents field because they are mapped to // their target node (the node they can fence), not the node where monitoring operations are scheduled. +// +kubebuilder:validation:XValidation:rule="!(self.name == 'Kubelet' || self.name == 'Etcd') || ['Healthy','InService','Managed','Enabled','Operational','Active','Started','Schedulable'].all(t, self.conditions.exists(c, c.type == t))",message="conditions must contain Healthy, InService, Managed, Enabled, Operational, Active, Started, and Schedulable for Kubelet and Etcd resources" +// +kubebuilder:validation:XValidation:rule="!(self.name == 'TaintAlertAgent' || self.name == 'UntaintAlertAgent') || ['Healthy','Enabled','Operational'].all(t, self.conditions.exists(c, c.type == t))",message="conditions must contain Healthy, Enabled, and Operational for TaintAlertAgent and UntaintAlertAgent resources" type PacemakerClusterResourceStatus struct { // conditions represent the observations of the resource's current state. // Known condition types are: "Healthy", "InService", "Managed", "Enabled", "Operational", @@ -693,26 +759,28 @@ type PacemakerClusterResourceStatus struct { // The "Active" condition tracks whether the resource is active (available to be used). // The "Started" condition tracks whether the resource is started. // The "Schedulable" condition tracks whether the resource is schedulable (not blocked). - // Each of these conditions is required, so the array must contain at least 8 items. + // Which conditions are required depends on the resource name: + // - For the pacemaker-managed resources "Kubelet" and "Etcd", all eight condition types listed + // above are required, so the array must contain at least 8 items. + // - For the alert-agent resources "TaintAlertAgent" and "UntaintAlertAgent", only "Healthy", + // "Enabled" (reason "ScriptConfigured"), and "Operational" (reason "ScriptPresent") are + // required, so the array must contain at least 3 items. The remaining condition types do not + // apply to an alert agent and may be omitted. + // The array must contain at least 3 items in all cases; the exact set of required conditions for + // each resource name is enforced by name-gated validation rules on this type. // +listType=map // +listMapKey=type - // +kubebuilder:validation:MinItems=8 + // +kubebuilder:validation:MinItems=3 // +kubebuilder:validation:MaxItems=16 - // +kubebuilder:validation:XValidation:rule="self.exists(c, c.type == 'Healthy')",message="conditions must contain a condition of type Healthy" - // +kubebuilder:validation:XValidation:rule="self.exists(c, c.type == 'InService')",message="conditions must contain a condition of type InService" - // +kubebuilder:validation:XValidation:rule="self.exists(c, c.type == 'Managed')",message="conditions must contain a condition of type Managed" - // +kubebuilder:validation:XValidation:rule="self.exists(c, c.type == 'Enabled')",message="conditions must contain a condition of type Enabled" - // +kubebuilder:validation:XValidation:rule="self.exists(c, c.type == 'Operational')",message="conditions must contain a condition of type Operational" - // +kubebuilder:validation:XValidation:rule="self.exists(c, c.type == 'Active')",message="conditions must contain a condition of type Active" - // +kubebuilder:validation:XValidation:rule="self.exists(c, c.type == 'Started')",message="conditions must contain a condition of type Started" - // +kubebuilder:validation:XValidation:rule="self.exists(c, c.type == 'Schedulable')",message="conditions must contain a condition of type Schedulable" // +required Conditions []metav1.Condition `json:"conditions,omitempty"` - // name is the name of the pacemaker resource. - // Valid values are "Kubelet" and "Etcd". + // name is the name of the resource. + // Valid values are "Kubelet", "Etcd", "TaintAlertAgent", and "UntaintAlertAgent". // The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. // The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. + // The TaintAlertAgent and UntaintAlertAgent entries track alert-agent configuration and per-node + // script presence rather than a pacemaker-managed resource. // Fencing agents are tracked separately in the node's fencingAgents field. // +required Name PacemakerClusterResourceName `json:"name,omitempty"` diff --git a/etcd/v1alpha1/zz_generated.crd-manifests/0000_25_etcd_01_pacemakerclusters.crd.yaml b/etcd/v1alpha1/zz_generated.crd-manifests/0000_25_etcd_01_pacemakerclusters.crd.yaml index e26afefbb18..928877f0de5 100644 --- a/etcd/v1alpha1/zz_generated.crd-manifests/0000_25_etcd_01_pacemakerclusters.crd.yaml +++ b/etcd/v1alpha1/zz_generated.crd-manifests/0000_25_etcd_01_pacemakerclusters.crd.yaml @@ -58,7 +58,7 @@ spec: The "Healthy" condition is an aggregate that tracks the overall health of the cluster. The "InService" condition tracks whether the cluster is in service (not in maintenance mode). The "NodeCountAsExpected" condition tracks whether the expected number of nodes are present. - Each of these conditions is required, so the array must contain at least 3 items. + Each of these three conditions is required, so the array must contain at least 3 items. items: description: Condition contains details for one aspect of the current state of this API Resource. @@ -452,21 +452,29 @@ spec: rule: '!format.dns1123Subdomain().validate(self).hasValue()' resources: description: |- - resources contains the status of pacemaker resources scheduled on this node. + resources contains the status of pacemaker resources tracked on this node. Each resource entry includes the resource name and its health conditions. - For Two Node OpenShift with Fencing, we track Kubelet and Etcd resources per node. - Both resources are required to be present, so the array must contain at least 2 items. - Valid resource names are "Kubelet" and "Etcd". + For Two Node OpenShift with Fencing, we track the Kubelet and Etcd pacemaker-managed resources + per node. Both are required to be present, so the array must contain at least 2 items. + The array may also contain optional alert-agent script entries named "TaintAlertAgent" and + "UntaintAlertAgent". Their entries track whether the alert agent is configured in the CIB + and whether its script is present on this node. Alert-agent entries are + optional and are omitted when a status collector version that predates them has not reported them. + Valid resource names are "Kubelet", "Etcd", "TaintAlertAgent", and "UntaintAlertAgent". Fencing agents are tracked separately in the fencingAgents field. items: description: |- - PacemakerClusterResourceStatus represents the status of a pacemaker resource scheduled on a node. + PacemakerClusterResourceStatus represents the status of a resource tracked on a node. A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. - For Two Node OpenShift with Fencing, we track two resources per node: + For Two Node OpenShift with Fencing, we track the following pacemaker-managed resources per node: - Kubelet (the Kubernetes node agent and a prerequisite for etcd) - Etcd (the distributed key-value store) + The same type is reused to track alert-agent scripts (TaintAlertAgent, UntaintAlertAgent). An alert + agent is not a pacemaker-managed resource, so only a subset of the pacemaker condition types applies + to it: the required conditions are enforced conditionally based on the resource name. + Fencing agents are tracked separately in the fencingAgents field because they are mapped to their target node (the node they can fence), not the node where monitoring operations are scheduled. properties: @@ -483,7 +491,15 @@ spec: The "Active" condition tracks whether the resource is active (available to be used). The "Started" condition tracks whether the resource is started. The "Schedulable" condition tracks whether the resource is schedulable (not blocked). - Each of these conditions is required, so the array must contain at least 8 items. + Which conditions are required depends on the resource name: + - For the pacemaker-managed resources "Kubelet" and "Etcd", all eight condition types listed + above are required, so the array must contain at least 8 items. + - For the alert-agent resources "TaintAlertAgent" and "UntaintAlertAgent", only "Healthy", + "Enabled" (reason "ScriptConfigured"), and "Operational" (reason "ScriptPresent") are + required, so the array must contain at least 3 items. The remaining condition types do not + apply to an alert agent and may be omitted. + The array must contain at least 3 items in all cases; the exact set of required conditions for + each resource name is enforced by name-gated validation rules on this type. items: description: Condition contains details for one aspect of the current state of this API Resource. @@ -541,51 +557,42 @@ spec: - type type: object maxItems: 16 - minItems: 8 + minItems: 3 type: array x-kubernetes-list-map-keys: - type x-kubernetes-list-type: map - x-kubernetes-validations: - - message: conditions must contain a condition of type - Healthy - rule: self.exists(c, c.type == 'Healthy') - - message: conditions must contain a condition of type - InService - rule: self.exists(c, c.type == 'InService') - - message: conditions must contain a condition of type - Managed - rule: self.exists(c, c.type == 'Managed') - - message: conditions must contain a condition of type - Enabled - rule: self.exists(c, c.type == 'Enabled') - - message: conditions must contain a condition of type - Operational - rule: self.exists(c, c.type == 'Operational') - - message: conditions must contain a condition of type - Active - rule: self.exists(c, c.type == 'Active') - - message: conditions must contain a condition of type - Started - rule: self.exists(c, c.type == 'Started') - - message: conditions must contain a condition of type - Schedulable - rule: self.exists(c, c.type == 'Schedulable') name: description: |- - name is the name of the pacemaker resource. - Valid values are "Kubelet" and "Etcd". + name is the name of the resource. + Valid values are "Kubelet", "Etcd", "TaintAlertAgent", and "UntaintAlertAgent". The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. + The TaintAlertAgent and UntaintAlertAgent entries track alert-agent configuration and per-node + script presence rather than a pacemaker-managed resource. Fencing agents are tracked separately in the node's fencingAgents field. enum: - Kubelet - Etcd + - TaintAlertAgent + - UntaintAlertAgent type: string required: - conditions - name type: object + x-kubernetes-validations: + - message: conditions must contain Healthy, InService, Managed, + Enabled, Operational, Active, Started, and Schedulable + for Kubelet and Etcd resources + rule: '!(self.name == ''Kubelet'' || self.name == ''Etcd'') + || [''Healthy'',''InService'',''Managed'',''Enabled'',''Operational'',''Active'',''Started'',''Schedulable''].all(t, + self.conditions.exists(c, c.type == t))' + - message: conditions must contain Healthy, Enabled, and Operational + for TaintAlertAgent and UntaintAlertAgent resources + rule: '!(self.name == ''TaintAlertAgent'' || self.name == + ''UntaintAlertAgent'') || [''Healthy'',''Enabled'',''Operational''].all(t, + self.conditions.exists(c, c.type == t))' maxItems: 8 minItems: 2 type: array diff --git a/etcd/v1alpha1/zz_generated.featuregated-crd-manifests/pacemakerclusters.etcd.openshift.io/DualReplica.yaml b/etcd/v1alpha1/zz_generated.featuregated-crd-manifests/pacemakerclusters.etcd.openshift.io/DualReplica.yaml index e867f16f114..8ae1d60291e 100644 --- a/etcd/v1alpha1/zz_generated.featuregated-crd-manifests/pacemakerclusters.etcd.openshift.io/DualReplica.yaml +++ b/etcd/v1alpha1/zz_generated.featuregated-crd-manifests/pacemakerclusters.etcd.openshift.io/DualReplica.yaml @@ -59,7 +59,7 @@ spec: The "Healthy" condition is an aggregate that tracks the overall health of the cluster. The "InService" condition tracks whether the cluster is in service (not in maintenance mode). The "NodeCountAsExpected" condition tracks whether the expected number of nodes are present. - Each of these conditions is required, so the array must contain at least 3 items. + Each of these three conditions is required, so the array must contain at least 3 items. items: description: Condition contains details for one aspect of the current state of this API Resource. @@ -453,21 +453,29 @@ spec: rule: '!format.dns1123Subdomain().validate(self).hasValue()' resources: description: |- - resources contains the status of pacemaker resources scheduled on this node. + resources contains the status of pacemaker resources tracked on this node. Each resource entry includes the resource name and its health conditions. - For Two Node OpenShift with Fencing, we track Kubelet and Etcd resources per node. - Both resources are required to be present, so the array must contain at least 2 items. - Valid resource names are "Kubelet" and "Etcd". + For Two Node OpenShift with Fencing, we track the Kubelet and Etcd pacemaker-managed resources + per node. Both are required to be present, so the array must contain at least 2 items. + The array may also contain optional alert-agent script entries named "TaintAlertAgent" and + "UntaintAlertAgent". Their entries track whether the alert agent is configured in the CIB + and whether its script is present on this node. Alert-agent entries are + optional and are omitted when a status collector version that predates them has not reported them. + Valid resource names are "Kubelet", "Etcd", "TaintAlertAgent", and "UntaintAlertAgent". Fencing agents are tracked separately in the fencingAgents field. items: description: |- - PacemakerClusterResourceStatus represents the status of a pacemaker resource scheduled on a node. + PacemakerClusterResourceStatus represents the status of a resource tracked on a node. A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. - For Two Node OpenShift with Fencing, we track two resources per node: + For Two Node OpenShift with Fencing, we track the following pacemaker-managed resources per node: - Kubelet (the Kubernetes node agent and a prerequisite for etcd) - Etcd (the distributed key-value store) + The same type is reused to track alert-agent scripts (TaintAlertAgent, UntaintAlertAgent). An alert + agent is not a pacemaker-managed resource, so only a subset of the pacemaker condition types applies + to it: the required conditions are enforced conditionally based on the resource name. + Fencing agents are tracked separately in the fencingAgents field because they are mapped to their target node (the node they can fence), not the node where monitoring operations are scheduled. properties: @@ -484,7 +492,15 @@ spec: The "Active" condition tracks whether the resource is active (available to be used). The "Started" condition tracks whether the resource is started. The "Schedulable" condition tracks whether the resource is schedulable (not blocked). - Each of these conditions is required, so the array must contain at least 8 items. + Which conditions are required depends on the resource name: + - For the pacemaker-managed resources "Kubelet" and "Etcd", all eight condition types listed + above are required, so the array must contain at least 8 items. + - For the alert-agent resources "TaintAlertAgent" and "UntaintAlertAgent", only "Healthy", + "Enabled" (reason "ScriptConfigured"), and "Operational" (reason "ScriptPresent") are + required, so the array must contain at least 3 items. The remaining condition types do not + apply to an alert agent and may be omitted. + The array must contain at least 3 items in all cases; the exact set of required conditions for + each resource name is enforced by name-gated validation rules on this type. items: description: Condition contains details for one aspect of the current state of this API Resource. @@ -542,51 +558,42 @@ spec: - type type: object maxItems: 16 - minItems: 8 + minItems: 3 type: array x-kubernetes-list-map-keys: - type x-kubernetes-list-type: map - x-kubernetes-validations: - - message: conditions must contain a condition of type - Healthy - rule: self.exists(c, c.type == 'Healthy') - - message: conditions must contain a condition of type - InService - rule: self.exists(c, c.type == 'InService') - - message: conditions must contain a condition of type - Managed - rule: self.exists(c, c.type == 'Managed') - - message: conditions must contain a condition of type - Enabled - rule: self.exists(c, c.type == 'Enabled') - - message: conditions must contain a condition of type - Operational - rule: self.exists(c, c.type == 'Operational') - - message: conditions must contain a condition of type - Active - rule: self.exists(c, c.type == 'Active') - - message: conditions must contain a condition of type - Started - rule: self.exists(c, c.type == 'Started') - - message: conditions must contain a condition of type - Schedulable - rule: self.exists(c, c.type == 'Schedulable') name: description: |- - name is the name of the pacemaker resource. - Valid values are "Kubelet" and "Etcd". + name is the name of the resource. + Valid values are "Kubelet", "Etcd", "TaintAlertAgent", and "UntaintAlertAgent". The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. + The TaintAlertAgent and UntaintAlertAgent entries track alert-agent configuration and per-node + script presence rather than a pacemaker-managed resource. Fencing agents are tracked separately in the node's fencingAgents field. enum: - Kubelet - Etcd + - TaintAlertAgent + - UntaintAlertAgent type: string required: - conditions - name type: object + x-kubernetes-validations: + - message: conditions must contain Healthy, InService, Managed, + Enabled, Operational, Active, Started, and Schedulable + for Kubelet and Etcd resources + rule: '!(self.name == ''Kubelet'' || self.name == ''Etcd'') + || [''Healthy'',''InService'',''Managed'',''Enabled'',''Operational'',''Active'',''Started'',''Schedulable''].all(t, + self.conditions.exists(c, c.type == t))' + - message: conditions must contain Healthy, Enabled, and Operational + for TaintAlertAgent and UntaintAlertAgent resources + rule: '!(self.name == ''TaintAlertAgent'' || self.name == + ''UntaintAlertAgent'') || [''Healthy'',''Enabled'',''Operational''].all(t, + self.conditions.exists(c, c.type == t))' maxItems: 8 minItems: 2 type: array diff --git a/etcd/v1alpha1/zz_generated.swagger_doc_generated.go b/etcd/v1alpha1/zz_generated.swagger_doc_generated.go index dc6f224288e..7ae6e138a2f 100644 --- a/etcd/v1alpha1/zz_generated.swagger_doc_generated.go +++ b/etcd/v1alpha1/zz_generated.swagger_doc_generated.go @@ -47,7 +47,7 @@ var map_PacemakerClusterNodeStatus = map[string]string{ "conditions": "conditions represent the observations of the node's current state. Known condition types are: \"Healthy\", \"Online\", \"InService\", \"Active\", \"Ready\", \"Clean\", \"Member\", \"FencingAvailable\", \"FencingHealthy\". The \"Healthy\" condition is an aggregate that tracks the overall health of the node. The \"Online\" condition tracks whether the node is online. The \"InService\" condition tracks whether the node is in service (not in maintenance mode). The \"Active\" condition tracks whether the node is active (not in standby mode). The \"Ready\" condition tracks whether the node is ready (not in a pending state). The \"Clean\" condition tracks whether the node is in a clean (status known) state. The \"Member\" condition tracks whether the node is a member of the cluster. The \"FencingAvailable\" condition tracks whether this node can be fenced by at least one healthy agent. The \"FencingHealthy\" condition tracks whether all fencing agents for this node are healthy. Each of these conditions is required, so the array must contain at least 9 items.", "nodeName": "nodeName is the name of the node. This is expected to match the Kubernetes node's name, which must be a lowercase RFC 1123 subdomain consisting of lowercase alphanumeric characters, '-' or '.', starting and ending with an alphanumeric character, and be at most 253 characters in length.", "addresses": "addresses is a list of IP addresses for the node. Pacemaker allows multiple IP addresses for Corosync communication between nodes. The first address in this list is used for IP-based peer URLs for etcd membership. Each address must be a valid global unicast IPv4 or IPv6 address in canonical form (e.g., \"192.168.1.1\" not \"192.168.001.001\", or \"2001:db8::1\" not \"2001:0db8::1\"). This excludes loopback, link-local, and multicast addresses.", - "resources": "resources contains the status of pacemaker resources scheduled on this node. Each resource entry includes the resource name and its health conditions. For Two Node OpenShift with Fencing, we track Kubelet and Etcd resources per node. Both resources are required to be present, so the array must contain at least 2 items. Valid resource names are \"Kubelet\" and \"Etcd\". Fencing agents are tracked separately in the fencingAgents field.", + "resources": "resources contains the status of pacemaker resources tracked on this node. Each resource entry includes the resource name and its health conditions. For Two Node OpenShift with Fencing, we track the Kubelet and Etcd pacemaker-managed resources per node. Both are required to be present, so the array must contain at least 2 items. The array may also contain optional alert-agent script entries named \"TaintAlertAgent\" and \"UntaintAlertAgent\". Their entries track whether the alert agent is configured in the CIB and whether its script is present on this node. Alert-agent entries are optional and are omitted when a status collector version that predates them has not reported them. Valid resource names are \"Kubelet\", \"Etcd\", \"TaintAlertAgent\", and \"UntaintAlertAgent\". Fencing agents are tracked separately in the fencingAgents field.", "fencingAgents": "fencingAgents contains the status of fencing agents that can fence this node. Unlike resources (which are scheduled to run on this node), fencing agents are mapped to the node they can fence (their target), not the node where monitoring operations run. Each fencing agent entry includes a unique name, fencing type, target node, and health conditions. A node is considered fence-capable if at least one fencing agent is healthy. A healthy node is expected to have at least 1 fencing agent, but the list may be empty when fencing agent discovery fails. Names must be unique within this array.", } @@ -56,9 +56,9 @@ func (PacemakerClusterNodeStatus) SwaggerDoc() map[string]string { } var map_PacemakerClusterResourceStatus = map[string]string{ - "": "PacemakerClusterResourceStatus represents the status of a pacemaker resource scheduled on a node. A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. For Two Node OpenShift with Fencing, we track two resources per node:\n - Kubelet (the Kubernetes node agent and a prerequisite for etcd)\n - Etcd (the distributed key-value store)\n\nFencing agents are tracked separately in the fencingAgents field because they are mapped to their target node (the node they can fence), not the node where monitoring operations are scheduled.", - "conditions": "conditions represent the observations of the resource's current state. Known condition types are: \"Healthy\", \"InService\", \"Managed\", \"Enabled\", \"Operational\", \"Active\", \"Started\", \"Schedulable\". The \"Healthy\" condition is an aggregate that tracks the overall health of the resource. The \"InService\" condition tracks whether the resource is in service (not in maintenance mode). The \"Managed\" condition tracks whether the resource is managed by pacemaker. The \"Enabled\" condition tracks whether the resource is enabled. The \"Operational\" condition tracks whether the resource is operational (not failed). The \"Active\" condition tracks whether the resource is active (available to be used). The \"Started\" condition tracks whether the resource is started. The \"Schedulable\" condition tracks whether the resource is schedulable (not blocked). Each of these conditions is required, so the array must contain at least 8 items.", - "name": "name is the name of the pacemaker resource. Valid values are \"Kubelet\" and \"Etcd\". The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. Fencing agents are tracked separately in the node's fencingAgents field.", + "": "PacemakerClusterResourceStatus represents the status of a resource tracked on a node. A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. For Two Node OpenShift with Fencing, we track the following pacemaker-managed resources per node:\n - Kubelet (the Kubernetes node agent and a prerequisite for etcd)\n - Etcd (the distributed key-value store)\n\nThe same type is reused to track alert-agent scripts (TaintAlertAgent, UntaintAlertAgent). An alert agent is not a pacemaker-managed resource, so only a subset of the pacemaker condition types applies to it: the required conditions are enforced conditionally based on the resource name.\n\nFencing agents are tracked separately in the fencingAgents field because they are mapped to their target node (the node they can fence), not the node where monitoring operations are scheduled.", + "conditions": "conditions represent the observations of the resource's current state. Known condition types are: \"Healthy\", \"InService\", \"Managed\", \"Enabled\", \"Operational\", \"Active\", \"Started\", \"Schedulable\". The \"Healthy\" condition is an aggregate that tracks the overall health of the resource. The \"InService\" condition tracks whether the resource is in service (not in maintenance mode). The \"Managed\" condition tracks whether the resource is managed by pacemaker. The \"Enabled\" condition tracks whether the resource is enabled. The \"Operational\" condition tracks whether the resource is operational (not failed). The \"Active\" condition tracks whether the resource is active (available to be used). The \"Started\" condition tracks whether the resource is started. The \"Schedulable\" condition tracks whether the resource is schedulable (not blocked). Which conditions are required depends on the resource name:\n - For the pacemaker-managed resources \"Kubelet\" and \"Etcd\", all eight condition types listed\n above are required, so the array must contain at least 8 items.\n - For the alert-agent resources \"TaintAlertAgent\" and \"UntaintAlertAgent\", only \"Healthy\",\n \"Enabled\" (reason \"ScriptConfigured\"), and \"Operational\" (reason \"ScriptPresent\") are\n required, so the array must contain at least 3 items. The remaining condition types do not\n apply to an alert agent and may be omitted.\nThe array must contain at least 3 items in all cases; the exact set of required conditions for each resource name is enforced by name-gated validation rules on this type.", + "name": "name is the name of the resource. Valid values are \"Kubelet\", \"Etcd\", \"TaintAlertAgent\", and \"UntaintAlertAgent\". The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. The TaintAlertAgent and UntaintAlertAgent entries track alert-agent configuration and per-node script presence rather than a pacemaker-managed resource. Fencing agents are tracked separately in the node's fencingAgents field.", } func (PacemakerClusterResourceStatus) SwaggerDoc() map[string]string { @@ -67,7 +67,7 @@ func (PacemakerClusterResourceStatus) SwaggerDoc() map[string]string { var map_PacemakerClusterStatus = map[string]string{ "": "PacemakerClusterStatus contains the actual pacemaker cluster status information. As part of validating the status object, we need to ensure that the lastUpdated timestamp may not be set to an earlier timestamp than the current value. The validation rule checks if oldSelf has lastUpdated before comparing, to handle the initial status creation case.", - "conditions": "conditions represent the observations of the pacemaker cluster's current state. Known condition types are: \"Healthy\", \"InService\", \"NodeCountAsExpected\". The \"Healthy\" condition is an aggregate that tracks the overall health of the cluster. The \"InService\" condition tracks whether the cluster is in service (not in maintenance mode). The \"NodeCountAsExpected\" condition tracks whether the expected number of nodes are present. Each of these conditions is required, so the array must contain at least 3 items.", + "conditions": "conditions represent the observations of the pacemaker cluster's current state. Known condition types are: \"Healthy\", \"InService\", \"NodeCountAsExpected\". The \"Healthy\" condition is an aggregate that tracks the overall health of the cluster. The \"InService\" condition tracks whether the cluster is in service (not in maintenance mode). The \"NodeCountAsExpected\" condition tracks whether the expected number of nodes are present. Each of these three conditions is required, so the array must contain at least 3 items.", "lastUpdated": "lastUpdated is the timestamp when this status was last updated. This is useful for identifying stale status reports. It must be a valid timestamp in RFC3339 format. Once set, this field cannot be removed and cannot be set to an earlier timestamp than the current value.", "nodes": "nodes provides detailed status for each control-plane node in the Pacemaker cluster. While Pacemaker supports up to 32 nodes, the limit is set to 5 (max OpenShift control-plane nodes). For Two Node OpenShift with Fencing, exactly 2 nodes are expected in a healthy cluster. An empty list indicates a catastrophic failure where Pacemaker reports no nodes.", } diff --git a/openapi/generated_openapi/zz_generated.openapi.go b/openapi/generated_openapi/zz_generated.openapi.go index 16f4728a9ae..078e9ae081f 100644 --- a/openapi/generated_openapi/zz_generated.openapi.go +++ b/openapi/generated_openapi/zz_generated.openapi.go @@ -30263,7 +30263,7 @@ func schema_openshift_api_etcd_v1_PacemakerClusterNodeStatus(ref common.Referenc }, }, SchemaProps: spec.SchemaProps{ - Description: "resources contains the status of pacemaker resources scheduled on this node. Each resource entry includes the resource name and its health conditions. For Two Node OpenShift with Fencing, we track Kubelet and Etcd resources per node. Both resources are required to be present, so the array must contain at least 2 items. Valid resource names are \"Kubelet\" and \"Etcd\". Fencing agents are tracked separately in the fencingAgents field.", + Description: "resources contains the status of pacemaker resources tracked on this node. Each resource entry includes the resource name and its health conditions. For Two Node OpenShift with Fencing, we track the Kubelet and Etcd pacemaker-managed resources per node. Both are required to be present, so the array must contain at least 2 items. The array may also contain optional alert-agent script entries named \"TaintAlertAgent\" and \"UntaintAlertAgent\". Their entries track whether the alert agent is configured in the CIB and whether its script is present on this node. Alert-agent entries are optional and are omitted when a status collector version that predates them has not reported them. Valid resource names are \"Kubelet\", \"Etcd\", \"TaintAlertAgent\", and \"UntaintAlertAgent\". Fencing agents are tracked separately in the fencingAgents field.", Type: []string{"array"}, Items: &spec.SchemaOrArray{ Schema: &spec.Schema{ @@ -30310,7 +30310,7 @@ func schema_openshift_api_etcd_v1_PacemakerClusterResourceStatus(ref common.Refe return common.OpenAPIDefinition{ Schema: spec.Schema{ SchemaProps: spec.SchemaProps{ - Description: "PacemakerClusterResourceStatus represents the status of a pacemaker resource scheduled on a node. A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. For Two Node OpenShift with Fencing, we track two resources per node:\n - Kubelet (the Kubernetes node agent and a prerequisite for etcd)\n - Etcd (the distributed key-value store)\n\nFencing agents are tracked separately in the fencingAgents field because they are mapped to their target node (the node they can fence), not the node where monitoring operations are scheduled.", + Description: "PacemakerClusterResourceStatus represents the status of a resource tracked on a node. A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. For Two Node OpenShift with Fencing, we track the following pacemaker-managed resources per node:\n - Kubelet (the Kubernetes node agent and a prerequisite for etcd)\n - Etcd (the distributed key-value store)\n\nThe same type is reused to track alert-agent scripts (TaintAlertAgent, UntaintAlertAgent). An alert agent is not a pacemaker-managed resource, so only a subset of the pacemaker condition types applies to it: the required conditions are enforced conditionally based on the resource name.\n\nFencing agents are tracked separately in the fencingAgents field because they are mapped to their target node (the node they can fence), not the node where monitoring operations are scheduled.", Type: []string{"object"}, Properties: map[string]spec.Schema{ "conditions": { @@ -30323,7 +30323,7 @@ func schema_openshift_api_etcd_v1_PacemakerClusterResourceStatus(ref common.Refe }, }, SchemaProps: spec.SchemaProps{ - Description: "conditions represent the observations of the resource's current state. Known condition types are: \"Healthy\", \"InService\", \"Managed\", \"Enabled\", \"Operational\", \"Active\", \"Started\", \"Schedulable\". The \"Healthy\" condition is an aggregate that tracks the overall health of the resource. The \"InService\" condition tracks whether the resource is in service (not in maintenance mode). The \"Managed\" condition tracks whether the resource is managed by pacemaker. The \"Enabled\" condition tracks whether the resource is enabled. The \"Operational\" condition tracks whether the resource is operational (not failed). The \"Active\" condition tracks whether the resource is active (available to be used). The \"Started\" condition tracks whether the resource is started. The \"Schedulable\" condition tracks whether the resource is schedulable (not blocked). Each of these conditions is required, so the array must contain at least 8 items.", + Description: "conditions represent the observations of the resource's current state. Known condition types are: \"Healthy\", \"InService\", \"Managed\", \"Enabled\", \"Operational\", \"Active\", \"Started\", \"Schedulable\". The \"Healthy\" condition is an aggregate that tracks the overall health of the resource. The \"InService\" condition tracks whether the resource is in service (not in maintenance mode). The \"Managed\" condition tracks whether the resource is managed by pacemaker. The \"Enabled\" condition tracks whether the resource is enabled. The \"Operational\" condition tracks whether the resource is operational (not failed). The \"Active\" condition tracks whether the resource is active (available to be used). The \"Started\" condition tracks whether the resource is started. The \"Schedulable\" condition tracks whether the resource is schedulable (not blocked). Which conditions are required depends on the resource name:\n - For the pacemaker-managed resources \"Kubelet\" and \"Etcd\", all eight condition types listed\n above are required, so the array must contain at least 8 items.\n - For the alert-agent resources \"TaintAlertAgent\" and \"UntaintAlertAgent\", only \"Healthy\",\n \"Enabled\" (reason \"ScriptConfigured\"), and \"Operational\" (reason \"ScriptPresent\") are\n required, so the array must contain at least 3 items. The remaining condition types do not\n apply to an alert agent and may be omitted.\nThe array must contain at least 3 items in all cases; the exact set of required conditions for each resource name is enforced by name-gated validation rules on this type.", Type: []string{"array"}, Items: &spec.SchemaOrArray{ Schema: &spec.Schema{ @@ -30337,10 +30337,10 @@ func schema_openshift_api_etcd_v1_PacemakerClusterResourceStatus(ref common.Refe }, "name": { SchemaProps: spec.SchemaProps{ - Description: "name is the name of the pacemaker resource. Valid values are \"Kubelet\" and \"Etcd\". The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. Fencing agents are tracked separately in the node's fencingAgents field.\n\nPossible enum values:\n - `\"Etcd\"` is the etcd pacemaker resource. The etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations.\n - `\"Kubelet\"` is the kubelet pacemaker resource. The kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments.", + Description: "name is the name of the resource. Valid values are \"Kubelet\", \"Etcd\", \"TaintAlertAgent\", and \"UntaintAlertAgent\". The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. The TaintAlertAgent and UntaintAlertAgent entries track alert-agent configuration and per-node script presence rather than a pacemaker-managed resource. Fencing agents are tracked separately in the node's fencingAgents field.\n\nPossible enum values:\n - `\"Etcd\"` is the etcd pacemaker resource. The etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations.\n - `\"Kubelet\"` is the kubelet pacemaker resource. The kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments.\n - `\"TaintAlertAgent\"` is the alert agent that taints a node after it is fenced. Its entry tracks the alert agent's configuration in the CIB and the presence of its script on this node.\n - `\"UntaintAlertAgent\"` is the alert agent that removes a node's taint once it rejoins the cluster. Its entry tracks the alert agent's configuration in the CIB and the presence of its script on this node.", Type: []string{"string"}, Format: "", - Enum: []interface{}{"Etcd", "Kubelet"}, + Enum: []interface{}{"Etcd", "Kubelet", "TaintAlertAgent", "UntaintAlertAgent"}, }, }, }, @@ -30369,7 +30369,7 @@ func schema_openshift_api_etcd_v1_PacemakerClusterStatus(ref common.ReferenceCal }, }, SchemaProps: spec.SchemaProps{ - Description: "conditions represent the observations of the pacemaker cluster's current state. Known condition types are: \"Healthy\", \"InService\", \"NodeCountAsExpected\". The \"Healthy\" condition is an aggregate that tracks the overall health of the cluster. The \"InService\" condition tracks whether the cluster is in service (not in maintenance mode). The \"NodeCountAsExpected\" condition tracks whether the expected number of nodes are present. Each of these conditions is required, so the array must contain at least 3 items.", + Description: "conditions represent the observations of the pacemaker cluster's current state. Known condition types are: \"Healthy\", \"InService\", \"NodeCountAsExpected\". The \"Healthy\" condition is an aggregate that tracks the overall health of the cluster. The \"InService\" condition tracks whether the cluster is in service (not in maintenance mode). The \"NodeCountAsExpected\" condition tracks whether the expected number of nodes are present. Each of these three conditions is required, so the array must contain at least 3 items.", Type: []string{"array"}, Items: &spec.SchemaOrArray{ Schema: &spec.Schema{ @@ -30660,7 +30660,7 @@ func schema_openshift_api_etcd_v1alpha1_PacemakerClusterNodeStatus(ref common.Re }, }, SchemaProps: spec.SchemaProps{ - Description: "resources contains the status of pacemaker resources scheduled on this node. Each resource entry includes the resource name and its health conditions. For Two Node OpenShift with Fencing, we track Kubelet and Etcd resources per node. Both resources are required to be present, so the array must contain at least 2 items. Valid resource names are \"Kubelet\" and \"Etcd\". Fencing agents are tracked separately in the fencingAgents field.", + Description: "resources contains the status of pacemaker resources tracked on this node. Each resource entry includes the resource name and its health conditions. For Two Node OpenShift with Fencing, we track the Kubelet and Etcd pacemaker-managed resources per node. Both are required to be present, so the array must contain at least 2 items. The array may also contain optional alert-agent script entries named \"TaintAlertAgent\" and \"UntaintAlertAgent\". Their entries track whether the alert agent is configured in the CIB and whether its script is present on this node. Alert-agent entries are optional and are omitted when a status collector version that predates them has not reported them. Valid resource names are \"Kubelet\", \"Etcd\", \"TaintAlertAgent\", and \"UntaintAlertAgent\". Fencing agents are tracked separately in the fencingAgents field.", Type: []string{"array"}, Items: &spec.SchemaOrArray{ Schema: &spec.Schema{ @@ -30707,7 +30707,7 @@ func schema_openshift_api_etcd_v1alpha1_PacemakerClusterResourceStatus(ref commo return common.OpenAPIDefinition{ Schema: spec.Schema{ SchemaProps: spec.SchemaProps{ - Description: "PacemakerClusterResourceStatus represents the status of a pacemaker resource scheduled on a node. A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. For Two Node OpenShift with Fencing, we track two resources per node:\n - Kubelet (the Kubernetes node agent and a prerequisite for etcd)\n - Etcd (the distributed key-value store)\n\nFencing agents are tracked separately in the fencingAgents field because they are mapped to their target node (the node they can fence), not the node where monitoring operations are scheduled.", + Description: "PacemakerClusterResourceStatus represents the status of a resource tracked on a node. A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. For Two Node OpenShift with Fencing, we track the following pacemaker-managed resources per node:\n - Kubelet (the Kubernetes node agent and a prerequisite for etcd)\n - Etcd (the distributed key-value store)\n\nThe same type is reused to track alert-agent scripts (TaintAlertAgent, UntaintAlertAgent). An alert agent is not a pacemaker-managed resource, so only a subset of the pacemaker condition types applies to it: the required conditions are enforced conditionally based on the resource name.\n\nFencing agents are tracked separately in the fencingAgents field because they are mapped to their target node (the node they can fence), not the node where monitoring operations are scheduled.", Type: []string{"object"}, Properties: map[string]spec.Schema{ "conditions": { @@ -30720,7 +30720,7 @@ func schema_openshift_api_etcd_v1alpha1_PacemakerClusterResourceStatus(ref commo }, }, SchemaProps: spec.SchemaProps{ - Description: "conditions represent the observations of the resource's current state. Known condition types are: \"Healthy\", \"InService\", \"Managed\", \"Enabled\", \"Operational\", \"Active\", \"Started\", \"Schedulable\". The \"Healthy\" condition is an aggregate that tracks the overall health of the resource. The \"InService\" condition tracks whether the resource is in service (not in maintenance mode). The \"Managed\" condition tracks whether the resource is managed by pacemaker. The \"Enabled\" condition tracks whether the resource is enabled. The \"Operational\" condition tracks whether the resource is operational (not failed). The \"Active\" condition tracks whether the resource is active (available to be used). The \"Started\" condition tracks whether the resource is started. The \"Schedulable\" condition tracks whether the resource is schedulable (not blocked). Each of these conditions is required, so the array must contain at least 8 items.", + Description: "conditions represent the observations of the resource's current state. Known condition types are: \"Healthy\", \"InService\", \"Managed\", \"Enabled\", \"Operational\", \"Active\", \"Started\", \"Schedulable\". The \"Healthy\" condition is an aggregate that tracks the overall health of the resource. The \"InService\" condition tracks whether the resource is in service (not in maintenance mode). The \"Managed\" condition tracks whether the resource is managed by pacemaker. The \"Enabled\" condition tracks whether the resource is enabled. The \"Operational\" condition tracks whether the resource is operational (not failed). The \"Active\" condition tracks whether the resource is active (available to be used). The \"Started\" condition tracks whether the resource is started. The \"Schedulable\" condition tracks whether the resource is schedulable (not blocked). Which conditions are required depends on the resource name:\n - For the pacemaker-managed resources \"Kubelet\" and \"Etcd\", all eight condition types listed\n above are required, so the array must contain at least 8 items.\n - For the alert-agent resources \"TaintAlertAgent\" and \"UntaintAlertAgent\", only \"Healthy\",\n \"Enabled\" (reason \"ScriptConfigured\"), and \"Operational\" (reason \"ScriptPresent\") are\n required, so the array must contain at least 3 items. The remaining condition types do not\n apply to an alert agent and may be omitted.\nThe array must contain at least 3 items in all cases; the exact set of required conditions for each resource name is enforced by name-gated validation rules on this type.", Type: []string{"array"}, Items: &spec.SchemaOrArray{ Schema: &spec.Schema{ @@ -30734,10 +30734,10 @@ func schema_openshift_api_etcd_v1alpha1_PacemakerClusterResourceStatus(ref commo }, "name": { SchemaProps: spec.SchemaProps{ - Description: "name is the name of the pacemaker resource. Valid values are \"Kubelet\" and \"Etcd\". The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. Fencing agents are tracked separately in the node's fencingAgents field.\n\nPossible enum values:\n - `\"Etcd\"` is the etcd pacemaker resource. The etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations.\n - `\"Kubelet\"` is the kubelet pacemaker resource. The kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments.", + Description: "name is the name of the resource. Valid values are \"Kubelet\", \"Etcd\", \"TaintAlertAgent\", and \"UntaintAlertAgent\". The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. The TaintAlertAgent and UntaintAlertAgent entries track alert-agent configuration and per-node script presence rather than a pacemaker-managed resource. Fencing agents are tracked separately in the node's fencingAgents field.\n\nPossible enum values:\n - `\"Etcd\"` is the etcd pacemaker resource. The etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations.\n - `\"Kubelet\"` is the kubelet pacemaker resource. The kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments.\n - `\"TaintAlertAgent\"` is the alert agent that taints a node after it is fenced. Its entry tracks the alert agent's configuration in the CIB and the presence of its script on this node.\n - `\"UntaintAlertAgent\"` is the alert agent that removes a node's taint once it rejoins the cluster. Its entry tracks the alert agent's configuration in the CIB and the presence of its script on this node.", Type: []string{"string"}, Format: "", - Enum: []interface{}{"Etcd", "Kubelet"}, + Enum: []interface{}{"Etcd", "Kubelet", "TaintAlertAgent", "UntaintAlertAgent"}, }, }, }, @@ -30766,7 +30766,7 @@ func schema_openshift_api_etcd_v1alpha1_PacemakerClusterStatus(ref common.Refere }, }, SchemaProps: spec.SchemaProps{ - Description: "conditions represent the observations of the pacemaker cluster's current state. Known condition types are: \"Healthy\", \"InService\", \"NodeCountAsExpected\". The \"Healthy\" condition is an aggregate that tracks the overall health of the cluster. The \"InService\" condition tracks whether the cluster is in service (not in maintenance mode). The \"NodeCountAsExpected\" condition tracks whether the expected number of nodes are present. Each of these conditions is required, so the array must contain at least 3 items.", + Description: "conditions represent the observations of the pacemaker cluster's current state. Known condition types are: \"Healthy\", \"InService\", \"NodeCountAsExpected\". The \"Healthy\" condition is an aggregate that tracks the overall health of the cluster. The \"InService\" condition tracks whether the cluster is in service (not in maintenance mode). The \"NodeCountAsExpected\" condition tracks whether the expected number of nodes are present. Each of these three conditions is required, so the array must contain at least 3 items.", Type: []string{"array"}, Items: &spec.SchemaOrArray{ Schema: &spec.Schema{ diff --git a/openapi/openapi.json b/openapi/openapi.json index 99f806ce8fe..c1e3fd1e5a1 100644 --- a/openapi/openapi.json +++ b/openapi/openapi.json @@ -16872,7 +16872,7 @@ "type": "string" }, "resources": { - "description": "resources contains the status of pacemaker resources scheduled on this node. Each resource entry includes the resource name and its health conditions. For Two Node OpenShift with Fencing, we track Kubelet and Etcd resources per node. Both resources are required to be present, so the array must contain at least 2 items. Valid resource names are \"Kubelet\" and \"Etcd\". Fencing agents are tracked separately in the fencingAgents field.", + "description": "resources contains the status of pacemaker resources tracked on this node. Each resource entry includes the resource name and its health conditions. For Two Node OpenShift with Fencing, we track the Kubelet and Etcd pacemaker-managed resources per node. Both are required to be present, so the array must contain at least 2 items. The array may also contain optional alert-agent script entries named \"TaintAlertAgent\" and \"UntaintAlertAgent\". Their entries track whether the alert agent is configured in the CIB and whether its script is present on this node. Alert-agent entries are optional and are omitted when a status collector version that predates them has not reported them. Valid resource names are \"Kubelet\", \"Etcd\", \"TaintAlertAgent\", and \"UntaintAlertAgent\". Fencing agents are tracked separately in the fencingAgents field.", "type": "array", "items": { "default": {}, @@ -16886,7 +16886,7 @@ } }, "com.github.openshift.api.etcd.v1.PacemakerClusterResourceStatus": { - "description": "PacemakerClusterResourceStatus represents the status of a pacemaker resource scheduled on a node. A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. For Two Node OpenShift with Fencing, we track two resources per node:\n - Kubelet (the Kubernetes node agent and a prerequisite for etcd)\n - Etcd (the distributed key-value store)\n\nFencing agents are tracked separately in the fencingAgents field because they are mapped to their target node (the node they can fence), not the node where monitoring operations are scheduled.", + "description": "PacemakerClusterResourceStatus represents the status of a resource tracked on a node. A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. For Two Node OpenShift with Fencing, we track the following pacemaker-managed resources per node:\n - Kubelet (the Kubernetes node agent and a prerequisite for etcd)\n - Etcd (the distributed key-value store)\n\nThe same type is reused to track alert-agent scripts (TaintAlertAgent, UntaintAlertAgent). An alert agent is not a pacemaker-managed resource, so only a subset of the pacemaker condition types applies to it: the required conditions are enforced conditionally based on the resource name.\n\nFencing agents are tracked separately in the fencingAgents field because they are mapped to their target node (the node they can fence), not the node where monitoring operations are scheduled.", "type": "object", "required": [ "conditions", @@ -16894,7 +16894,7 @@ ], "properties": { "conditions": { - "description": "conditions represent the observations of the resource's current state. Known condition types are: \"Healthy\", \"InService\", \"Managed\", \"Enabled\", \"Operational\", \"Active\", \"Started\", \"Schedulable\". The \"Healthy\" condition is an aggregate that tracks the overall health of the resource. The \"InService\" condition tracks whether the resource is in service (not in maintenance mode). The \"Managed\" condition tracks whether the resource is managed by pacemaker. The \"Enabled\" condition tracks whether the resource is enabled. The \"Operational\" condition tracks whether the resource is operational (not failed). The \"Active\" condition tracks whether the resource is active (available to be used). The \"Started\" condition tracks whether the resource is started. The \"Schedulable\" condition tracks whether the resource is schedulable (not blocked). Each of these conditions is required, so the array must contain at least 8 items.", + "description": "conditions represent the observations of the resource's current state. Known condition types are: \"Healthy\", \"InService\", \"Managed\", \"Enabled\", \"Operational\", \"Active\", \"Started\", \"Schedulable\". The \"Healthy\" condition is an aggregate that tracks the overall health of the resource. The \"InService\" condition tracks whether the resource is in service (not in maintenance mode). The \"Managed\" condition tracks whether the resource is managed by pacemaker. The \"Enabled\" condition tracks whether the resource is enabled. The \"Operational\" condition tracks whether the resource is operational (not failed). The \"Active\" condition tracks whether the resource is active (available to be used). The \"Started\" condition tracks whether the resource is started. The \"Schedulable\" condition tracks whether the resource is schedulable (not blocked). Which conditions are required depends on the resource name:\n - For the pacemaker-managed resources \"Kubelet\" and \"Etcd\", all eight condition types listed\n above are required, so the array must contain at least 8 items.\n - For the alert-agent resources \"TaintAlertAgent\" and \"UntaintAlertAgent\", only \"Healthy\",\n \"Enabled\" (reason \"ScriptConfigured\"), and \"Operational\" (reason \"ScriptPresent\") are\n required, so the array must contain at least 3 items. The remaining condition types do not\n apply to an alert agent and may be omitted.\nThe array must contain at least 3 items in all cases; the exact set of required conditions for each resource name is enforced by name-gated validation rules on this type.", "type": "array", "items": { "default": {}, @@ -16906,11 +16906,13 @@ "x-kubernetes-list-type": "map" }, "name": { - "description": "name is the name of the pacemaker resource. Valid values are \"Kubelet\" and \"Etcd\". The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. Fencing agents are tracked separately in the node's fencingAgents field.\n\nPossible enum values:\n - `\"Etcd\"` is the etcd pacemaker resource. The etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations.\n - `\"Kubelet\"` is the kubelet pacemaker resource. The kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments.", + "description": "name is the name of the resource. Valid values are \"Kubelet\", \"Etcd\", \"TaintAlertAgent\", and \"UntaintAlertAgent\". The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. The TaintAlertAgent and UntaintAlertAgent entries track alert-agent configuration and per-node script presence rather than a pacemaker-managed resource. Fencing agents are tracked separately in the node's fencingAgents field.\n\nPossible enum values:\n - `\"Etcd\"` is the etcd pacemaker resource. The etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations.\n - `\"Kubelet\"` is the kubelet pacemaker resource. The kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments.\n - `\"TaintAlertAgent\"` is the alert agent that taints a node after it is fenced. Its entry tracks the alert agent's configuration in the CIB and the presence of its script on this node.\n - `\"UntaintAlertAgent\"` is the alert agent that removes a node's taint once it rejoins the cluster. Its entry tracks the alert agent's configuration in the CIB and the presence of its script on this node.", "type": "string", "enum": [ "Etcd", - "Kubelet" + "Kubelet", + "TaintAlertAgent", + "UntaintAlertAgent" ] } } @@ -16925,7 +16927,7 @@ ], "properties": { "conditions": { - "description": "conditions represent the observations of the pacemaker cluster's current state. Known condition types are: \"Healthy\", \"InService\", \"NodeCountAsExpected\". The \"Healthy\" condition is an aggregate that tracks the overall health of the cluster. The \"InService\" condition tracks whether the cluster is in service (not in maintenance mode). The \"NodeCountAsExpected\" condition tracks whether the expected number of nodes are present. Each of these conditions is required, so the array must contain at least 3 items.", + "description": "conditions represent the observations of the pacemaker cluster's current state. Known condition types are: \"Healthy\", \"InService\", \"NodeCountAsExpected\". The \"Healthy\" condition is an aggregate that tracks the overall health of the cluster. The \"InService\" condition tracks whether the cluster is in service (not in maintenance mode). The \"NodeCountAsExpected\" condition tracks whether the expected number of nodes are present. Each of these three conditions is required, so the array must contain at least 3 items.", "type": "array", "items": { "default": {}, @@ -17116,7 +17118,7 @@ "type": "string" }, "resources": { - "description": "resources contains the status of pacemaker resources scheduled on this node. Each resource entry includes the resource name and its health conditions. For Two Node OpenShift with Fencing, we track Kubelet and Etcd resources per node. Both resources are required to be present, so the array must contain at least 2 items. Valid resource names are \"Kubelet\" and \"Etcd\". Fencing agents are tracked separately in the fencingAgents field.", + "description": "resources contains the status of pacemaker resources tracked on this node. Each resource entry includes the resource name and its health conditions. For Two Node OpenShift with Fencing, we track the Kubelet and Etcd pacemaker-managed resources per node. Both are required to be present, so the array must contain at least 2 items. The array may also contain optional alert-agent script entries named \"TaintAlertAgent\" and \"UntaintAlertAgent\". Their entries track whether the alert agent is configured in the CIB and whether its script is present on this node. Alert-agent entries are optional and are omitted when a status collector version that predates them has not reported them. Valid resource names are \"Kubelet\", \"Etcd\", \"TaintAlertAgent\", and \"UntaintAlertAgent\". Fencing agents are tracked separately in the fencingAgents field.", "type": "array", "items": { "default": {}, @@ -17130,7 +17132,7 @@ } }, "com.github.openshift.api.etcd.v1alpha1.PacemakerClusterResourceStatus": { - "description": "PacemakerClusterResourceStatus represents the status of a pacemaker resource scheduled on a node. A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. For Two Node OpenShift with Fencing, we track two resources per node:\n - Kubelet (the Kubernetes node agent and a prerequisite for etcd)\n - Etcd (the distributed key-value store)\n\nFencing agents are tracked separately in the fencingAgents field because they are mapped to their target node (the node they can fence), not the node where monitoring operations are scheduled.", + "description": "PacemakerClusterResourceStatus represents the status of a resource tracked on a node. A pacemaker resource is a unit of work managed by pacemaker. In pacemaker terminology, resources are services or applications that pacemaker monitors, starts, stops, and moves between nodes to maintain high availability. For Two Node OpenShift with Fencing, we track the following pacemaker-managed resources per node:\n - Kubelet (the Kubernetes node agent and a prerequisite for etcd)\n - Etcd (the distributed key-value store)\n\nThe same type is reused to track alert-agent scripts (TaintAlertAgent, UntaintAlertAgent). An alert agent is not a pacemaker-managed resource, so only a subset of the pacemaker condition types applies to it: the required conditions are enforced conditionally based on the resource name.\n\nFencing agents are tracked separately in the fencingAgents field because they are mapped to their target node (the node they can fence), not the node where monitoring operations are scheduled.", "type": "object", "required": [ "conditions", @@ -17138,7 +17140,7 @@ ], "properties": { "conditions": { - "description": "conditions represent the observations of the resource's current state. Known condition types are: \"Healthy\", \"InService\", \"Managed\", \"Enabled\", \"Operational\", \"Active\", \"Started\", \"Schedulable\". The \"Healthy\" condition is an aggregate that tracks the overall health of the resource. The \"InService\" condition tracks whether the resource is in service (not in maintenance mode). The \"Managed\" condition tracks whether the resource is managed by pacemaker. The \"Enabled\" condition tracks whether the resource is enabled. The \"Operational\" condition tracks whether the resource is operational (not failed). The \"Active\" condition tracks whether the resource is active (available to be used). The \"Started\" condition tracks whether the resource is started. The \"Schedulable\" condition tracks whether the resource is schedulable (not blocked). Each of these conditions is required, so the array must contain at least 8 items.", + "description": "conditions represent the observations of the resource's current state. Known condition types are: \"Healthy\", \"InService\", \"Managed\", \"Enabled\", \"Operational\", \"Active\", \"Started\", \"Schedulable\". The \"Healthy\" condition is an aggregate that tracks the overall health of the resource. The \"InService\" condition tracks whether the resource is in service (not in maintenance mode). The \"Managed\" condition tracks whether the resource is managed by pacemaker. The \"Enabled\" condition tracks whether the resource is enabled. The \"Operational\" condition tracks whether the resource is operational (not failed). The \"Active\" condition tracks whether the resource is active (available to be used). The \"Started\" condition tracks whether the resource is started. The \"Schedulable\" condition tracks whether the resource is schedulable (not blocked). Which conditions are required depends on the resource name:\n - For the pacemaker-managed resources \"Kubelet\" and \"Etcd\", all eight condition types listed\n above are required, so the array must contain at least 8 items.\n - For the alert-agent resources \"TaintAlertAgent\" and \"UntaintAlertAgent\", only \"Healthy\",\n \"Enabled\" (reason \"ScriptConfigured\"), and \"Operational\" (reason \"ScriptPresent\") are\n required, so the array must contain at least 3 items. The remaining condition types do not\n apply to an alert agent and may be omitted.\nThe array must contain at least 3 items in all cases; the exact set of required conditions for each resource name is enforced by name-gated validation rules on this type.", "type": "array", "items": { "default": {}, @@ -17150,11 +17152,13 @@ "x-kubernetes-list-type": "map" }, "name": { - "description": "name is the name of the pacemaker resource. Valid values are \"Kubelet\" and \"Etcd\". The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. Fencing agents are tracked separately in the node's fencingAgents field.\n\nPossible enum values:\n - `\"Etcd\"` is the etcd pacemaker resource. The etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations.\n - `\"Kubelet\"` is the kubelet pacemaker resource. The kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments.", + "description": "name is the name of the resource. Valid values are \"Kubelet\", \"Etcd\", \"TaintAlertAgent\", and \"UntaintAlertAgent\". The Kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments. The Etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations. The TaintAlertAgent and UntaintAlertAgent entries track alert-agent configuration and per-node script presence rather than a pacemaker-managed resource. Fencing agents are tracked separately in the node's fencingAgents field.\n\nPossible enum values:\n - `\"Etcd\"` is the etcd pacemaker resource. The etcd resource may temporarily transition to stopped during pacemaker quorum-recovery operations.\n - `\"Kubelet\"` is the kubelet pacemaker resource. The kubelet resource is a prerequisite for etcd in Two Node OpenShift with Fencing deployments.\n - `\"TaintAlertAgent\"` is the alert agent that taints a node after it is fenced. Its entry tracks the alert agent's configuration in the CIB and the presence of its script on this node.\n - `\"UntaintAlertAgent\"` is the alert agent that removes a node's taint once it rejoins the cluster. Its entry tracks the alert agent's configuration in the CIB and the presence of its script on this node.", "type": "string", "enum": [ "Etcd", - "Kubelet" + "Kubelet", + "TaintAlertAgent", + "UntaintAlertAgent" ] } } @@ -17169,7 +17173,7 @@ ], "properties": { "conditions": { - "description": "conditions represent the observations of the pacemaker cluster's current state. Known condition types are: \"Healthy\", \"InService\", \"NodeCountAsExpected\". The \"Healthy\" condition is an aggregate that tracks the overall health of the cluster. The \"InService\" condition tracks whether the cluster is in service (not in maintenance mode). The \"NodeCountAsExpected\" condition tracks whether the expected number of nodes are present. Each of these conditions is required, so the array must contain at least 3 items.", + "description": "conditions represent the observations of the pacemaker cluster's current state. Known condition types are: \"Healthy\", \"InService\", \"NodeCountAsExpected\". The \"Healthy\" condition is an aggregate that tracks the overall health of the cluster. The \"InService\" condition tracks whether the cluster is in service (not in maintenance mode). The \"NodeCountAsExpected\" condition tracks whether the expected number of nodes are present. Each of these three conditions is required, so the array must contain at least 3 items.", "type": "array", "items": { "default": {},