On this page
Context and Problem Statement
The long-term design assumes Advanced Cluster Management (ACM) for cluster lifecycle management, but initial hardware constraints may limit how many hubs and service clusters we can deploy in 1H2026. We need to decide whether ACM is mandatory for initial deployment or can be phased in.
Decision Drivers
- Alignment with the target single Advanced Cluster Management (ACM) hub operating model with one Red Hat Advanced Cluster Management installation per hub cluster (see 10-topics/platform/topic-phase-1-investigation-backlog-investigation).
- Hardware capacity and timeline constraints for 1H2026 delivery.
- Desire for automated provisioning and Automation-driven day-2 operations.
Considered Options
- Build an Advanced Cluster Management (ACM) hub (with one Red Hat Advanced Cluster Management installation on that hub cluster) and use ACM to deploy storage and workload clusters.
- Deploy storage and workload clusters with Ansible/Automation first, then import into Advanced Cluster Management (ACM) when capacity is available.
Decision Outcome
Chosen option: Build an Advanced Cluster Management (ACM) hub (with one Red Hat Advanced Cluster Management installation on that hub cluster) and use ACM as the production governance plane.
Implementation baseline:
- Production-bound clusters must be onboarded into ACM before they are considered rollout-ready.
- Temporary bootstrap automation outside ACM is allowed only for non-production bootstrap windows and must converge back into ACM before cutover gates are approved.
- Fleet policy compliance, placement, and lifecycle workflows are measured from ACM artifacts as rollout evidence.
- Exception scope: only the standalone Location A transitional OpenShift AI lane is out of ACM lifecycle scope; no additional production lanes are covered by this exception.
- Exception owner: Platform Architecture (governance owner) and Platform Operations (execution owner) jointly maintain the exception record.
- Exception evidence: each release bundle must include lane-boundary evidence, control-compliance evidence, and an explicit statement that Location A lane capacity is excluded from Location B cutover-readiness claims.
- Exception expiry: the exception record must carry an explicit expiry date and renewal decision, and it is automatically retired when Location B cutover is approved and executed.
Pros and Cons of the Options
Option 1 – Advanced Cluster Management-first deployment (chosen)
Pros
- Aligns immediately with the long-term operating model.
- Consistent governance and policy enforcement from day one.
Cons
- Requires more initial hardware and operational setup.
Option 2 – Phased Advanced Cluster Management adoption
Pros
- Enables early delivery with limited hardware.
- Allows Automation/Ansible to unblock initial clusters.
Cons
- Requires later migration and onboarding work into Advanced Cluster Management (ACM).