RWS Architecture article

ADR-004 Single-active phased datacenter rollout with Datacenter Services clusters

The platform will launch across two physical locations but will not operate in an active-active posture initially. We need a rollout strategy that keeps production stable, accelera

  1. Typeadr
  2. Statusaccepted
  3. Domainplatform
On this page
  1. Context and Problem Statement
  2. Decision Drivers
  3. Considered Options
  4. Decision Outcome
  5. Rollout Model (Implemented Baseline)
  6. Readiness and Cutover Controls (Normative)
  7. Cluster Naming and Governance
  8. Positive Consequences
  9. Negative Consequences
  10. Pros and Cons of the Options
  11. Option 1 – Single-active, phased rollout (chosen)
  12. Option 2 – Active-active from day one

Context and Problem Statement

The platform will launch across two physical locations but will not operate in an active-active posture initially. We need a rollout strategy that keeps production stable, accelerates DevOps automation with OpenShift AI coding agents, and clarifies the cluster role responsible for shared dependencies per site. The rollout intent is captured in topic-datacenter-rollout-intent.

Decision Drivers

  • Minimize operational risk by avoiding active-active complexity during early rollout.
  • Use OpenShift AI coding agents to accelerate automation and validation of the target platform.
  • Keep Location A network scope tightly bounded to the AI lane while Location B matures the target-state EVPN/VXLAN platform.
  • Maintain explicit, location-scoped ownership of shared services while preserving centralized governance in Advanced Cluster Management (ACM).
  • Provide a clear cutover path once the dynamic platform reaches parity.

Considered Options

  1. Single-active, phased rollout with Location A as the initial baseline and Location B as the dynamic build-out target.
  2. Active-active from day one across both locations.
  3. Single-location launch with later expansion, without staged cutover planning.

Decision Outcome

Chosen option: Single-active, phased rollout with a controlled cutover to Location B.

Rollout Model (Implemented Baseline)

  • Location A (static intent baseline): Keep a minimal-change baseline for initial production continuity; the Location A ACI lane remains strictly AI-cluster-only and is not a general tenant platform lane.
  • Location B (dynamic target): Build and validate the full dynamic platform (EVPN/VXLAN, multi-cluster automation, Datacenter Services internal/external profiles, tiered Red Hat Quay) prior to production cutover.
  • Cutover and fallback: Execute cutover only through the approved change procedure with rollback controls, keeping Location A as bounded fallback during transition.

Readiness and Cutover Controls (Normative)

  • Cutover decision requires measurable release evidence approved by named owners from network, platform, security, and BC/DR.
  • Rollback objectives for critical services are validated in delivery exercises before cutover approval.
  • Location A AI-lane boundary controls are enforced through topic-network-aci-location-a-ai-fasttrack and openshift-ai-cluster-setup.

Cluster Naming and Governance

The intent for this rollout model and naming change is captured in the reference record (see topic-datacenter-rollout-intent).

Positive Consequences

  • Reduced initial operational complexity compared to active-active deployments.
  • Clear sequencing for platform build-out and cutover activities.
  • Explicit location-scoped separation of shared services with centralized governance.

Negative Consequences

  • Requires cutover planning and validation criteria before switching locations.
  • Temporarily maintains two environments with different maturity levels.

Pros and Cons of the Options

Option 1 – Single-active, phased rollout (chosen)

  • Good: Reduces early operational risk and coordination overhead.
  • Good: Allows AI-enabled automation to mature before cutover.
  • Bad: Requires a defined cutover window and readiness gates.

Option 2 – Active-active from day one

  • Good: Immediate dual-site resiliency.
  • Bad: Higher complexity and policy drift risk early on.

Option 3 – Single-location launch, later expansion

  • Good: Simplifies initial rollout.
  • Bad: Defers the Location B build-out and delays cutover readiness.