RWS Architecture article

ADR-022 Transitional production exception for OpenShift AI in DC-A

The program needs production-visible progress early, while the broader DC 3.0 target architecture is still being engineered and validated. A dedicated OpenShift AI cluster in DC-A

  1. Typeadr
  2. Statusdraft
  3. Domainplatform
On this page
  1. Context and Problem Statement
  2. Decision Drivers
  3. Considered Options
  4. Decision Outcome
  5. Guardrails
  6. Transitional dependency selections (explicit)
  7. Transitional service boundary (explicit)
  8. Transitional model-governance posture (explicit)
  9. Transitional backup and restore posture (explicit)
  10. Control ownership mapping (explicit)
  11. Observability and expansion posture (explicit)
  12. Exit Conditions

Context and Problem Statement

The program needs production-visible progress early, while the broader DC 3.0 target architecture is still being engineered and validated. A dedicated OpenShift AI cluster in DC-A (AM1) can deliver near-term value by reusing selected IST dependencies, but this deviates from the intended SOLL operating model in DC-B (AM4).

Without an explicit exception decision, the transitional lane can be misread as a second permanent baseline, which increases architectural drift risk and governance ambiguity.

Decision Drivers

  • Deliver production value quickly to support program confidence and continuity.
  • Keep the DC-A lane tightly scoped to AI acceleration and avoid broad platform sprawl.
  • Preserve clear separation between transitional architecture and DC 3.0 target-state design.
  • Make caveats, risks, and retirement conditions explicit and auditable.
  • Maintain traceability from architecture to delivery and governance controls.

Considered Options

  1. Delay production until DC 3.0 target architecture is ready.
  2. Launch OpenShift AI in DC-A as an unrestricted production platform.
  3. Launch OpenShift AI in DC-A as a controlled transitional production exception.

Decision Outcome

Chosen option: Option 3 - controlled transitional production exception.

The DC-A OpenShift AI cluster is approved as a temporary production lane with strict scope boundaries, explicit caveats, and retirement criteria. This lane is not a replacement for the DC 3.0 target architecture.

Guardrails

  • Single cluster in DC-A only.
  • Standalone cluster lifecycle with bootstrap installation (no ACM-managed deployment path in this lane).
  • Bootstrap host remains persistent for this lane (jump host plus approved automation runner).
  • Scope limited to AI acceleration services and required platform dependencies.
  • No general tenant expansion or broad platform service onboarding.
  • Full BIO2 compliance and formal BIO2 certification are not target outcomes for this transitional lane.
  • Dependency reuse from IST must be documented with owner and retirement trigger.
  • Transitional posture remains under periodic architecture governance review.
  • The full DC-A transitional production control catalog is mandatory for production onboarding and ongoing operation.

Transitional dependency selections (explicit)

  • IAM: use existing enterprise AD connectivity via OIDC through ADFS for this lane; authorization is based on a curated AD group allowlist with two role personas (platform-admin with full cluster-admin rights in this lane and service-consumer as consume-only); outage mode is fail-closed with audited break-glass access only (two local emergency admin accounts in primary/backup model, two-person approval, and post-incident credential rotation); Keycloak is not part of the DC-A transitional implementation.
  • Registry: use existing Harbor pull-through cache as the primary image pull path for most well-known registries in this lane; direct access exception is explicitly allowed for Red Hat registries; additional direct external pulls are operationally controlled and no formal exception record is mandatory in this phase; shared cluster-wide pull secret is accepted for transitional operations; no hard scanning/signing gate is enforced at day-1 onboarding; Quay is not part of the DC-A transitional implementation.
  • Networking: consume IST-provided ACI networking as-is for this lane; no platform-side network automation is delivered in this phase.
  • Storage: use a production Ceph cluster in AM1, aligned with Storage Architecture, with deployment automation executed from the persistent bootstrap host and integrated through ODF external mode.

These selections are lane-specific exceptions and do not change the long-term DC 3.0 target direction.

Transitional service boundary (explicit)

  • Service model is single-tier production for this lane.
  • Day-1 scope includes Notebooks-as-a-Service (NaaS) and Models-as-a-Service (MaaS).
  • Platform usage also includes LLM-backend support for OpenShift Lightspeed for this AM1 lane and for DC 3.0 development activities in AM4.
  • Managed model training and broad multi-tenant self-service are out of scope in this transitional phase.
  • Standard support window is business hours with formal escalation path for major incidents.
  • Service SLO posture is business-hours response targets.
  • New team onboarding requires technical readiness and architecture sign-off (no formal approval response SLA in this phase).

Transitional model-governance posture (explicit)

  • LLMeval and LMjudge are treated as stretch-goal capabilities in this phase.
  • If implemented, they run on each release candidate and on a weekly scheduled regression cycle.
  • If implemented, they remain advisory (non-blocking).
  • No numeric threshold bands are fixed yet; outputs are used for qualitative release-risk assessment.
  • Formal acceptance records for poor outcomes are optional (recommended for significant regressions).

Transitional backup and restore posture (explicit)

  • Mandatory backup scope includes AI namespaces and platform configuration.
  • Restore drills run monthly and target original namespaces.
  • Minimum drill depth is metadata plus representative sample PVC data restore.
  • Drill failures require remediation on a case-by-case timeline and do not block releases in this phase.
  • Restore-drill evidence is retained for 12 months.

Control ownership mapping (explicit)

  • Platform Ops owns access and identity, image and supply chain, network and egress, data and storage, observability and operations, and platform change controls.
  • Backup Ops owns backup and recovery controls.
  • AI Platform Team owns model-governance controls.
  • Dedicated AI operations ownership remains pending in this phase; interim handoff boundaries are tracked in Workpackage 21 runbooks.

Observability and expansion posture (explicit)

  • Operations paging is triggered by service-impact events only.
  • Mandatory paging channels for service-impact events are ticket plus chat.
  • Expansion beyond pilot users requires core control automation to be in place.

Exit Conditions

  • Primary trigger is DC 3.0 target-state in DC-B proven in production and approved through readiness gates.
  • AI services migrate to DC-B after production-proof approval, if migration is required.
  • Migration runs as a joint effort by platform and customer teams in planned service windows with best-effort downtime posture.
  • DC-A transitional lane is retired after service migration completion.
  • DC-A is rebuilt as DC 3.0 production-ready site.
  • Transitional dependency exceptions are resolved or explicitly replaced by target-state integrations.

Positive Consequences

  • Enables near-term production delivery while larger platform work continues.
  • Creates explicit governance boundaries that reduce accidental scope creep.
  • Improves traceability of temporary decisions and their retirement path.

Negative Consequences

  • Additional operational complexity due to transitional dependency mix.
  • Requires active governance effort to prevent temporary patterns from becoming permanent.
  • May provide lower standardization than the target-state operating model.
  • Absence of BIO2 certification in this lane can limit use cases that require formal certification.
  • Pending dedicated AI operations ownership can create handoff friction if not resolved early.

Pros and Cons of the Options

Option 1 - Delay production until DC 3.0 target architecture is ready

  • Good: Strong architectural purity and lower transitional complexity.
  • Bad: Delays business-visible outcomes and learning cycles.

Option 2 - Unrestricted production platform in DC-A

  • Good: Fastest path to broad service enablement in one location.
  • Bad: High risk of architectural drift and duplicate baseline creation.

Option 3 - Controlled transitional production exception (chosen)

  • Good: Balances delivery urgency with explicit governance controls.
  • Bad: Still requires temporary complexity management and disciplined retirement.