On this page
- Context and Problem Statement
- Decision Drivers
- Considered Options
- Decision Outcome
- Guardrails
- Transitional dependency selections (explicit)
- Transitional service boundary (explicit)
- Transitional model-governance posture (explicit)
- Transitional backup and restore posture (explicit)
- Control ownership mapping (explicit)
- Observability and expansion posture (explicit)
- Exit Conditions
Context and Problem Statement
The program needs production-visible progress early, while the broader DC 3.0 target architecture is still being engineered and validated. A dedicated OpenShift AI cluster in DC-A (AM1) can deliver near-term value by reusing selected IST dependencies, but this deviates from the intended SOLL operating model in DC-B (AM4).
Without an explicit exception decision, the transitional lane can be misread as a second permanent baseline, which increases architectural drift risk and governance ambiguity.
Decision Drivers
- Deliver production value quickly to support program confidence and continuity.
- Keep the DC-A lane tightly scoped to AI acceleration and avoid broad platform sprawl.
- Preserve clear separation between transitional architecture and DC 3.0 target-state design.
- Make caveats, risks, and retirement conditions explicit and auditable.
- Maintain traceability from architecture to delivery and governance controls.
Considered Options
- Delay production until DC 3.0 target architecture is ready.
- Launch OpenShift AI in DC-A as an unrestricted production platform.
- Launch OpenShift AI in DC-A as a controlled transitional production exception.
Decision Outcome
Chosen option: Option 3 - controlled transitional production exception.
The DC-A OpenShift AI cluster is approved as a temporary production lane with strict scope boundaries, explicit caveats, and retirement criteria. This lane is not a replacement for the DC 3.0 target architecture.
Guardrails
- Single cluster in DC-A only.
- Standalone cluster lifecycle with bootstrap installation (no ACM-managed deployment path in this lane).
- Bootstrap host remains persistent for this lane (jump host plus approved automation runner).
- Scope limited to AI acceleration services and required platform dependencies.
- No general tenant expansion or broad platform service onboarding.
- Full BIO2 compliance and formal BIO2 certification are not target outcomes for this transitional lane.
- Dependency reuse from IST must be documented with owner and retirement trigger.
- Transitional posture remains under periodic architecture governance review.
- The full DC-A transitional production control catalog is mandatory for production onboarding and ongoing operation.
Transitional dependency selections (explicit)
- IAM: use existing enterprise AD connectivity via OIDC through ADFS for this lane; authorization is based on a curated AD group allowlist with two role personas (
platform-adminwith full cluster-admin rights in this lane andservice-consumeras consume-only); outage mode is fail-closed with audited break-glass access only (two local emergency admin accounts in primary/backup model, two-person approval, and post-incident credential rotation); Keycloak is not part of the DC-A transitional implementation. - Registry: use existing Harbor pull-through cache as the primary image pull path for most well-known registries in this lane; direct access exception is explicitly allowed for Red Hat registries; additional direct external pulls are operationally controlled and no formal exception record is mandatory in this phase; shared cluster-wide pull secret is accepted for transitional operations; no hard scanning/signing gate is enforced at day-1 onboarding; Quay is not part of the DC-A transitional implementation.
- Networking: consume IST-provided ACI networking as-is for this lane; no platform-side network automation is delivered in this phase.
- Storage: use a production Ceph cluster in AM1, aligned with Storage Architecture, with deployment automation executed from the persistent bootstrap host and integrated through ODF external mode.
These selections are lane-specific exceptions and do not change the long-term DC 3.0 target direction.
Transitional service boundary (explicit)
- Service model is single-tier
productionfor this lane. - Day-1 scope includes Notebooks-as-a-Service (NaaS) and Models-as-a-Service (MaaS).
- Platform usage also includes LLM-backend support for OpenShift Lightspeed for this AM1 lane and for DC 3.0 development activities in AM4.
- Managed model training and broad multi-tenant self-service are out of scope in this transitional phase.
- Standard support window is business hours with formal escalation path for major incidents.
- Service SLO posture is business-hours response targets.
- New team onboarding requires technical readiness and architecture sign-off (no formal approval response SLA in this phase).
Transitional model-governance posture (explicit)
- LLMeval and LMjudge are treated as stretch-goal capabilities in this phase.
- If implemented, they run on each release candidate and on a weekly scheduled regression cycle.
- If implemented, they remain advisory (non-blocking).
- No numeric threshold bands are fixed yet; outputs are used for qualitative release-risk assessment.
- Formal acceptance records for poor outcomes are optional (recommended for significant regressions).
Transitional backup and restore posture (explicit)
- Mandatory backup scope includes AI namespaces and platform configuration.
- Restore drills run monthly and target original namespaces.
- Minimum drill depth is metadata plus representative sample PVC data restore.
- Drill failures require remediation on a case-by-case timeline and do not block releases in this phase.
- Restore-drill evidence is retained for 12 months.
Control ownership mapping (explicit)
- Platform Ops owns access and identity, image and supply chain, network and egress, data and storage, observability and operations, and platform change controls.
- Backup Ops owns backup and recovery controls.
- AI Platform Team owns model-governance controls.
- Dedicated AI operations ownership remains pending in this phase; interim handoff boundaries are tracked in Workpackage 21 runbooks.
Observability and expansion posture (explicit)
- Operations paging is triggered by service-impact events only.
- Mandatory paging channels for service-impact events are ticket plus chat.
- Expansion beyond pilot users requires core control automation to be in place.
Exit Conditions
- Primary trigger is DC 3.0 target-state in DC-B proven in production and approved through readiness gates.
- AI services migrate to DC-B after production-proof approval, if migration is required.
- Migration runs as a joint effort by platform and customer teams in planned service windows with best-effort downtime posture.
- DC-A transitional lane is retired after service migration completion.
- DC-A is rebuilt as DC 3.0 production-ready site.
- Transitional dependency exceptions are resolved or explicitly replaced by target-state integrations.
Positive Consequences
- Enables near-term production delivery while larger platform work continues.
- Creates explicit governance boundaries that reduce accidental scope creep.
- Improves traceability of temporary decisions and their retirement path.
Negative Consequences
- Additional operational complexity due to transitional dependency mix.
- Requires active governance effort to prevent temporary patterns from becoming permanent.
- May provide lower standardization than the target-state operating model.
- Absence of BIO2 certification in this lane can limit use cases that require formal certification.
- Pending dedicated AI operations ownership can create handoff friction if not resolved early.
Pros and Cons of the Options
Option 1 - Delay production until DC 3.0 target architecture is ready
- Good: Strong architectural purity and lower transitional complexity.
- Bad: Delays business-visible outcomes and learning cycles.
Option 2 - Unrestricted production platform in DC-A
- Good: Fastest path to broad service enablement in one location.
- Bad: High risk of architectural drift and duplicate baseline creation.
Option 3 - Controlled transitional production exception (chosen)
- Good: Balances delivery urgency with explicit governance controls.
- Bad: Still requires temporary complexity management and disciplined retirement.
Links and References
- Architecture context: openshift-ai-cluster-setup
- Delivery stream: wp-21-openshift-ai-dc-a-transitional-production-lane
- Rollout context: adr-004-single-active-phased-rollout, adr-019-hybrid-network-rollout-aci-a-ai-only-evpn-b-target
- Connectivity context: adr-011-connected-vs-disconnected