RWS Architecture article

OpenShift AI DC-A Transitional Production Lane

Deliver a standalone, production-ready OpenShift AI implementation in DC-A (AM1) as a transitional exception lane, with explicit caveats, operational controls, and a governed trans

  1. Typeworkpackage
  2. Statusplanned
  3. Domainplatform
On this page
  1. WP-21 OpenShift AI DC-A Transitional Production Lane
  2. Summary
  3. Scope
  4. Handover Boundary
  5. Implementation assumptions and constraints
  6. Engineering Clarifications
  7. Out of Scope
  8. Architecture Context
  9. Decision Context
  10. Dependencies
  11. Technical Operating Profiles and Baseline
  12. Engineering Handover Package

WP-21 OpenShift AI DC-A Transitional Production Lane

Summary

Deliver a standalone, production-ready OpenShift AI implementation in DC-A (AM1) as a transitional exception lane, with explicit caveats, operational controls, and a governed transition path to DC-B (AM4) target-state operations when needed.

Scope

  • Build and operate one OpenShift AI cluster in DC-A with a single production service tier.
  • Implement this as a standalone cluster (not managed or deployed by ACM).
  • Install the cluster through the bootstrap installation process, with a persistent bootstrap host used as jump host and automation runner.
  • Deploy the production Ceph cluster through approved automation run from the bootstrap host, and validate readiness before ODF external integration.
  • Deliver NaaS (Notebooks-as-a-Service) and MaaS (Models-as-a-Service) services for approved teams under transitional controls.
  • Enable LLM-backend usage for OpenShift Lightspeed for this AM1 lane and DC 3.0 development activities in AM4.
  • Implement the mandatory production control catalog and produce audit-ready evidence.
  • Provide all engineering handover content directly in this workpackage as the implementation contract.
  • Prepare and execute transition readiness for migration to DC-B once DC 3.0 production gates are met, if migration is required.

Handover Boundary

  • This workpackage is the engineering handoff contract for implementation scope, sequencing, controls, and acceptance.
  • Detailed LLD-specific installation inputs (exact DNS records, CIDRs, domain values, proxy/no-proxy lists, and host-level bootstrap artifacts) are owned by engineering implementation design and are not duplicated in architecture notes.

Implementation assumptions and constraints

  • AM1 is the delivery location for this transitional lane; AM4 is the target location for DC 3.0 production readiness and potential transition.
  • Networking uses IST-provided ACI as-is for this lane, with no platform-side network automation scope.
  • Full BIO2 compliance and formal BIO2 certification are not target outcomes for this transitional lane.
  • Dedicated AI operations team assignment is pending; interim operational ownership and handoff model must be defined during delivery.
  • The bootstrap host is persistent for this lane and is used for jump access and approved automation execution.

Engineering Clarifications

  • Finalize exact installation inputs and network values in engineering LLD before starting Phase 2.
  • Finalize bootstrap host operating baseline (access model, hardening posture, and toolchain ownership) before installation execution.
  • Treat Ceph automation outputs as a hard readiness input for platform storage integration.
  • Treat AD group naming and mapping details as a joint implementation task with IAM team during Phase 3.
  • Keep registry behavior explicit in runbooks: Harbor pull-through cache default path, Red Hat direct exception path, and operational control of additional direct pulls.
  • Use role-based ownership for control evidence and phase exits.
  • Pin stable versions at install time and record them in Package A for reproducibility.
  • Use fix-forward operations between phases; do not assume rollback as default execution path.

Out of Scope

  • Managed model training service commitments.
  • Broad multi-tenant self-service onboarding.
  • Alignment of this lane to full DC 3.0 target-state parity before initial production use.
  • BIO2 certification delivery in this transitional lane.

Architecture Context

Decision Context

Dependencies

  • No dependency on other workpackages.
  • Delivery uses external service dependencies defined in architecture and ADR documents (AD/ADFS, Harbor, Infoblox, NetScaler, Cohesity, Ceph).

Technical Operating Profiles and Baseline

AreaImplementation profile for engineering
Cluster topologyOne OpenShift AI cluster in DC-A only; single service tier production.
Location modelDelivery location is AM1 (DC-A); transition target location is AM4 (DC-B).
Cluster lifecycle modelStandalone cluster lifecycle; no ACM-based deployment or lifecycle management for this lane.
Installation methodBootstrap installation flow (assisted by approved bootstrap automation where applicable).
Bootstrap host modelBootstrap host remains persistent and serves as jump host and approved automation runner.
Ceph delivery modelProduction Ceph deployment is executed via approved automation from bootstrap and validated before ODF external integration.
Hardware profileBare-metal control plane with 3 control-plane nodes and 6 GPU worker nodes; use the standard platform hardware profile from OpenShift Platform Hardware Specs.
In-scope servicesNaaS (Notebooks-as-a-Service) and MaaS (Models-as-a-Service) enabled for approved teams.
OpenShift Lightspeed backendLLM-backend usage is supported for this AM1 lane and DC 3.0 development workflows in AM4.
IAMOIDC via ADFS, AD allowlist for platform-admin and service-consumer; platform-admin has cluster-admin rights; service-consumer has consume-only access.
Break-glassTwo local emergency admin accounts (primary and backup), two-person approval, rotate credentials after incident closure.
RegistryHarbor pull-through cache is the primary path for most well-known registries; direct path is limited to approved Red Hat registries; any additional direct pull is operationally controlled by Platform Ops runtime policy.
StorageProduction Ceph backend aligned to DC 3.0 storage direction via ODF external integration for RBD/CephFS and object where required.
Network automation boundaryACI networking is consumed from IST as-is; no platform-side network automation is delivered in this lane.
Backup/restoreBackup scope includes AI namespaces and platform configuration; monthly restore drills in original namespace with metadata plus sample PVC data; evidence retained 12 months.
Model governanceLLMeval/LMjudge is a stretch goal for this lane; when enabled it is advisory and non-blocking.
ObservabilityService-impact paging only; mandatory paging channels are ticket and chat.
Expansion gateExpansion beyond pilot users requires core control automation outcomes to be in place.
Version policyUse stable, explicitly pinned versions for OpenShift, operators, and dependencies at install time.
Compliance boundaryBIO2 certification is not a delivery objective for this transitional lane.

This baseline aligns with the operating profiles in openshift-ai-cluster-setup and serves as the execution profile for implementation and handover.

Engineering Handover Package

Required engineering inputs (SME-owned)

  • Final AM1 installation inputs: DNS/LB records, CIDRs, domain values, egress/proxy/no-proxy lists, and host access details.
  • Bootstrap host build inputs: OS baseline, hardening controls, identity model, and approved automation toolchain.
  • Ceph deployment inputs: node inventory, network mappings, storage profiles, and readiness criteria for ODF external integration.
  • IAM and access inputs: ADFS endpoint details, AD group mappings, and break-glass operational approvals.
  • Version selection inputs: pinned OpenShift, operator, and dependency versions with compatibility rationale.

Required engineering outputs (SME-owned)

  • LLD deliverable set with final implementation decisions, sequence, and rollback/fix-forward posture.
  • Executable deployment runbook and validation checkpoints for bootstrap host, Ceph, OpenShift, and OpenShift AI bring-up.
  • Operations runbook with ownership model, incident handling, backup/restore drills, and transition decision gates.
  • Evidence bundle mapped to phase gates and mandatory control families.

Package A - Implementation specification

  • Capture final cluster, service, IAM, registry, storage, and observability configuration values used for deployment.
  • Capture approved namespace model, access mappings, and onboarding entry criteria.

Package B - Deployment runbook

  • Define ordered bootstrap and configuration sequence from cluster baseline to service enablement.
  • Define validation checkpoints after each major setup stage.

Package C - Operations runbook

  • Define operational procedures for incident handling, break-glass access, routine change, backup, and restore.
  • Define escalation flow and service-impact handling using ticket and chat.

Package D - Validation and evidence profile

  • Define mandatory verification checks per control family.
  • Define evidence artifacts and retention expectations for governance and audit.

Phase Gate Matrix

PhaseEntry criteriaExit criteriaEvidenceOwner (role)
Phase 1 - PreparationADR and architecture baseline approved; external dependency owners identifiedAll prerequisites confirmed, hardware readiness validated, bootstrap host baseline approved, install prerequisites signed offPrerequisite checklist, dependency confirmations, hardware readiness record, bootstrap host baseline recordPlatform Ops
Phase 2 - Cluster installationPhase 1 exit signedCeph automation executed from bootstrap with readiness validated; standalone OpenShift installation complete; control-plane quorum and worker registration healthyCeph automation report, Ceph readiness checks, install log, cluster health checks, bootstrap completion reportPlatform Ops
Phase 3 - Initial cluster configurationPhase 2 exit signedAuth/proxy/registry baseline configured; break-glass controls validatedConfig manifests, auth validation, break-glass validation recordPlatform Ops
Phase 4 - Platform capability deploymentPhase 3 exit signedStorage/GPU/operator prerequisites installed and validated for OAI; ODF external integration validated against production CephOperator status report, GPU readiness checks, ODF-Ceph integration validationPlatform Ops
Phase 5 - OpenShift AI finalizationPhase 4 exit signedOAI services configured (NaaS/MaaS), go-live controls validated, onboarding gate approvedOAI validation report, control evidence bundle, onboarding sign-offPlatform Ops + Architecture

Implementation Workstreams

Phase 1 - Preparation

  • Confirm prerequisite external dependencies are available and approved for use (AD/ADFS, Harbor, Infoblox, NetScaler, Cohesity, Ceph).
  • Confirm hardware readiness for 3 control-plane and 6 GPU worker bare-metal nodes.
  • Confirm network lane prerequisites for DC-A AI-only scope.
  • Confirm installation inputs, access, and bootstrap artifacts are complete.
  • Confirm persistent bootstrap host baseline and approved automation toolchain.

Phase 2 - Cluster installation

  • Build and harden the persistent bootstrap host and enable approved automation runtime.
  • Execute Ceph deployment automation from bootstrap and validate Ceph readiness outputs.
  • Install standalone OpenShift cluster via bootstrap process.
  • Validate control-plane quorum, worker registration, and baseline cluster health.

Phase 3 - Initial cluster configuration

  • Configure initial cluster settings, including authentication, proxy, registry access, and baseline operational controls.
  • Implement OIDC via ADFS and curated AD allowlist role mappings.
  • Finalize AD group names and mapping details in coordination with IAM team.
  • Implement break-glass primary and backup controls.
  • Apply lane network controls under the IST-provided ACI model (no platform-side network automation).

Phase 4 - Platform capability deployment

  • Deploy required OpenShift platform prerequisites for AI workloads (storage integration, GPU runtime support, and required operators).
  • Integrate ODF external mode with the production Ceph cluster validated in Phase 2.
  • Validate cluster readiness for OpenShift AI deployment.

Phase 5 - OpenShift AI finalization

  • Deploy OpenShift AI and configure NaaS and MaaS services in the single production tier.
  • Configure and validate LLM-backend exposure pattern for OpenShift Lightspeed use in AM1 and DC 3.0 development workflows in AM4.
  • If implemented, enable advisory LLMeval/LMjudge execution cadence for release candidates and weekly regressions (stretch goal).
  • Finalize production control evidence, restore drill readiness, and transition-readiness documentation.
  • Produce and approve Package A through D handover content during each phase.

Cross-phase operating rule

  • Execution uses a fix-forward model; rollback is not the default strategy.

Delivery Sequence

  1. Complete Phase 1 preparation and prerequisite sign-off.
  2. Execute Phase 2 bootstrap-host baseline, Ceph automation, and OpenShift installation.
  3. Execute Phase 3 initial cluster configuration.
  4. Execute Phase 4 platform capability deployment.
  5. Execute Phase 5 OpenShift AI finalization and production-readiness evidence.
  6. Start production onboarding for approved teams.
  7. Execute migration and rebuild sequence when DC-B readiness trigger is approved and migration is required.

Acceptance Criteria

  • OpenShift AI DC-A lane is in production with NaaS and MaaS available for approved teams.
  • Cluster is installed as a standalone cluster via bootstrap process, without ACM-managed deployment.
  • Persistent bootstrap host role is operational (jump host plus approved automation runner).
  • Production Ceph deployment from bootstrap automation is validated and integrated through ODF external mode.
  • Hardware baseline is implemented and verified: 3 control-plane bare-metal nodes and 6 GPU worker nodes.
  • Mandatory production control catalog is implemented with assigned owners and evidence artifacts.
  • OpenShift Lightspeed LLM-backend usage path is implemented for AM1 lane use and AM4 development support.
  • LLMeval/LMjudge stretch goal status is explicitly documented (implemented or deferred) with rationale.
  • Backup and monthly restore drill process is operational with retained evidence.
  • Engineering handover package sections in this workpackage are complete, reviewed, and implementation-usable by engineering teams.
  • Transition trigger, migration runbook (joint platform and customer execution, if needed), and DC-A rebuild initiation criteria are documented and approved.
  • Initial production onboarding requires both technical readiness and architecture sign-off.

Restore Drill Success Criteria

  • Restore completes in the original namespace for selected in-scope targets.
  • Restored resources are healthy and reachable for the defined validation path.
  • Sample PVC data integrity check passes for selected workloads.
  • Drill evidence and findings are recorded and retained according to policy.

Required Evidence

  • Deployment checkpoint evidence for each workstream.
  • IAM, registry, and break-glass control evidence.
  • Backup runs and monthly restore drill evidence with retention proof.
  • LLMeval/LMjudge execution evidence where stretch goal is implemented, or explicit defer rationale where not implemented.
  • Onboarding approvals for in-scope teams.
  • Migration readiness evidence and approved transition decision.

Risks and Caveats

  • Transitional dependency reuse introduces operational variance versus target-state controls.
  • IST-provided ACI with no platform-side network automation increases dependency on external operational lead times.
  • Advisory model-governance posture increases reliance on operational judgment for release risk.
  • Shared pull-secret model and non-blocking image controls are transitional and require strict operational discipline.
  • Dedicated AI operations ownership is not final and can create handoff friction if left unresolved.
  • No BIO2 certification target in this lane may constrain use cases that require formal certification.

Transition and Exit

  • Transition begins only when DC-B is proven in production and readiness gates are approved.
  • AI services migrate from DC-A to DC-B in planned windows with best-effort downtime posture, if migration is required.
  • Migration execution is a joint effort by platform and customer teams.
  • DC-A transitional lane is retired after migration completion.
  • DC-A rebuild to DC 3.0 production-ready state starts after migration closure.

Future Work

  • Keep this workpackage synchronized with architecture note updates for operating boundaries, caveats, and migration conditions.
  • Refine evidence templates and ownership mappings as the dedicated AI operations function is formalized.
  • Add post-go-live optimization backlog items for operations efficiency and control automation maturity.

Open Questions

  • Which team takes final steady-state ownership for OpenShift Lightspeed LLM-backend operations?
  • Which readiness threshold defines when migration to AM4 is required versus continued AM1 operation?
  • Which minimum additional controls are required before scaling from pilot onboarding to broader production onboarding?