RWS Architecture article

OpenShift AI Cluster Setup (DC-A Transitional Production Lane)

This note defines the transitional production architecture for a single OpenShift AI cluster in DC-A (AM1), created to deliver production value quickly by reusing selected IST

  1. Typearchitecture
  2. Statusdraft
  3. Domainplatform
On this page
  1. Summary
  2. Applicable Principles
  3. Architecture
  4. Decision snapshot
  5. Physical architecture
  6. Logical architecture
  7. View navigation
  8. Intent and positioning
  9. Scope lock (hard boundaries)
  10. Cluster role and service intent
  11. Dependency and integration baseline (DC-A transitional)
  12. In-cluster implementation stack (initial posture)

Summary

hero-right

This note defines the transitional production architecture for a single OpenShift AI cluster in DC-A (AM1), created to deliver production value quickly by reusing selected IST (legacy) services. It is intentionally outside the long-term DC 3.0 SOLL target and is treated as a time-boxed exception with explicit caveats, controls, and exit criteria.

The lane runs as a standalone bootstrap-installed cluster (no ACM lifecycle management) with a fixed hardware profile of three bare-metal control-plane nodes and six GPU worker nodes. Engineering delivery is anchored in Workpackage 21, while governance and exception boundaries are anchored in ADR 022. It uses a production Ceph cluster in AM1, created in line with the DC 3.0 Storage Architecture.

AspectDecision
Lane typeTransitional production exception in DC-A
Delivery anchorWorkpackage 21
Governance anchorADR 022
Service modelSingle tier production with NaaS and MaaS
IAMAD/ADFS OIDC, curated allowlist, break-glass controls
RegistryHarbor pull-through cache primary, Red Hat direct exception
Model governanceLLMeval/LMjudge stretch goal; if enabled, advisory and non-blocking
Storage baselineProduction Ceph in AM1 aligned to DC 3.0 Storage Architecture
Exit triggerDC-B (AM4) proven production under DC 3.0 readiness gates

Applicable Principles

Architecture

Decision snapshot

  • Why this lane exists: deliver production value quickly while DC 3.0 target-state is still being built.
  • What is explicitly constrained: single cluster in DC-A, no broad self-service expansion, no target-state precedent.
  • How delivery is governed: mandatory control catalog with named owners and evidence artifacts.
  • How the lane exits: migrate to DC-B after readiness gates; rebuild DC-A as DC 3.0 production site.

Physical architecture

OpenShift AI - Prod - Physical

assets/rendered/openshift/OpenShift-AI-prod-physical.svg
assets/rendered/openshift/OpenShift-AI-prod-physical.svg

Logical architecture

OpenShift AI - Prod - Logical

assets/rendered/openshift/OpenShift-AI-prod-logical.svg
assets/rendered/openshift/OpenShift-AI-prod-logical.svg

View navigation

  • Architecture intent and controls: this note.
  • Implementation baseline for engineering: Workpackage 21.
  • Exception governance decision: ADR 022.

Intent and positioning

  • This cluster exists to accelerate production delivery and management confidence while full DC 3.0 platform capabilities are built in Location B.
  • The lane is deliberately transitional: production use is allowed, but capability growth is constrained to avoid creating a second permanent platform baseline.
  • The long-term target architecture remains openshift-concept-overview and associated ADRs.

Scope lock (hard boundaries)

  • Exactly one OpenShift AI cluster in DC-A.
  • Standalone cluster lifecycle for this lane (no ACM-managed deployment path).
  • No broad multi-tenant platform onboarding in this lane.
  • No expansion into a generic CaaS/VMaaS platform role.
  • Full BIO2 compliance and formal BIO2 certification are not in scope for this transitional lane due to delivery timeline constraints.
  • No policy exception may be treated as precedent for DC 3.0 target-state design.

Cluster role and service intent

  • Deliver internal DevOps acceleration capabilities using coding-agent workflows.
  • Production-facing service set:
  • Notebooks-as-a-Service (Jupyter-based workflows).
  • Models-as-a-Service (LLM inference endpoints).
  • Target additional platform usage as an LLM backend for OpenShift Lightspeed, serving both this DC-A OpenShift AI cluster and DC 3.0 development activities in DC-B (AM4).
  • Evaluation remains a stretch-goal capability in this lane; if implemented, outputs are advisory (non-blocking).

Dependency and integration baseline (DC-A transitional)

Capability areaCurrent DC-A transitional postureLong-term direction
Network laneSingle ACI pod, AI-only scope (adr-019-hybrid-network-rollout-aci-a-ai-only-evpn-b-target)EVPN/VXLAN target lane in Location B
Connectivity profileRestricted network baseline (adr-011-connected-vs-disconnected)Target controls in DC 3.0 platform lanes
External DNS/DHCP/NTPIST dependency via infoblox integration contractsStandardized platform-managed dependency contracts
North-south load balancingExternal ADC dependency via netscaler-blx lane constraintsDatacenter Services baseline per adr-005-datacenter-services-baseline
IAM and authN/authZDirect integration with existing enterprise Active Directory (AD) for this transitional lane; no Keycloak in DC-A laneConsolidated IAM model via adr-012-single-sign-on and adr-015-keycloak-identity-access-management
Registry pathExisting Harbor registry is used in this transitional lane; no Quay in DC-A laneTiered Quay target architecture
StorageExternal Ceph via ODF integration (RBD/CephFS; object path by service need)Same backend pattern with broader DC 3.0 lifecycle controls
Backup/restoreOADP plus Cohesity policy model per topic-openshift-backup-and-restoreSame model with full DR orchestration maturity

In-cluster implementation stack (initial posture)

  • Platform core: OpenShift plus OpenShift AI, Git-governed configuration, and controlled day-2 operations.
  • Installation posture: bootstrap-based installation flow for a standalone cluster.
  • Hardware footprint: bare-metal topology with 3 control-plane nodes and 6 GPU worker nodes.
  • GPU runtime: GPU worker pools with explicit capacity boundaries for inference and evaluation workloads.
  • Storage integration: PVC profiles mapped to workload classes and Ceph-backed storage classes.
  • Security and policy controls: baseline IAM/RBAC, namespace boundaries, and admission policy controls aligned with platform defaults.
  • Observability controls: capture evaluation evidence, serving telemetry, and operational events required for release and audit.

Operating profiles (DC-A transitional)

The sections below define the normative operating baseline for this lane.

Service profile

AspectProfile detailOwnerEvidence
BaselineSingle service tier production with NaaS and MaaSPlatform OpsService readiness checks
Support windowBusiness-hours support for standard operations; major incidents follow escalation processPlatform OpsIncident response records
SLO postureBusiness-hours response SLO applies in this phasePlatform OpsService-level reports
BoundaryNo managed training service and no broad multi-tenant self-service in this phasePlatform OpsScope compliance review
Onboarding gateTechnical readiness and architecture sign-off are both requiredPlatform OpsApproved onboarding records

IAM profile

AspectProfile detailOwnerEvidence
BaselineOIDC via ADFS with curated AD allowlist mapped to OpenShift RBACPlatform OpsOAuth and login validation output
Access modelplatform-admin is cluster-admin; service-consumer is consume-onlyPlatform OpsGroup-to-role mapping records
Outage postureFail closed for normal SSO access; break-glass accounts only during AD/ADFS outagePlatform OpsOutage and access audit trail
Break-glass controlsPrimary/backup local emergency admins; two-person approval; post-incident credential rotationPlatform OpsBreak-glass approvals and rotation logs
BoundaryNo Keycloak capabilities in this lane; AD group naming finalized with IAM team during implementationPlatform OpsIAM alignment notes

Registry profile

AspectProfile detailOwnerEvidence
BaselineHarbor pull-through cache is primary image path; Red Hat registries are direct-access exceptionPlatform OpsImage source inventory and pull tests
BoundaryAdditional direct pulls are operationally controlled; shared pull-secret model is accepted in this phasePlatform OpsRuntime access configuration snapshots
Exception record postureNo formal exception record is mandatory for additional direct pulls in this phasePlatform OpsOperational control evidence
Control postureNo hard image scanning/signing gate at day-1 onboardingPlatform OpsPolicy configuration review

Backup and recovery profile

AspectProfile detailOwnerEvidence
BaselineBackup scope includes AI namespaces and platform configurationBackup OpsBackup configuration and run reports
Drill postureMonthly restore drills run in original namespace with metadata plus sample PVC data validationBackup OpsRestore drill records
Failure policyRemediation after failed drills is case-by-case and does not block releasesBackup OpsRemediation tracking notes
Evidence retentionRestore-drill evidence is retained for 12 monthsBackup OpsRetention audit records

Observability profile

AspectProfile detailOwnerEvidence
BaselineService-impact paging uses ticket and chat channelsPlatform OpsPaging route tests and sample alerts
BoundaryWarning alerts are signal-only by default and do not pagePlatform OpsAlert policy configuration

Model governance profile

AspectProfile detailOwnerEvidence
BaselineLLMeval/LMjudge is a stretch-goal capabilityAI Platform TeamImplementation or defer decision record
If implementedExecution is advisory and non-blocking with no fixed numeric thresholds in this phaseAI Platform TeamGate policy record
Cadence (if implemented)Run on each release candidate and weekly scheduled regressionsAI Platform TeamExecution logs and scorecards
Poor-result handlingFormal acceptance records are optional and recommended for significant regressionsAI Platform TeamAcceptance records when used

Continuous evaluation workflow (LLMeval + LMjudge)

This workflow is a stretch-goal capability in the DC-A transitional lane. When implemented, it remains advisory and non-blocking.

  1. Candidate model is promoted to an evaluation namespace and registered in the model registry.
  2. LLMeval executes coding-agent benchmark suites (pipeline generation, IaC updates, code-review tasks) and records quality and runtime metrics.
  3. LMjudge applies rubric scoring for correctness, safety, and policy compliance.
  4. Metrics and artifacts are retained in approved storage targets with traceable release decisions.
  5. Release workflows consume scorecards to inform promotion decisions and risk discussions.

Production caveats (explicit)

  • This lane is production but not full DC 3.0 target-state parity.
  • Legacy dependency reuse increases operational variance versus target-state automation.
  • Networking for this lane uses IST-provided ACI as-is, with no platform-side network automation in scope.
  • IAM for this lane depends on existing AD connectivity and does not include Keycloak capabilities.
  • Registry supply chain for this lane depends on existing Harbor and does not include Quay controls.
  • The dedicated AI operations team assignment is still pending for this transitional phase.
  • No BIO2 certification is planned for this transitional lane.
  • Some controls are transitional and require additional guardrails until cutover parity is reached.
  • Service catalog breadth is intentionally constrained to prevent uncontrolled scope expansion.

Mandatory production control catalog (DC-A transitional)

All controls below are mandatory before and during production onboarding in this lane.

Control familyMandatory controlEnforcement postureEvidence artifactOwner
Access and identityOIDC via ADFS for cluster authenticationMandatory nowOAuth configuration snapshot and auth test evidencePlatform Ops
Access and identityCurated AD group allowlist to OpenShift RBAC mappingsMandatory nowApproved group-to-role mapping recordPlatform Ops
Access and identityBreak-glass flow uses primary/backup accounts with two-person approval and post-incident rotationMandatory nowBreak-glass runbook, dual-approval records, audit event trail, and rotation evidencePlatform Ops
Image and supply chainImage pulls use Harbor pull-through cache as default pathMandatory nowWorkload image source inventory reportPlatform Ops
Image and supply chainDirect path is limited to Red Hat registries; additional direct external pulls are operationally controlledMandatory nowRuntime access control configuration evidencePlatform Ops
Image and supply chainShared cluster-wide pull secret lifecycle is controlledMandatory nowSecret rotation and access review recordsPlatform Ops
Network and egressRestricted-network egress allowlist is enforcedMandatory nowFirewall and proxy allowlist evidencePlatform Ops
Network and egressDNS/DHCP/NTP dependencies are defined and testedMandatory nowDependency contract checklist and test logsPlatform Ops
Data and storageStorage classes for NaaS/MaaS are approved and documentedMandatory nowStorage profile matrix and approval recordPlatform Ops
Data and storageDataset and artifact retention classes are definedMandatory nowData retention policy mapping for AI workloadsPlatform Ops
Backup and recoveryOADP schedules are active for AI namespaces plus platform configuration scopeMandatory nowOADP configuration and backup run reportsBackup Ops
Backup and recoveryMonthly restore drills run in original namespace with metadata plus sample PVC dataMandatory nowRestore drill results, remediation tracking records, and retained evidence (12 months)Backup Ops
Observability and operationsLogs, metrics, and alerts are wired to approved platformsMandatory nowMonitoring and alert route validation evidencePlatform Ops
Observability and operationsIncident and escalation runbooks are approvedMandatory nowRunbook approval records and drill evidencePlatform Ops
Model governanceLLMeval and LMjudge operate as advisory stretch-goal capability (if enabled)Stretch goalGate execution reports and scorecard evidence, or defer rationaleAI Platform Team
Model governancePoor-result acceptance record process exists (optional use in this phase)Stretch goalAcceptance record template and any captured approvalsAI Platform Team
Platform change controlProduction changes use approved Git workflowMandatory nowPR history and deployment traceabilityPlatform Ops
Platform change controlEmergency change path is controlled and reviewedMandatory nowEmergency change records and postmortemsPlatform Ops

Expansion automation gate (pilot to broader onboarding)

  • Expansion beyond pilot users requires core controls to be automated.
  • Core controls are tracked as mandatory automation outcomes, without a fixed hardcoded checklist in this phase.

Exit and retirement criteria

  • Primary transition trigger is met: DC 3.0 target-state in DC-B is proven in production with approved readiness evidence.
  • AI services are migrated from DC-A to DC-B after production-proof gate approval.
  • Migration (if needed) is executed as a joint effort by platform and customer teams, in planned service windows with best-effort downtime posture.
  • DC-A transitional lane is sunset after service migration completion.
  • DC-A site is rebuilt to DC 3.0 production-ready target-state after migration.

Control retirement mapping (separate from control catalog)

  • Retirement mapping is maintained as a separate migration artifact that links each transitional control to its DC 3.0 target-state replacement pattern.
  • Retirement mapping includes transition trigger, dependency, migration owner, and validation evidence.
  • Retirement mapping must explicitly track AD/ADFS and Harbor transitional dependencies until replaced by target-state patterns.

Execution focus

This architecture is implemented through Workpackage 21, which contains the handoff-ready engineering package, workstreams, acceptance criteria, and evidence requirements.

Decisions

Delivery Context

Future Work

  • Finalize the detailed installation runbook sequence, including persistent bootstrap host responsibilities and Ceph automation flow from bootstrap.
  • Keep service boundary, control catalog, and OpenShift Lightspeed backend usage boundaries synchronized with implemented operations.
  • Complete the operating model for AI operations ownership and handoff between Platform Ops and the dedicated AI operations function.
  • Add a linked evidence profile for release gates, caveat controls, and override approvals, and keep Workpackage 21 implementation content current.

Open Questions

  • Which retirement trigger and sequence applies for AD- and Harbor-specific transitional dependencies?
  • Which minimum hard-gate controls are required at scale-up checkpoints, given the transitional posture and no BIO2 certification target?
  • Which party owns OpenShift Lightspeed LLM-backend lifecycle operations in each phase (platform, AI operations, and customer teams)?
  • Which migration decision gates and readiness checks determine whether migration to DC-B is required versus continued operation in DC-A?