On this page
WP-21 OpenShift AI DC-A Transitional Production Lane
Summary
Deliver a standalone, production-ready OpenShift AI implementation in DC-A (AM1) as a transitional exception lane, with explicit caveats, operational controls, and a governed transition path to DC-B (AM4) target-state operations when needed.
Scope
- Build and operate one OpenShift AI cluster in DC-A with a single
productionservice tier. - Implement this as a standalone cluster (not managed or deployed by ACM).
- Install the cluster through the bootstrap installation process, with a persistent bootstrap host used as jump host and automation runner.
- Deploy the production Ceph cluster through approved automation run from the bootstrap host, and validate readiness before ODF external integration.
- Deliver NaaS (Notebooks-as-a-Service) and MaaS (Models-as-a-Service) services for approved teams under transitional controls.
- Enable LLM-backend usage for OpenShift Lightspeed for this AM1 lane and DC 3.0 development activities in AM4.
- Implement the mandatory production control catalog and produce audit-ready evidence.
- Provide all engineering handover content directly in this workpackage as the implementation contract.
- Prepare and execute transition readiness for migration to DC-B once DC 3.0 production gates are met, if migration is required.
Handover Boundary
- This workpackage is the engineering handoff contract for implementation scope, sequencing, controls, and acceptance.
- Detailed LLD-specific installation inputs (exact DNS records, CIDRs, domain values, proxy/no-proxy lists, and host-level bootstrap artifacts) are owned by engineering implementation design and are not duplicated in architecture notes.
Implementation assumptions and constraints
- AM1 is the delivery location for this transitional lane; AM4 is the target location for DC 3.0 production readiness and potential transition.
- Networking uses IST-provided ACI as-is for this lane, with no platform-side network automation scope.
- Full BIO2 compliance and formal BIO2 certification are not target outcomes for this transitional lane.
- Dedicated AI operations team assignment is pending; interim operational ownership and handoff model must be defined during delivery.
- The bootstrap host is persistent for this lane and is used for jump access and approved automation execution.
Engineering Clarifications
- Finalize exact installation inputs and network values in engineering LLD before starting Phase 2.
- Finalize bootstrap host operating baseline (access model, hardening posture, and toolchain ownership) before installation execution.
- Treat Ceph automation outputs as a hard readiness input for platform storage integration.
- Treat AD group naming and mapping details as a joint implementation task with IAM team during Phase 3.
- Keep registry behavior explicit in runbooks: Harbor pull-through cache default path, Red Hat direct exception path, and operational control of additional direct pulls.
- Use role-based ownership for control evidence and phase exits.
- Pin stable versions at install time and record them in Package A for reproducibility.
- Use fix-forward operations between phases; do not assume rollback as default execution path.
Out of Scope
- Managed model training service commitments.
- Broad multi-tenant self-service onboarding.
- Alignment of this lane to full DC 3.0 target-state parity before initial production use.
- BIO2 certification delivery in this transitional lane.
Architecture Context
- Primary architecture note: openshift-ai-cluster-setup
- Platform domain context: Platform Architecture
- Storage and Ceph baseline context: Ceph Architecture
Decision Context
- adr-022-openshift-ai-dc-a-transitional-production-exception
- adr-004-single-active-phased-rollout
- adr-011-connected-vs-disconnected
- adr-019-hybrid-network-rollout-aci-a-ai-only-evpn-b-target
Dependencies
- No dependency on other workpackages.
- Delivery uses external service dependencies defined in architecture and ADR documents (AD/ADFS, Harbor, Infoblox, NetScaler, Cohesity, Ceph).
Technical Operating Profiles and Baseline
| Area | Implementation profile for engineering |
|---|---|
| Cluster topology | One OpenShift AI cluster in DC-A only; single service tier production. |
| Location model | Delivery location is AM1 (DC-A); transition target location is AM4 (DC-B). |
| Cluster lifecycle model | Standalone cluster lifecycle; no ACM-based deployment or lifecycle management for this lane. |
| Installation method | Bootstrap installation flow (assisted by approved bootstrap automation where applicable). |
| Bootstrap host model | Bootstrap host remains persistent and serves as jump host and approved automation runner. |
| Ceph delivery model | Production Ceph deployment is executed via approved automation from bootstrap and validated before ODF external integration. |
| Hardware profile | Bare-metal control plane with 3 control-plane nodes and 6 GPU worker nodes; use the standard platform hardware profile from OpenShift Platform Hardware Specs. |
| In-scope services | NaaS (Notebooks-as-a-Service) and MaaS (Models-as-a-Service) enabled for approved teams. |
| OpenShift Lightspeed backend | LLM-backend usage is supported for this AM1 lane and DC 3.0 development workflows in AM4. |
| IAM | OIDC via ADFS, AD allowlist for platform-admin and service-consumer; platform-admin has cluster-admin rights; service-consumer has consume-only access. |
| Break-glass | Two local emergency admin accounts (primary and backup), two-person approval, rotate credentials after incident closure. |
| Registry | Harbor pull-through cache is the primary path for most well-known registries; direct path is limited to approved Red Hat registries; any additional direct pull is operationally controlled by Platform Ops runtime policy. |
| Storage | Production Ceph backend aligned to DC 3.0 storage direction via ODF external integration for RBD/CephFS and object where required. |
| Network automation boundary | ACI networking is consumed from IST as-is; no platform-side network automation is delivered in this lane. |
| Backup/restore | Backup scope includes AI namespaces and platform configuration; monthly restore drills in original namespace with metadata plus sample PVC data; evidence retained 12 months. |
| Model governance | LLMeval/LMjudge is a stretch goal for this lane; when enabled it is advisory and non-blocking. |
| Observability | Service-impact paging only; mandatory paging channels are ticket and chat. |
| Expansion gate | Expansion beyond pilot users requires core control automation outcomes to be in place. |
| Version policy | Use stable, explicitly pinned versions for OpenShift, operators, and dependencies at install time. |
| Compliance boundary | BIO2 certification is not a delivery objective for this transitional lane. |
This baseline aligns with the operating profiles in openshift-ai-cluster-setup and serves as the execution profile for implementation and handover.
Engineering Handover Package
Required engineering inputs (SME-owned)
- Final AM1 installation inputs: DNS/LB records, CIDRs, domain values, egress/proxy/no-proxy lists, and host access details.
- Bootstrap host build inputs: OS baseline, hardening controls, identity model, and approved automation toolchain.
- Ceph deployment inputs: node inventory, network mappings, storage profiles, and readiness criteria for ODF external integration.
- IAM and access inputs: ADFS endpoint details, AD group mappings, and break-glass operational approvals.
- Version selection inputs: pinned OpenShift, operator, and dependency versions with compatibility rationale.
Required engineering outputs (SME-owned)
- LLD deliverable set with final implementation decisions, sequence, and rollback/fix-forward posture.
- Executable deployment runbook and validation checkpoints for bootstrap host, Ceph, OpenShift, and OpenShift AI bring-up.
- Operations runbook with ownership model, incident handling, backup/restore drills, and transition decision gates.
- Evidence bundle mapped to phase gates and mandatory control families.
Package A - Implementation specification
- Capture final cluster, service, IAM, registry, storage, and observability configuration values used for deployment.
- Capture approved namespace model, access mappings, and onboarding entry criteria.
Package B - Deployment runbook
- Define ordered bootstrap and configuration sequence from cluster baseline to service enablement.
- Define validation checkpoints after each major setup stage.
Package C - Operations runbook
- Define operational procedures for incident handling, break-glass access, routine change, backup, and restore.
- Define escalation flow and service-impact handling using ticket and chat.
Package D - Validation and evidence profile
- Define mandatory verification checks per control family.
- Define evidence artifacts and retention expectations for governance and audit.
Phase Gate Matrix
| Phase | Entry criteria | Exit criteria | Evidence | Owner (role) |
|---|---|---|---|---|
| Phase 1 - Preparation | ADR and architecture baseline approved; external dependency owners identified | All prerequisites confirmed, hardware readiness validated, bootstrap host baseline approved, install prerequisites signed off | Prerequisite checklist, dependency confirmations, hardware readiness record, bootstrap host baseline record | Platform Ops |
| Phase 2 - Cluster installation | Phase 1 exit signed | Ceph automation executed from bootstrap with readiness validated; standalone OpenShift installation complete; control-plane quorum and worker registration healthy | Ceph automation report, Ceph readiness checks, install log, cluster health checks, bootstrap completion report | Platform Ops |
| Phase 3 - Initial cluster configuration | Phase 2 exit signed | Auth/proxy/registry baseline configured; break-glass controls validated | Config manifests, auth validation, break-glass validation record | Platform Ops |
| Phase 4 - Platform capability deployment | Phase 3 exit signed | Storage/GPU/operator prerequisites installed and validated for OAI; ODF external integration validated against production Ceph | Operator status report, GPU readiness checks, ODF-Ceph integration validation | Platform Ops |
| Phase 5 - OpenShift AI finalization | Phase 4 exit signed | OAI services configured (NaaS/MaaS), go-live controls validated, onboarding gate approved | OAI validation report, control evidence bundle, onboarding sign-off | Platform Ops + Architecture |
Implementation Workstreams
Phase 1 - Preparation
- Confirm prerequisite external dependencies are available and approved for use (AD/ADFS, Harbor, Infoblox, NetScaler, Cohesity, Ceph).
- Confirm hardware readiness for 3 control-plane and 6 GPU worker bare-metal nodes.
- Confirm network lane prerequisites for DC-A AI-only scope.
- Confirm installation inputs, access, and bootstrap artifacts are complete.
- Confirm persistent bootstrap host baseline and approved automation toolchain.
Phase 2 - Cluster installation
- Build and harden the persistent bootstrap host and enable approved automation runtime.
- Execute Ceph deployment automation from bootstrap and validate Ceph readiness outputs.
- Install standalone OpenShift cluster via bootstrap process.
- Validate control-plane quorum, worker registration, and baseline cluster health.
Phase 3 - Initial cluster configuration
- Configure initial cluster settings, including authentication, proxy, registry access, and baseline operational controls.
- Implement OIDC via ADFS and curated AD allowlist role mappings.
- Finalize AD group names and mapping details in coordination with IAM team.
- Implement break-glass primary and backup controls.
- Apply lane network controls under the IST-provided ACI model (no platform-side network automation).
Phase 4 - Platform capability deployment
- Deploy required OpenShift platform prerequisites for AI workloads (storage integration, GPU runtime support, and required operators).
- Integrate ODF external mode with the production Ceph cluster validated in Phase 2.
- Validate cluster readiness for OpenShift AI deployment.
Phase 5 - OpenShift AI finalization
- Deploy OpenShift AI and configure NaaS and MaaS services in the single production tier.
- Configure and validate LLM-backend exposure pattern for OpenShift Lightspeed use in AM1 and DC 3.0 development workflows in AM4.
- If implemented, enable advisory LLMeval/LMjudge execution cadence for release candidates and weekly regressions (stretch goal).
- Finalize production control evidence, restore drill readiness, and transition-readiness documentation.
- Produce and approve Package A through D handover content during each phase.
Cross-phase operating rule
- Execution uses a fix-forward model; rollback is not the default strategy.
Delivery Sequence
- Complete Phase 1 preparation and prerequisite sign-off.
- Execute Phase 2 bootstrap-host baseline, Ceph automation, and OpenShift installation.
- Execute Phase 3 initial cluster configuration.
- Execute Phase 4 platform capability deployment.
- Execute Phase 5 OpenShift AI finalization and production-readiness evidence.
- Start production onboarding for approved teams.
- Execute migration and rebuild sequence when DC-B readiness trigger is approved and migration is required.
Acceptance Criteria
- OpenShift AI DC-A lane is in production with NaaS and MaaS available for approved teams.
- Cluster is installed as a standalone cluster via bootstrap process, without ACM-managed deployment.
- Persistent bootstrap host role is operational (jump host plus approved automation runner).
- Production Ceph deployment from bootstrap automation is validated and integrated through ODF external mode.
- Hardware baseline is implemented and verified: 3 control-plane bare-metal nodes and 6 GPU worker nodes.
- Mandatory production control catalog is implemented with assigned owners and evidence artifacts.
- OpenShift Lightspeed LLM-backend usage path is implemented for AM1 lane use and AM4 development support.
- LLMeval/LMjudge stretch goal status is explicitly documented (implemented or deferred) with rationale.
- Backup and monthly restore drill process is operational with retained evidence.
- Engineering handover package sections in this workpackage are complete, reviewed, and implementation-usable by engineering teams.
- Transition trigger, migration runbook (joint platform and customer execution, if needed), and DC-A rebuild initiation criteria are documented and approved.
- Initial production onboarding requires both technical readiness and architecture sign-off.
Restore Drill Success Criteria
- Restore completes in the original namespace for selected in-scope targets.
- Restored resources are healthy and reachable for the defined validation path.
- Sample PVC data integrity check passes for selected workloads.
- Drill evidence and findings are recorded and retained according to policy.
Required Evidence
- Deployment checkpoint evidence for each workstream.
- IAM, registry, and break-glass control evidence.
- Backup runs and monthly restore drill evidence with retention proof.
- LLMeval/LMjudge execution evidence where stretch goal is implemented, or explicit defer rationale where not implemented.
- Onboarding approvals for in-scope teams.
- Migration readiness evidence and approved transition decision.
Risks and Caveats
- Transitional dependency reuse introduces operational variance versus target-state controls.
- IST-provided ACI with no platform-side network automation increases dependency on external operational lead times.
- Advisory model-governance posture increases reliance on operational judgment for release risk.
- Shared pull-secret model and non-blocking image controls are transitional and require strict operational discipline.
- Dedicated AI operations ownership is not final and can create handoff friction if left unresolved.
- No BIO2 certification target in this lane may constrain use cases that require formal certification.
Transition and Exit
- Transition begins only when DC-B is proven in production and readiness gates are approved.
- AI services migrate from DC-A to DC-B in planned windows with best-effort downtime posture, if migration is required.
- Migration execution is a joint effort by platform and customer teams.
- DC-A transitional lane is retired after migration completion.
- DC-A rebuild to DC 3.0 production-ready state starts after migration closure.
Future Work
- Keep this workpackage synchronized with architecture note updates for operating boundaries, caveats, and migration conditions.
- Refine evidence templates and ownership mappings as the dedicated AI operations function is formalized.
- Add post-go-live optimization backlog items for operations efficiency and control automation maturity.
Open Questions
- Which team takes final steady-state ownership for OpenShift Lightspeed LLM-backend operations?
- Which readiness threshold defines when migration to AM4 is required versus continued AM1 operation?
- Which minimum additional controls are required before scaling from pilot onboarding to broader production onboarding?