On this page
- Context and Problem Statement
- Decision Drivers
- Considered Options
- Decision Outcome
- Positive Consequences
- Negative Consequences
- Operational Guardrails for the Hybrid Phase
- Pros and Cons of the Options
- Option 1 - Single EVPN/VXLAN baseline immediately
- Option 2 - ACI baseline for both locations
- Option 3 - Hybrid approach (chosen)
- Links and References
Context and Problem Statement
The rollout requires fast delivery of an AI cluster in Location A while the multi-tenant, multi-cluster platform is built and validated in Location B. A single network baseline across both locations increases short-term delivery risk for the AI lane. We need an explicit transitional model that keeps long-term target-state alignment clear.
Decision Drivers
- Fast-track the AI cluster deployment in Location A with minimal lead time.
- Keep the long-term datacenter network target aligned to EVPN/VXLAN in Location B.
- Prevent uncontrolled dual-operating-model sprawl by hard-scoping the ACI lane.
- Preserve auditability and cutover governance while both lanes exist.
- Avoid ambiguity in platform onboarding scope per location.
Considered Options
- Single EVPN/VXLAN baseline immediately across both locations.
- ACI baseline for both locations in phase 1.
- Hybrid approach: ACI in Location A for AI-only, EVPN/VXLAN in Location B for target platform.
Decision Outcome
Chosen option: Option 3 - Hybrid approach (transitional).
Location A uses a single ACI pod only for the AI cluster use case. Location B remains the target-state build lane for EVPN/VXLAN multi-tenant networking and the multi-cluster platform.
This decision supersedes the single-fabric assumption in adr-016-network-topology-2026 for the current rollout phase.
Positive Consequences
- Reduces short-term delivery risk while the EVPN/VXLAN blueprint is completed for Location B.
- Provides a bounded acceleration lane for AI without blocking Location B platform build-out.
- Keeps target-state investment focused on EVPN/VXLAN in the location that will host the multi-cluster baseline.
Negative Consequences
- Dual-operating models increase complexity and make validation/auditing harder.
- Creates avoidable migration work while not improving long-term alignment.
- Requires strict scope controls so non-AI workloads do not expand into the ACI lane.
Operational Guardrails for the Hybrid Phase
- Location A remains AI-only; non-AI tenant onboarding is blocked by policy and reviewed as an exception.
- Location B EVPN/VXLAN lane must pass readiness evidence before expanding critical workloads (BGP/EVPN health, MTU validation, and route-leak negative tests).
- DC-to-backbone connectivity follows the L3 interconnect model in adr-021-dc-evpn-to-backbone-l3vpn-interconnect; L2 cross-site extension is exception-only.
- Backbone service/transport semantics follow adr-020-backbone-service-vs-transport-l3vpn-sr-mpls to avoid mixed-layer design decisions.
Pros and Cons of the Options
Option 1 - Single EVPN/VXLAN baseline immediately
- Good: Fastest path to one operating model.
- Bad: Higher delivery risk if EVPN/VXLAN readiness for the AI timeline is not met.
Option 2 - ACI baseline for both locations
- Good: Simplifies near-term delivery if ACI capabilities are already available.
- Bad: Delays target-state EVPN/VXLAN maturity and increases future migration effort.
Option 3 - Hybrid approach (chosen)
- Good: Reduces short-term delivery risk while the EVPN/VXLAN blueprint is completed for Location B.
- Bad: Dual-operating models increase complexity and make validation/auditing harder.
- Bad: Creates avoidable migration work while not improving long-term alignment.
Links and References
- Related ADRs: adr-016-network-topology-2026, adr-004-single-active-phased-rollout
- Supporting ADRs: adr-020-backbone-service-vs-transport-l3vpn-sr-mpls, adr-021-dc-evpn-to-backbone-l3vpn-interconnect
- Related topics: topic-datacenter-rollout-intent, topic-network-nxos-evpn-implementation, topic-datacenter-3-0