Introduction
Over the last several months I have been spending a considerable amount of my study time on AI infrastructure. My background is predominantly network engineering, datacenter, automation and infrastructure architecture, subsequently when I initially started looking into AI infrastructure my natural focus was around the physical infrastructure. The GPU servers, high speed backend network, RoCEv2, storage and how all these different infrastructure components connect together.
However as I started going deeper into the actual architecture, Kubernetes(K8s) continually appeared. Whether I was reading about inference, training, GPU clusters or deploying inference engines such as vLLM, Kubernetes was almost always somewhere within the architecture. Subsequently for me the important aspect was not simply learning Kubernetes configuration, but actually understanding why we need Kubernetes within AI infrastructure in the first place.
This is something I have always found important when studying technology. Before concentrating on commands and actual implementation, I like to understand the purpose and theoretical aspect behind the technology. Once I understand the problem the technology is trying to solve, the actual implementation usually starts making considerably more sense.
Within AI infrastructure we already have a significant amount of technologies. We have the GPU providing the compute, CUDA allowing applications to utilise that GPU, vLLM potentially providing the inference engine, NCCL providing GPU communication and technologies such as RoCEv2 or InfiniBand providing the high performance network transport. Therefore my initial question was relatively simple, where exactly does Kubernetes fit into all of this?
Fortunately as a network professional Kubernetes isn’t alien to me. I studied and own a Kubernetes cluster lab. It’s a simple bare metal Kubernetes cluster with control plane node and three worker nodes.
The easiest way I have found to understand it is that Kubernetes provides the orchestration layer around the AI workload. It does not perform the actual inference, it does not train the model and it certainly does not replace the physical network infrastructure underneath it. Kubernetes provides a platform where we can deploy, schedule, scale and manage the workloads that consume all these different infrastructure resources.
Where Kubernetes fits within AI Infrastructure

Once I looked at Kubernetes from this perspective the overall architecture became considerably clearer. Kubernetes is not replacing any of these technologies, it essentially sits above the infrastructure and provides a common method of controlling where and how the applications consume those underlying resources.
Starting with the Physical GPU Infrastructure
Coming from a network and datacenter background, I find it considerably easier to understand the Kubernetes architecture if I start at the physical infrastructure and work my way upwards. Let’s imagine we have four physical GPU servers and every server contains eight NVIDIA GPUs, subsequently we have a total of 32 GPUs available within our AI infrastructure.
There is absolutely nothing stopping us from installing Ubuntu directly onto those GPU servers, installing our NVIDIA drivers, CUDA, Python libraries and an inference engine such as vLLM. We could subsequently SSH directly into a server, load an LLM and begin using the GPUs for inference. In fact for a small lab environment this is probably one of the easiest methods to initially learn how all the different AI components work.
However let’s expand this beyond the lab. We now have multiple development teams using the infrastructure, several different models, RAG applications, inference workloads and potentially some fine-tuning or training workloads. At this stage manually deciding which application runs on which physical GPU server starts becoming considerably more difficult.
This is where Kubernetes starts to become extremely useful. The GPU servers can essentially become Kubernetes worker nodes, while the Kubernetes control plane is responsible for controlling and orchestrating workloads across the cluster. The control plane itself does not inherently need the expensive GPUs because its primary purpose is not performing the AI computation, it is managing the state of the cluster and deciding where workloads should operate.
GPU Servers as Kubernetes Worker Nodes

The key component for me here is that Kubernetes does not somehow create the GPU infrastructure, the physical resources still need to exist underneath it. Kubernetes essentially gives us a mechanism of viewing and consuming those resources as part of a larger platform instead of treating every physical server as an individual entity.
Kubernetes provides stable support for managing GPUs through the device-plugin framework. Once the appropriate vendor plugin is installed, resources such as NVIDIA GPUs can be exposed to Kubernetes as schedulable resources and subsequently requested by a container in a similar manner to requesting CPU or memory. (Kubernetes)
Therefore instead of me manually looking at GPU Server 3 and deciding that four GPUs are currently available, the workload can declare the GPU resources it requires and Kubernetes can make a scheduling decision based on the available infrastructure.
Where does the Container and Pod fit?
At this stage I now understand that my physical GPU servers can become Kubernetes worker nodes, however the next part I wanted to understand was how the actual AI application gets onto those GPUs. This is where containers and Kubernetes Pods come into the architecture.
Let’s use vLLM as an example. vLLM provides an inference engine which can load and serve an LLM, subsequently the purpose of vLLM is completely different to Kubernetes. Kubernetes is responsible for orchestrating the workload, while vLLM is actually responsible for efficiently performing the inference workload.
We can package the required application and dependencies into a container image and Kubernetes subsequently runs that container inside a Pod. That Pod is then scheduled onto an appropriate Kubernetes worker node which has access to the required GPU resources.
The architecture therefore starts to look something like this.

For me this separation was important to understand because when initially learning AI infrastructure there are a significant amount of technologies and it is relatively easy to start mixing up what each one is responsible for. The physical GPU provides the compute, CUDA provides the GPU compute platform, vLLM provides the inference engine and Kubernetes provides the orchestration around the application.
The NVIDIA GPU Operator provides a good example of how this is implemented at scale. NVIDIA uses the Kubernetes operator framework to automate software components required on GPU nodes, including NVIDIA drivers, the GPU device plugin, NVIDIA Container Toolkit, GPU node labelling and DCGM based monitoring. (NVIDIA Docs)
This is where we can start seeing the difference between simply having some GPU servers and actually building a GPU platform.
Why do we actually need Kubernetes?
Let’s go back to our previous example of four physical GPU servers containing eight GPUs per server. Subsequently we have 32 GPUs available within our infrastructure, however now several different teams need to use them.
One team may be running a RAG application using an LLM for inference, another team could be testing a completely different model, while another team may be performing fine-tuning. Without any type of orchestration somebody inherently needs to decide where each workload should run, how many GPUs each team is allowed to use and what happens when those applications are stopped, upgraded or fail.
In a small environment this can probably be managed manually, however as the amount of infrastructure starts increasing the operational complexity also increases. We have to determine which GPUs are available, whether a particular server has sufficient GPU memory, where another model should be placed and potentially what should happen if a physical node fails.
Kubernetes gives us the ability to describe what resources the workload requires and allow the platform to schedule that workload onto the available infrastructure. Subsequently we are moving away from thinking purely in terms of individual GPU servers and instead starting to treat the overall environment as a shared compute platform.
I think this is an extremely important aspect when discussing enterprise AI infrastructure. Simply purchasing several powerful GPU servers does not inherently give us an AI platform. We still require a method of efficiently consuming, controlling and operating those resources, especially when multiple users and applications need access to the same infrastructure.
Kubernetes and the Network
This is where the topic becomes particularly interesting for me personally because my primary background is network and datacenter infrastructure. Kubernetes can make a decision regarding where a particular workload is scheduled, however that decision can inherently affect where the network traffic will flow.
Let’s imagine we have an LLM running across eight GPUs and all eight GPUs are located inside the same physical server. Depending on the hardware platform, the GPUs may be interconnected using high bandwidth technologies such as NVLink or NVSwitch, subsequently a significant amount of the GPU-to-GPU communication can remain inside the physical server.
However what happens when our model or workload requires sixteen GPUs?
If our physical servers only contain eight GPUs each, the workload may subsequently need to span multiple physical GPU servers. At this stage the GPU communication can no longer remain completely within one server because traffic now has to leave the server, traverse the backend network and reach GPUs located within another physical server.
This is where Kubernetes, compute and networking start becoming extremely interconnected.
When the AI workload leaves the server

At this stage the network is no longer simply carrying conventional application traffic. Depending on the workload, we may have extremely high bandwidth GPU-to-GPU communication traversing the backend fabric and technologies such as NCCL may be performing collective communication across GPUs residing within multiple physical servers.
If we are using a RoCEv2 based backend network, technologies such as ECN, PFC and DCQCN can become extremely important for congestion management and maintaining the characteristics required by RDMA traffic. Kubernetes may have been responsible for orchestrating the workload, however it certainly cannot compensate for a poorly designed physical backend network.
This is why I think the network becomes even more important within distributed AI infrastructure rather than less important. The abstraction provided by Kubernetes does not mean the physical network somehow disappears, the packets still need to physically move between the GPU servers.
Frontend and Backend AI Networks
Another aspect which I found useful when learning AI infrastructure was separating the frontend AI network from the backend AI network. Although both networks ultimately support the same AI platform, the traffic characteristics and purpose are fundamentally different.
The frontend network is essentially where users, applications or AI agents access the AI service. For example an enterprise application may send an API request towards an inference service running vLLM inside a Kubernetes environment. This traffic is fundamentally north-south and may pass through traditional infrastructure components such as load balancers, API gateways, firewalls and application security controls.
The backend AI network has a completely different purpose. This is where communication between GPU servers occurs when a workload is distributed across the infrastructure. During something such as distributed training there can be a significant amount of east-west traffic, while large distributed inference workloads using Tensor Parallelism can also create GPU-to-GPU communication across the backend fabric.
Frontend vs Backend AI Network

Subsequently when we say Kubernetes schedules an AI workload, from a network architecture perspective we should also be thinking about what network traffic is going to be generated as a result of that decision.
For example if the entire inference workload remains inside one GPU server, the backend network may see relatively little GPU communication for that particular workload. However if that same workload is now distributed across four physical GPU servers, the traffic matrix has fundamentally changed even though from an application perspective it may still appear to be the same model.
This is why I continually come back to one particular point when studying AI infrastructure, the workload dictates the traffic flow. We cannot design the network correctly without understanding what the workload is actually doing.
Kubernetes Scheduling and GPU Topology
At this stage another important aspect becomes apparent, not every group of eight GPUs is necessarily equal from an infrastructure perspective.
Let’s imagine Kubernetes needs to allocate eight GPUs to a workload. Eight GPUs located inside one physical server can potentially communicate using the internal GPU interconnects of that server, while eight GPUs spread across multiple physical servers inherently require more communication across the network.
Both scenarios may technically provide the application with eight GPUs, however the underlying communication topology is significantly different.

From a network perspective this is extremely important because workload placement can directly influence how much east-west traffic appears on the backend fabric. Subsequently Kubernetes scheduling within AI infrastructure can become considerably more advanced than simply asking whether a node has sufficient CPU or RAM.
The physical GPU topology, available GPU memory, network placement and potentially the locality of storage can all become important considerations when deciding where the workload should run.
This also demonstrates why AI infrastructure cannot really be treated as individual silos. The compute architects cannot completely ignore the network, while equally the network architects cannot design the fabric without understanding the compute topology and workloads operating on top of it.
Kubernetes and Scaling vLLM Inference
Another area where Kubernetes becomes extremely useful is inference scaling. Let’s imagine we have deployed an enterprise chatbot and the LLM is being served using vLLM. Initially perhaps one inference instance is enough to support the number of users accessing the platform, however as adoption increases we may eventually reach the point where additional inference capacity is required.
Instead of manually installing another instance onto another physical GPU server, Kubernetes allows us to deploy and manage multiple replicas of the inference service. Traffic can subsequently be distributed across those instances, giving us a considerably more scalable architecture.
There is however another distinction which I think is important to understand, scaling the amount of inference instances is not necessarily the same thing as splitting one model across several GPUs.
If the model itself is too large for one GPU, we may use something such as Tensor Parallelism(TP) to split that model across multiple GPUs. If we then require more capacity to support additional concurrent users, we can potentially create additional model replicas.
Therefore we could theoretically have several vLLM Pods, while each individual vLLM instance itself uses several GPUs.
Scaling an AI Inference Service

For me this is where we start moving away from simply running an LLM and start talking about designing an actual AI inference platform.
A single engineer running a model on one GPU server is relatively straightforward. When we start supporting hundreds or potentially thousands of requests, multiple models, different development teams and numerous GPU servers, we inherently need a platform capable of orchestrating all those different components.
Kubernetes subsequently becomes one of the technologies helping us achieve that.
What about Training?
Although I have concentrated heavily on inference so far, the same orchestration concept also applies to training workloads.
Let’s imagine we have a distributed training job requiring four physical GPU servers with eight GPUs per server. Subsequently our workload needs a total of 32 GPUs and those GPUs need to work together to perform the training.
Kubernetes can assist in orchestrating and placing those workloads across the available worker nodes, however the actual model training is still performed by the machine learning framework while technologies such as NCCL provide the communication between the GPUs.
Once again the network becomes a critical component because distributed training can generate enormous amounts of east-west traffic. The Kubernetes layer may orchestrate where the job operates, but underneath Kubernetes our backend network still has to physically transport the communication between those different servers.
Therefore the overall architecture starts to become relatively easy to separate.

Every component has a different purpose, however collectively they form part of the same AI infrastructure architecture.
Storage is also Part of the Architecture
We also cannot ignore storage because AI workloads inherently require access to data. Training workloads may need to process very large datasets, inference environments need access to model weights and RAG based applications require access to external enterprise data and vector databases.
Again Kubernetes does not replace the underlying storage infrastructure. Instead it provides mechanisms which allow applications running inside the cluster to consume persistent storage resources.
This becomes extremely important when we start thinking about the lifecycle of containers. A container can be recreated or moved without us inherently wanting the underlying data to disappear with it, subsequently persistent storage gives us the separation between the application lifecycle and the actual data.
For training this can become particularly important because model checkpoints may need to be retained throughout a long-running training process. For inference we may also need fast access to very large model files when instances are initially created.
Therefore just like the network, storage has a direct relationship with the workload operating within Kubernetes.
Kubernetes does not replace the Infrastructure
For me this is probably one of the most important aspects to take away from the entire topic. Kubernetes is extremely powerful, however Kubernetes does not replace good infrastructure design.
If the GPU does not have sufficient memory to load the model, Kubernetes cannot fix that. If our backend fabric is congested and the GPU servers cannot communicate efficiently, Kubernetes cannot magically provide additional network bandwidth. Equally if our storage infrastructure cannot deliver the required data fast enough, orchestration alone will not resolve the underlying performance problem.
Kubernetes essentially provides the mechanism of bringing these resources together and operating workloads on top of them.
I think this is where the role of the modern infrastructure architect becomes particularly interesting because many technologies that were traditionally treated as completely separate infrastructure domains are becoming extremely interconnected within AI.
The model can dictate the GPU memory requirements, the GPU topology can dictate the communication requirements, the communication requirements can dictate the network architecture and Kubernetes can influence where those workloads are actually placed.
Subsequently a decision made at one layer can have a direct impact on another layer.
Why Network Engineers should understand Kubernetes
Personally I do not believe every network engineer suddenly needs to become a Kubernetes administrator simply because AI infrastructure is becoming more common. However I do think we need to understand the fundamentals because Kubernetes can directly influence where workloads operate and subsequently where network traffic flows.
As network engineers we have traditionally spent significant amounts of time understanding endpoints, routing, switching, traffic flows and how applications communicate. Kubernetes does not fundamentally remove this requirement, it simply introduces another layer that can dynamically control where those endpoints and workloads are placed.
From an AI infrastructure perspective this becomes even more important because the traffic between GPU servers can be significantly different from the conventional application traffic we are used to carrying across enterprise networks.
Understanding that Kubernetes has scheduled a distributed workload across four GPU nodes immediately provides useful information to the network engineer. We can now start asking whether the GPUs are communicating across the backend fabric, which leaf switches are involved, whether NCCL performance is normal, whether ECN marking is increasing or perhaps whether PFC pause frames are appearing on a particular path.
This is where Kubernetes knowledge starts becoming directly relevant to network operations and observability rather than simply being a subject owned by the application team.
Conclusion
When I initially started looking at Kubernetes within AI infrastructure I simply wanted to understand where it actually sits within the overall architecture. Once I understood its purpose, many of the other components started fitting together considerably more logically.
The GPU provides the physical compute, CUDA provides the GPU computing platform, the model provides the actual intelligence, vLLM can provide the inference engine, NCCL can provide communication between GPUs and RoCEv2 or InfiniBand provides the high performance backend transport. Kubernetes subsequently provides the orchestration layer which helps deploy, schedule, scale and manage those workloads across the available infrastructure.
The key component for me personally is understanding that Kubernetes does not replace the infrastructure underneath it, it orchestrates how that infrastructure is consumed.
For network engineers this is particularly relevant because where Kubernetes places a workload can inherently dictate where network traffic appears. If a workload remains within one physical GPU server then much of the GPU communication may remain local, however once that same workload starts spanning multiple GPU servers the backend network becomes part of the actual AI compute path.
This is why I think AI infrastructure is such an interesting area for infrastructure engineers. Compute, storage, networking, automation, Kubernetes and AI software are no longer completely isolated topics, they are increasingly becoming different layers of the same architecture.
Personally this is exactly how I want to continue learning AI infrastructure. Understand the theory, understand the purpose, build it within my lab environment and subsequently observe what actually happens to the compute and network when the workloads start running. As with most technologies, once you actually implement it and see the traffic flow for yourself, the theory starts making considerably more sense.




Leave a comment