F5 Hardened Release 1 is available. Staying current is one of the most important steps you can take to protect your environment.Learn more

The next era of AI is sharing. Kimi K3 just made it expensive.

Industry Trends | July 24, 2026

Kimi K3 may be another Sputnik moment for AI. Its most important lesson is not simply that open weights can approach the proprietary frontier. Building the model is only half the race; serving it efficiently is the other half. K3 changes both the infrastructure required to deliver frontier intelligence and the economics of who can do so profitably.

Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model. Only 16 of its 896 experts are activated for each token, but efficiently distributing those experts becomes a first-order systems challenge at this scale. Moonshot AI recommends supernode configurations with at least 64 accelerators connected through a high-bandwidth communication domain.

The model may be open. Serving it reliably, quickly, and cost-effectively is where the real technical and commercial challenge begins.

For GPU providers, that gap between model capability and service delivery is where both risk and opportunity emerge. That pressure will only intensify as users, enterprises, and autonomous agents adopt frontier-scale open models. This shift changes the business model for GPU providers. Success depends not only on acquiring accelerators, but on combining them into services that deliver predictable latency, enforce tenant boundaries, preserve host capacity, and generate more customer value from every rack.

Kimi K3 shows that generating tokens efficiently is only part of the challenge. F5 BIG-IP Next for Kubernetes helps providers deliver those tokens securely, reliably, and profitably.

The minimum unit gets bigger

Much of the GPU cloud market has been built around fragmented capacity: individual GPUs and small allocations of one, four, or eight accelerators. Those configurations remain useful for smaller models, development, fine-tuning, and many enterprise workloads. But they are not sufficient to serve K3 efficiently as a unified frontier-scale endpoint.

At four bits per parameter, the weights alone require roughly 1.4 terabytes of storage before accounting for runtime memory, KV cache, and other serving overhead. K3 therefore pushes providers toward larger guaranteed GPU allocations, high-bandwidth scale-up domains, rack-scale systems such as NVL72, topology-aware placement, disaggregated prefill and decode, and continuous KV-cache management.

This raises the minimum viable serving unit from an individual GPU to a tightly coordinated supernode. Capacity that was previously rented in small increments may need to be reorganized into larger, dedicated pools. The weights are shared, but the infrastructure becomes more concentrated and expensive.

Scale-up is only half the problem

Modern frontier-model architectures rely on sophisticated techniques to distribute model execution across many accelerators. Expert parallelism, disaggregated serving, topology-aware placement, and KV-cache management can improve internal model-serving efficiency.

These technologies optimize model execution and east-west communication across the accelerated compute fabric. But every inference request must still enter and leave the serving environment. It must be accepted from a user or agent, authenticated, secured, assigned to a tenant, routed to an available endpoint, observed, metered, and delivered against a service-level objective.

As models such as Kimi K3 increase inference throughput and overall service traffic, the north-south service path becomes an increasingly important control point for scalable, secure, Kubernetes-based AI delivery.

The hidden host tax

Host processors still support networking, proxies, security, Kubernetes services, request handling, telemetry, and other functions surrounding the GPU workload. Every CPU core consumed by infrastructure is a core unavailable to model-serving software, orchestration, or services that keep expensive GPUs productive.

F5 BIG-IP Next for Kubernetes addresses this challenge by consolidating L4-L7 traffic processing, encryption, security, and policy enforcement into the application delivery path. This reduces the burden on host infrastructure and returns more capacity to model-serving software, KV-cache management, storage services, and inference coordination.

BIG-IP Next for Kubernetes also provides accelerated traffic management, segmentation, tenant isolation, firewall, DDoS protection, and policy enforcement. Broader F5 services can extend this protection with WAF, API security, and AI security across the end-to-end environment.

From capacity to customer service

Wide Expert Parallelism and expert load balancing improve internal model-serving efficiency. BIG-IP Next for Kubernetes addresses a different but complementary challenge: turning that capacity into reliable output for customers.

By improving traffic distribution, reducing queueing, enforcing tenant policies, and protecting the service path, it helps operators deliver more consistent inference performance from the same infrastructure.

Even so, the findings demonstrate that north-south delivery architecture can materially improve output-token throughput, time to first token, and cost per token, helping operators produce more reliable, revenue-generating inference from the same GPU environment.

This distinction is particularly important for agents. Agentic systems create chains of model calls, tool invocations, and dependent decisions. Latency at the beginning of one inference step delays every subsequent step. A 20-second delay may frustrate a human; across dozens of sequential agent calls, it can make the service operationally unusable.

Frontier-scale multi-tenancy

K3’s rapid adoption highlights another challenge: model capability can grow faster than the infrastructure available to deliver it. Providers will need to share large serving systems across multiple customers to achieve acceptable economics, but sharing a frontier-scale supernode requires more than dividing GPU capacity.

Each tenant must be isolated. Traffic must be authenticated and governed. Service policies must be consistently enforced. Operators need visibility into consumption and whether one workload is degrading another.

BIG-IP Next for Kubernetes brings Kubernetes-native traffic management, network segmentation, tenant isolation, policy, and observability into the accelerated infrastructure path. This helps transform an expensive rack from a pool of unmanaged compute into a secure, differentiated AI service.

A service plane for inference

K3 demonstrates why AI factories need an AI service plane. The scale-up fabric optimizes GPU communication. Model-serving software executes inference. Schedulers place workloads and allocate GPUs. The AI service plane governs how applications, agents, and tenants consume the infrastructure.

BIG-IP Next for Kubernetes serves as a core traffic-control engine for that service plane, providing secure tenant access, inference traffic management, endpoint resilience, service-level enforcement, consumption visibility, policy, and reliable delivery.

Kimi K3 does not prove that every AI deployment requires a DPU. It does show that frontier inference is becoming a full-stack systems problem and creates exactly the conditions in which dedicated infrastructure processing becomes economically important.

Generating tokens efficiently is only part of the challenge. F5 BIG-IP Next for Kubernetes helps providers turn raw inference capacity into a dependable, secure, and econ

The model may be open. Serving it is not cheap.

Whoever operationalizes the rack and turns its raw capacity into a dependable multi-tenant service captures the margin.

To learn more, explore our F5 BIG-IP Next for Kubernetes and our AI infrastructure solutions webpages.

Share

About the Author

Ahmed Guetari
Ahmed GuetariSenior Vice President, Product Management – Service Provider | F5

More blogs by Ahmed Guetari

Related Blog Posts

Securing the new control points in the AI journey
Industry Trends | 07/01/2026

Securing the new control points in the AI journey

AI architecture is fundamentally different than traditional IT environments and requires a different security strategy to protect critical AI workloads.

The patch window has closed. Here is how F5 is built for what comes next.
Industry Trends | 04/27/2026

The patch window has closed. Here is how F5 is built for what comes next.

As AI models have changed software security, the industry needs to adapt.

Best practices for optimizing AI infrastructure at scale
Industry Trends | 01/21/2026

Best practices for optimizing AI infrastructure at scale

Optimizing AI infrastructure isn’t about chasing peak performance benchmarks. It’s about designing for stability, resiliency, security, and operational clarity

Datos Insights: Securing APIs and multicloud in financial services
Industry Trends | 12/23/2025

Datos Insights: Securing APIs and multicloud in financial services

New threat analysis from Datos Insights highlights actionable recommendations for API and web application security in the financial services sector

Secrets to scaling AI-ready, secure SaaS
Industry Trends | 12/12/2025

Secrets to scaling AI-ready, secure SaaS

Learn how secure SaaS scales with application delivery, security, observability, and XOps.

How AI inference changes application delivery
Industry Trends | 11/19/2025

How AI inference changes application delivery

Learn how AI inference reshapes application delivery by redefining performance, availability, and reliability, and why traditional approaches no longer suffice.