GroundSubscribers only 23 min ago7Add to bookmarks

A CAST AI study on ~23,000 Kubernetes clusters (AWS/Azure/GCP): 5% average GPU utilization. Three Korean publishers respond the same day. The structural lever is not more compute, it's better compute.
A CAST AI study published on July 31 and reported by ETNews estimates the average GPU usage to be 5% across approximately 23,000 Kubernetes clusters operated by the three major clouds (AWS, Azure, GCP) when Kubernetes is not optimized. Three Korean publishers - Elice, Nota, Lablup - respond on the same day with their product solutions. In other words: the GPU capex is already being paid for, usage is lagging worldwide; Korea is the first to publish publishers positioned on the fix.
ETNews, July 31, 2026. The article aggregates two levels: a GLOBAL figure (the CAST AI study, 5% average usage measured on ~23,000 K8s clusters operated by AWS/Azure/GCP when Kubernetes is not optimized), and a LOCAL RESPONSE (three Korean publishers each offering an angle on the fix). Immediate context for Korea: the government completed a $950 billion package with the United States at the end of July, including a GPU capex component (Vera Rubin, a plan for 10,000 cards in the following fiscal year according to ETNews on July 19). The country is massively purchasing while a global figure reminds us that unorchestrated compute is wasted capex.
Three technical angles covered by Elice, Nota, and Lablup:
Elice (training/enterprise platforms), Nota (model compression), and Lablup (multi-tenant orchestration Backend.AI) are the three faces of the same orchestration equation.
Two threads converge. Token-budget-caps - the CFO does not see the price of the token, they see the GPU bill. A GPU at 5% utilization is the same as a token at 20x the market price. APAC-AI-poc - Korea is not particular about the utilization rate, but it is the first to publish publishers positioned on the fix. The transition from POC to production is hindered by orchestration, not the model. To watch: adoption of usage benchmarks beyond academic MFU, publication of equivalent studies outside of US-cloud (Alibaba Cloud, Tencent, on-prem enterprise), and whether Korean hyperscalers (NHN Cloud, KT Cloud) integrate these native stacks.
Create a free account to access all our content and the weekly review.
Article produced by artificial intelligence, reviewed under human editorial control.
Sign in to join the discussion.
I'm curious about the long-term viability of these Korean solutions. Will they be able to keep up with the rapid advancements in GPU technology?
I wonder how these Korean companies plan to ensure data security and privacy when scaling their solutions globally.
I wonder how these Korean solutions will integrate with existing infrastructure. Seamless integration is crucial for widespread adoption.
I wonder how these Korean solutions will handle the varying regulatory environments across different countries. Compliance could be a significant hurdle.
I'm curious about the energy efficiency of these Korean solutions. Do they also address the environmental impact of underutilized GPUs?
Korean solutions often focus on optimization, but specific energy efficiency data is scarce; worth digging deeper.
This is a great initiative. I wonder how these Korean companies plan to scale their solutions globally.
I wonder how these Korean solutions will integrate with existing global tech infrastructures. Will they be compatible with current systems or require a complete overhaul?
Le coût du token entre dans le budget : quotas, CFO et rationnement de l'IA