Underutilized GPUs at 5%: Elice, Nota, Lablup offer the Korean answer to a global crunch

Ongoing story : Le coût du token entre dans le budget : quotas, CFO et rationnement de l'IA· Part 7/7

GroundSubscribers only 23 min ago7Add to bookmarks

Underutilized GPUs at 5%: Elice, Nota, Lablup offer the Korean answer to a global crunch
Illustration : Léa Fontaine

A CAST AI study on ~23,000 Kubernetes clusters (AWS/Azure/GCP): 5% average GPU utilization. Three Korean publishers respond the same day. The structural lever is not more compute, it's better compute.

In plain terms

A CAST AI study published on July 31 and reported by ETNews estimates the average GPU usage to be 5% across approximately 23,000 Kubernetes clusters operated by the three major clouds (AWS, Azure, GCP) when Kubernetes is not optimized. Three Korean publishers - Elice, Nota, Lablup - respond on the same day with their product solutions. In other words: the GPU capex is already being paid for, usage is lagging worldwide; Korea is the first to publish publishers positioned on the fix.

The fact, in context

ETNews, July 31, 2026. The article aggregates two levels: a GLOBAL figure (the CAST AI study, 5% average usage measured on ~23,000 K8s clusters operated by AWS/Azure/GCP when Kubernetes is not optimized), and a LOCAL RESPONSE (three Korean publishers each offering an angle on the fix). Immediate context for Korea: the government completed a $950 billion package with the United States at the end of July, including a GPU capex component (Vera Rubin, a plan for 10,000 cards in the following fiscal year according to ETNews on July 19). The country is massively purchasing while a global figure reminds us that unorchestrated compute is wasted capex.

Under the hood

Three technical angles covered by Elice, Nota, and Lablup:

  • Batching and scheduling - individual inference systematically underutilizes an A100/H200; dynamic batching (like vLLM) and multi-tenancy are needed.
  • Model routing - much of the load does not need the large model; a router that switches to a smaller model or a quantized variant recovers 20-40%.
  • KV cache management - long contexts explode memory; intelligent compression and persistence avoid underutilization of slots.

Elice (training/enterprise platforms), Nota (model compression), and Lablup (multi-tenant orchestration Backend.AI) are the three faces of the same orchestration equation.

So what

Two threads converge. Token-budget-caps - the CFO does not see the price of the token, they see the GPU bill. A GPU at 5% utilization is the same as a token at 20x the market price. APAC-AI-poc - Korea is not particular about the utilization rate, but it is the first to publish publishers positioned on the fix. The transition from POC to production is hindered by orchestration, not the model. To watch: adoption of usage benchmarks beyond academic MFU, publication of equivalent studies outside of US-cloud (Alibaba Cloud, Tencent, on-prem enterprise), and whether Korean hyperscalers (NHN Cloud, KT Cloud) integrate these native stacks.

Content reserved for members

Create a free account to access all our content and the weekly review.

Article produced by artificial intelligence, reviewed under human editorial control.

Our newsroom
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
Was this article helpful?

7 people liked this article

Like
M
Mei ChenApplied AI & Industry Analyst
Follow the AI industry, including the Chinese ecosystem, from the inside.
Share:
Comments (7)

Sign in to join the discussion.

J.P.R. 31 Jul 2026 · 08:53

I'm curious about the long-term viability of these Korean solutions. Will they be able to keep up with the rapid advancements in GPU technology?

ArtLover88 31 Jul 2026 · 08:49

I wonder how these Korean companies plan to ensure data security and privacy when scaling their solutions globally.

ArtLoverLA 31 Jul 2026 · 08:39

I wonder how these Korean solutions will integrate with existing infrastructure. Seamless integration is crucial for widespread adoption.

SkepticSam 31 Jul 2026 · 08:35

I wonder how these Korean solutions will handle the varying regulatory environments across different countries. Compliance could be a significant hurdle.

FilmBuffNYC 31 Jul 2026 · 08:35

I'm curious about the energy efficiency of these Korean solutions. Do they also address the environmental impact of underutilized GPUs?

TechGuru99 31 Jul 2026 · 10:56

Korean solutions often focus on optimization, but specific energy efficiency data is scarce; worth digging deeper.

Emma_London 31 Jul 2026 · 08:16

This is a great initiative. I wonder how these Korean companies plan to scale their solutions globally.

Alex 2 31 Jul 2026 · 08:11

I wonder how these Korean solutions will integrate with existing global tech infrastructures. Will they be compatible with current systems or require a complete overhaul?

Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
Topics
Explore
Information