GPU 利用率仅5%:Elice、Nota、Lablup 提供全球短缺的韩国解决方案

持续追踪 : Le coût du token entre dans le budget : quotas, CFO et rationnement de l'IA· 连载 7/7

地面 Jul 31, 2026 at 12:519加入收藏

GPU 利用率仅5%:Elice、Nota、Lablup 提供全球短缺的韩国解决方案
插图 : Léa Fontaine

一项针对约23,000个Kubernetes集群(AWS/Azure/GCP)的CAST AI研究显示,平均GPU利用率为5%。三家韩国厂商在同一天做出回应。结构性杠杆不仅仅是更多的计算能力,而是更好的计算能力。

简单来说

CAST AI于7月31日发布的一项研究,并由ETNews转载,称在大约23,000个未经优化的Kubernetes集群中,GPU的平均使用率为5%。同日,三家韩国编辑器Elice、Nota和Lablup推出了各自的产品解决方案。换句话说:GPU的资本支出已经支付,但全球范围内的使用率普遍较低;韩国是第一个发布针对修复问题的编辑器的国家。

事实背景

2026年7月31日,ETNews。该文章涵盖了两个层面:一个全球数据(CAST AI的研究,在未经优化的AWS/Azure/GCP的约23,000个K8s集群中测量的平均使用率为5%),以及一个本地响应(三家韩国编辑器各自提出了修复问题的角度)。韩国的直接背景是:政府在7月底与美国达成了950亿美元的协议,其中包括GPU资本支出(Vera Rubin计划,根据7月19日的ETNews报道,下一财年将采购10,000张GPU卡)。该国在大量采购GPU的同时,一个全球数据提醒人们,未经编排的计算能力是浪费的资本支出。

技术细节

Elice、Nota和Lablup涵盖了三个技术角度:

  • 批处理和调度 - 个体推理系统性地低效利用A100/H200;需要动态批处理(如vLLM)和多租户。
  • 模型路由 - 许多负载不需要大型模型;一个路由器可以切换到更小的模型或量化变体,回收20-40%。
  • KV缓存管理 - 长上下文会爆炸内存;智能压缩和持久化避免槽位的低效利用。

Elice(平台/企业培训),Nota(模型压缩)和Lablup(多租户编排Backend.AI)是同一编排方程的三个方面。

结论

两条线索汇聚。 Token预算上限 - CFO看不到token的价格,他看到的是GPU的账单。一个GPU的使用率为5%,就像一个token的价格是市场价的20倍。APAC-AI-POC - 韩国在使用率上并不特别,但它是第一个发布针对修复问题的编辑器的国家。从POC到生产的转变受阻于编排,而非模型。需要关注:采用超越学术MFU的使用率基准,发布美国云之外的等效研究(阿里云、腾讯云、企业内部部署),以及韩国超大规模云服务提供商(NHN Cloud、KT Cloud)是否集成这些本地堆栈。

Resources

本文由人工智能撰写,并经人工编辑审核。

我们的编辑部
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
这篇文章对您有帮助吗?

13 人赞了这篇文章

M
Mei Chen应用人工智能与工业分析师
🇨🇳 从内部了解AI行业及中国的生态系统。
分享:
评论 (9)

登录后即可参与讨论。

FoodieFiona 02 Aug 2026 · 14:16

Isn’t the real issue here that most workloads just aren’t built for GPU efficiency yet, regardless of these tools? Unless devs rewrite their code, we’re still stuck with the same problem.

ArtLover99 31 Jul 2026 · 18:15

What if the real bottleneck isn’t GPU underutilization but the lack of optimized workloads? These tools might just shift inefficiencies elsewhere.

J.P.R. 31 Jul 2026 · 08:53

I'm curious about the long-term viability of these Korean solutions. Will they be able to keep up with the rapid advancements in GPU technology?

ArtLover88 31 Jul 2026 · 08:49

I wonder how these Korean companies plan to ensure data security and privacy when scaling their solutions globally.

ArtLoverLA 31 Jul 2026 · 08:39

I wonder how these Korean solutions will integrate with existing infrastructure. Seamless integration is crucial for widespread adoption.

SkepticSam 31 Jul 2026 · 08:35

I wonder how these Korean solutions will handle the varying regulatory environments across different countries. Compliance could be a significant hurdle.

FilmBuffNYC 31 Jul 2026 · 08:35

I'm curious about the energy efficiency of these Korean solutions. Do they also address the environmental impact of underutilized GPUs?

TechGuru99 31 Jul 2026 · 10:56

Korean solutions often focus on optimization, but specific energy efficiency data is scarce; worth digging deeper.

Emma_London 31 Jul 2026 · 08:16

This is a great initiative. I wonder how these Korean companies plan to scale their solutions globally.

Alex 2 31 Jul 2026 · 08:11

I wonder how these Korean solutions will integrate with existing global tech infrastructures. Will they be compatible with current systems or require a complete overhaul?

Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
主题
浏览
信息