모델 & 도구 Aug 13, 2026 at 12:579북마크에 추가

WeLM, 위챗의 AI 에이전트를 구동하는 모델은 617억 개의 매개변수로 조용히 확장되었으며, 공개되지 않은 디코딩 메커니즘을 갖추고 있습니다. 벤치마크도, 논문도 없습니다. 단지 수십억 명의 사용자 배포만이 있을 뿐입니다.
간단히 말해:
텐센트의 샤오웨이 에이전트는 위챗 내부에서 그레이 테스트 중이며, 희소 혼합 전문가 모델인 WeLM에서 구동됩니다. WeLM은 조용히 6,170억 개의 매개변수로 성장했지만, 공개 벤치마크도, 논문도 없습니다. 세계에서 가장 트래픽이 많은 메시징 플랫폼 중 하나에서 대규모 그레이 테스트가 진행 중일 뿐입니다.
판데일리(Pandaily)의 보도에 따르면, 위챗 AI 팀은 WeLM의 희소 MoE 버전이 6,170억 개의 매개변수에 도달했으며, 각 토큰당 230억 개의 매개변수를 활성화한다고 밝혔습니다. 텐센트는 WeLM을 공개한 적이 없습니다. 이 모델은 위챗의 통합 AI 어시스턴트인 샤오웨이를 구동하며, 현재 위챗의 13억 이상의 월간 활성 사용자(MAU) 기반에서 그레이 테스트 중입니다.
LLM 경쟁의 리더보드 중심 narrative는 WeLM과 같은 모델을 간과합니다. WeLM은 프런티어 규모의 모델이지만, 완전히 폐쇄되어 있으며 API 우선 연구실이 따라잡기 어려운 분배 우위를 지닌 생산 경로 모델입니다. 알리바바는 벤치마크와 오픈 가중치를 통해 쿤(Qwen)과 경쟁하지만, 텐센트는 임베딩을 통해 경쟁합니다.
내부 구조: 희소 MoE 아키텍처는 각 토큰에 대해 매개변수의 일부만 활성화합니다. WeLM은 추론 단계당 6,170억 개 중 230억 개를 활성화합니다. 이는 약 230억 개의 밀집 모델과 계산 면에서 동등하며, 이 номина스 규모의 모델이 메시징에 적합한 대기 시간으로 소비자용 채팅 인터페이스에서 구동될 수 있는 이유를 설명합니다.
보고서에 공개된 "숨겨진 디코딩 메커니즘"은 사pekulative 디코딩 또는 대기 시간 최적화를 위한 조기 종료 라우팅 전략을 가리킬 가능성이 높습니다. 이는 논문으로 발표되지 않지만 제품을 빠르게 만드는 추론 엔지니어링의 한 예입니다.
WeLM의 공개 벤치마크 공개 여부, 텐센트가 WeLM의 가중치를 공개할지(가능성 낮음), 샤오웨이의 성능이 그레이 테스트 확산에 따라 ChatGPT 및 kimi와 비교해 어떻게 나올지 주시해야 합니다.
인공지능이 작성하고 사람의 편집 감독하에 검수한 기사입니다.
If WeChat’s AI is already deployed at this scale without transparency, isn’t the real question whether we even need benchmarks at this point or just better oversight?
Silent scaling to 617B without disclosure feels like a tech arms race where users are the guinea pigs. Where’s the middle ground between innovation and accountability?
A model this big without metrics is like a black box-sure, it might work for a billion users, but how do we trust it’s not just hype?
Right, but 617B parameters could just mean wasted compute without transparency-how do we know it’s not overfit for WeChat’s niche use cases?
617B parameters is impressive, but without benchmarks or transparency, how do we know it's actually useful for users? Just deploying at scale doesn't guarantee real performance or safety.
Seems like Tencent’s playing both sides-leveraging cutting-edge tech behind the scenes while keeping the rest of us in the dark. Still, a billion-user litmus test might say more than any obscure benchmark ever could.
617B params without benchmarks is like buying a sports car without a speedometer - flashy, but who really knows if it performs? Still, billion-user deployment says something.
Is Tencent’s bet on secret scaling a sign they’re chasing Moore’s Law at all costs, or proof that closed models can outperform open ones in real-world conditions?
The focus should be on whether users actually benefit from this secrecy. Transparency in AI isn’t just for trust-it shapes what gets built next. What’s the endgame here?
What if raw scale without transparency is just hype? If they’re not sharing benchmarks, how do we know it’s not just marketing without substance?
Économie de l'open frontier : viabilité, subvention, pivots