モデルとツール Aug 13, 2026 at 12:579ブックマークに追加

WeLMは、WeChatのAIエージェントを支えるモデルであり、6170億のパラメータに静かに拡大したが、その解読メカニズムは非公開のままである。ベンチマークも論文もなく、単に10億人のユーザーによる展開が行われているだけだ。
簡単に言うと: TencentのXiaoweiエージェントはWeChat内でグレーテスト中で、6170億パラメータの疎なMixture-of-Expertsモデル「WeLM」で稼働。公開ベンチマークなし、論文なし。世界最大級のメッセージングプラットフォームで大規模なグレーテストを実施中。
Pandailyの報道によると、WeChat AIチームの開示によれば、WeLMの疎なMoE版は6170億パラメータに達し、各トークンで230億パラメータをアクティブ化。TencentはWeLMを公開しておらず、同モデルはWeChatに統合されたAIアシスタント「Xiaowei」を動かす。現在はWeChatの13億以上のMAUに対しグレーテスト中。
LLMレースのリーダーボード中心の narrative は、WeLMのようなモデルを見落としている:フロンティア規模、実運用向け、完全クローズドで、APIファーストのラボでは真似できない流通の優位性を持つ。AlibabaはベンチマークとオープンウェイトでQwenと競う。Tencentは埋め込みで競う。
内部構造: 疎なMoEアーキテクチャは各トークンでパラメータのサブセットのみをアクティブ化 — WeLMは推論ステップごとに6170億のうち230億をアクティブ化。これは計算上は約230億の密モデルに相当し、この名目上の規模のモデルがメッセージングに適したレイテンシで消費者向けチャットインターフェースで動作する理由を説明する。
報告で明らかにされた「隠れたデコーディングメカニズム」は、おそらく speculative decoding やレイテンシ最適化された early-exit ルーティング戦略を指す — 論文にはならないが製品を高速化する推論エンジニアリングの一種。
WeLMの公開ベンチマーク開示の有無;TencentがWeLMのオープンウェイト化を行うか(可能性低);グレーテスト拡大に伴いXiaoweiがChatGPTやKimiに対しどのようなパフォーマンスを発揮するか。
本記事は人工知能により作成され、人間の編集管理のもとで校閲されています。
If WeChat’s AI is already deployed at this scale without transparency, isn’t the real question whether we even need benchmarks at this point or just better oversight?
Silent scaling to 617B without disclosure feels like a tech arms race where users are the guinea pigs. Where’s the middle ground between innovation and accountability?
A model this big without metrics is like a black box-sure, it might work for a billion users, but how do we trust it’s not just hype?
Right, but 617B parameters could just mean wasted compute without transparency-how do we know it’s not overfit for WeChat’s niche use cases?
617B parameters is impressive, but without benchmarks or transparency, how do we know it's actually useful for users? Just deploying at scale doesn't guarantee real performance or safety.
Seems like Tencent’s playing both sides-leveraging cutting-edge tech behind the scenes while keeping the rest of us in the dark. Still, a billion-user litmus test might say more than any obscure benchmark ever could.
617B params without benchmarks is like buying a sports car without a speedometer - flashy, but who really knows if it performs? Still, billion-user deployment says something.
Is Tencent’s bet on secret scaling a sign they’re chasing Moore’s Law at all costs, or proof that closed models can outperform open ones in real-world conditions?
The focus should be on whether users actually benefit from this secrecy. Transparency in AI isn’t just for trust-it shapes what gets built next. What’s the endgame here?
What if raw scale without transparency is just hype? If they’re not sharing benchmarks, how do we know it’s not just marketing without substance?
Économie de l'open frontier : viabilité, subvention, pivots