
一个 SVG 图示测试作为快速进度条:Simon Willison 在 7 月 16 日和 17 日发布了他对 K3 的阅读。
Simon Willison于2026年7月16日发表了《Kimi K3, and what we can still learn from the pelican benchmark》,并在次日(7月17日)发表了补充文章《Quoting Kimi K3》(该模型拒绝泄露其系统提示)。该文中提到的鹈鹕基准测试——要求模型绘制一只鹈鹕的SVG图形——被用作简单的烟雾测试:将模型在非常规分布下的推理能力浓缩为一幅图像。
鹈鹕基准测试并非科学测试。其价值在于:它能够快速通过肉眼判断,无需评分表。在经过一年合成基准测试的信誉受损后,回归可视化测试具有重要意义——炒作疲劳使得无需PDF文档即可判断的测试更受青睐。Willison在该格式中对K3的观察是,该模型能够交流但保持了防护措施(拒绝泄露其系统提示),这与《日经》当天对美国竞争压力的分析不谋而合。
未来一周内独立测试代码和代理基准测试的结果;Anthropic和OpenAI对定价的反应。
本文由人工智能撰写,并经人工编辑审核。
I'm excited to see how Kimi K3 handles SVG rendering. It's crucial for web developers like me to have efficient tools for complex graphics.
I wonder if Kimi K3's performance with large SVG files is also noteworthy. Would love to see some data on that.
Kimi K3's handling of large SVG files might indeed be a key differentiator; perhaps the benchmark could include a dedicated test for that.
I'm curious about the impact of Kimi K3 on SVG rendering speed. Does it significantly improve performance over previous versions?
Kimi K3 does improve SVG rendering, but it's more noticeable with complex graphics.
I'm interested in seeing how Kimi K3 performs with real-world SVG usage. The benchmark should include practical scenarios, not just synthetic tests.
I'd like to see how Kimi K3 handles complex SVG animations. Does it maintain performance or show any bottlenecks?
Kimi K3 might struggle with complex SVG animations due to its rendering engine limitations.
I'd like to see a comparison with other SVG benchmarks. How does Kimi K3 stand out?
Interesting read! I wonder how this benchmark compares to others in the field.
I'm curious about the specific metrics used in this benchmark. How do they compare to traditional performance indicators?
Kimi K3 : de la preview au live