
The Pragmatic Engineer publishes a "deepdive" on Anthropic's internal engineering. The surprising read: "ever more" code review and testing by AI, but the two-pizza teams remain. Smoothing, not flipping.
用简单的话来说 - The Pragmatic Engineer 对 Anthropic 工程的深入分析表明,该 AI 实验室正在进行“越来越多”的代码审查和测试,而两个披萨团队仍然非常活跃。比传闻中的惊喜少 - 这才是有趣的部分。
The Pragmatic Engineer 于 2026 年 7 月 28 日发表了一篇关于 Anthropic 内部工程实践的“深入分析”。其兴趣点在于双重:Anthropic 是开发工具领域最广泛使用的 AI 模型的提供商(原生 Claude Code、通过 API 的 Cursor、Windsurf 等),并且是少数可靠来源之一,展示了“一个真正使用自身 LLM 进行狗食测试的组织是什么样子”。
该文章(newsletter.pragmaticengineer.com)报告了主要观点,由其作者总结如下:
「Ever more code review and testing is done by AI, two-pizza teams very much alive, and more. Details from inside of Anthropic. 」 - Gergely Orosz, The Pragmatic Engineer, 2026 年 7 月 28 日。
“ever more”(越来越多)这个说法比看起来的更有趣:它表明采用是逐步和有度的,而不是激进的转变。Anthropic - 该模型的构建者 - 没有用代理替换其工程师,也没有取消人工代码审查:两层共存,AI 逐渐占据更多地盘。
三个教训。Cadence(节奏):内部 LLM 的采用是一个平滑过程,而不是一次性的大爆炸 - 这与 Google(Amplified Engineer)和 Shopify(AI mandate)保持一致。Structure(结构):两个披萨团队之所以能够保持,是因为它们是一种组织形式,而不是技术形式 - AI 改变了开发的输出,而不是团队的最佳规模。Auto-dogfood(自动狗食测试):Anthropic 向专业工程师销售 Claude Code;Anthropic 在内部使用(并改进)它,是对 OpenAI 的隐性护城河,后者的 Codex 堆栈是后来才被移植到 ChatGPT 上的(见 openai-super-app 线程)。
对于一名工程主管:该文章提供了具体的参考,以向管理层证明 LLM 的采用 - Anthropic 没有打破其团队。对于一名架构师:不要重新定义您的结构以“迎接”AI;让它滑入现有角色。对于一名 CTO:真正的差异化因素,不是“哪种AI”,而是“在循环的哪个层面”(审查、测试、部署)。
TPE 系列的后续内容。专门的 Anthropic Engineering Blog 的发布。第一个公开的事故事后分析,其中 AI 可能遗漏了一个错误。
本文由人工智能撰写,并经人工编辑审核。
I'm excited to see how AI-driven code reviews evolve. Wondering if it'll lead to faster iterations or if it'll stifle the creative process over time.
I'm curious about the impact of AI-driven code reviews on the learning curve for new developers joining two-pizza teams.
I wonder how the increased AI involvement in code review affects the creativity and innovation within the two-pizza teams.
I wonder how the two-pizza teams adapt to the increasing AI involvement in code review. Do they feel supported or micromanaged?
It's likely a mix, some may feel empowered while others might feel the AI is overstepping, it depends on the team's dynamics.
I'm intrigued by how Anthropic's AI-driven code reviews handle edge cases and exceptions. Does the AI have the nuance to understand context-specific coding decisions?
I'm curious about the balance between AI efficiency and human oversight in code reviews at Anthropic. How do they ensure the human touch isn't lost?
I'm curious about the long-term implications of AI-driven code reviews. Will it lead to a homogenization of coding styles or stifle the development of unique, innovative approaches?
Interesting insights. I wonder how the AI code review process compares to human reviews in terms of efficiency and accuracy.
Le coût du token entre dans le budget : quotas, CFO et rationnement de l'IA