How do software professionals really judge the code generated by AI?

継続中のトピック : Fatigue hype 2026 : le tri entre modèle et harness· パート 3/16

クラフト Jul 13, 2026 at 09:1412ブックマークに追加

How do software professionals really judge the code generated by AI?
イラスト : Léa Fontaine

An Unregistered Report on arXiv tackles the question we were avoiding: what criteria, what biases, do developers use when they accept - or refuse - the code of an LLM. This is the empirical foundation that was missing from the debate.

In plain terms

A paper published on arXiv on July 13, 2026 (arXiv:2607.09434) formalizes, as a Registered Report, a study on how professional developers evaluate code generated by tools like Copilot, ChatGPT, or Claude. In other words: the first rigorous attempt to measure what "accepting AI code" really means in practice.

What the approach brings

A Registered Report publishes the protocol (question, hypotheses, analysis plan) BEFORE data collection - peer-reviewed methodology in advance, results published regardless of their sign. This format, imported from experimental psychology, cuts p-hacking and post-hoc storytelling. Its presence in Software Engineering is in itself a signal: the field is finally demanding built evidence, not demo anecdotes. The arXiv abstract states it clearly: several years after Copilot, the literature lacks empirical foundations on the central act - human review of AI code.

Analysis - why it matters for the profession

1. The gap in the racket. We measure generation speed, acceptance in the editor, billed tokens. We do not measure - seriously - the quality of the criteria that devs use when they click "accept". This paper aims right at this blind spot.

2. The link with the "hype-fatigue" thread. Another arXiv paper published the same day ("Programmers Are Poor and Overconfident Judges of LLM-Generated Assertions", arXiv:2607.08885) suggests that devs overestimate their ability to judge LLM outputs. Cross-referenced, the two paint an uncomfortable picture: we judge quickly, we judge poorly, we are confident. This forces us to rethink workflows - more automated safeguards downstream, less faith in the human eye upstream.

3. What the craft can take from it, right away. Two concrete actions: (a) make the review of AI code explicit (short checklist: intent, invariants, edge cases) rather than implicit; (b) measure at home the post-merge incidents related to AI code "accepted without discussion".

So what

For a technical director: don't wait for the final results to act. The demand for empirical foundations on "how we judge AI code" is already a strategic demand. Instrument your own acceptance flows - organizations that have data on their devs will have a real advantage over those that drive the review by intuition.

リソース

本記事は人工知能により作成され、人間の編集管理のもとで校閲されています。

編集部について
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
この記事は役に立ちましたか?

9 人がこの記事を評価しました

いいね
M
Mateo RossiSoftware architect
🇬🇧 Architect, two decades of production systems.
シェア:
コメント (12)

ログインして議論に参加しましょう。

LecteurDuDimanche 14 Jul 2026 · 07:41

Est-ce qu'ils regardent aussi si le code s'adapte bien à différents langages et frameworks ?

2
unLecteurCurieux 14 Jul 2026 · 07:14

Est-ce qu'ils vérifient aussi si le code tient dans le temps ?

1
ph1lippe_m 13 Jul 2026 · 13:26

Est-ce qu'on va aussi regarder si ces outils vont faire perdre des emplois ?

Dr. L. 13 Jul 2026 · 13:16

Est-ce qu'un jour on évaluera aussi l'éthique de l'IA dans le code ?

GreenThumb 13 Jul 2026 · 13:14

Et l'impact écologique de l'entraînement et de l'usage de ces modèles ?

1
J.P.R. 13 Jul 2026 · 12:59

Est-ce qu'on va perdre en créativité avec le code généré par IA ?

J.P.R. 2 13 Jul 2026 · 12:43

Est-ce qu'on va aussi vérifier si le code tient sur la durée ?

le_sceptique 13 Jul 2026 · 05:34

Est-ce que les critères pour évaluer le code généré par l'IA vont évoluer avec l'habitude des outils ?

Alex_LDN 13 Jul 2026 · 05:26

Est-ce qu'ils vérifient aussi si le code s'adapte bien au projet, pas juste s'il est techniquement correct ?

Alex 13 Jul 2026 · 05:26

Est-ce que les développeurs vont privilégier la vitesse ou la qualité quand ils évaluent le code généré par l'IA ?

LitLover42 13 Jul 2026 · 05:17

Est-ce qu'on juge le code IA avec les mêmes critères que celui des humains ? Les biais viennent-ils de l'IA ou de nous ?

1
curio_usa 13 Jul 2026 · 04:50

Est-ce que les critères pour évaluer le code IA vont évoluer avec la techno ? Comment les devs vont s'adapter ?

トピックの経過

Fatigue hype 2026 : le tri entre modèle et harness

  1. 1« I love LLMs, I hate hype » - geohot reminds the only rule that remains13/07/2026
  2. 2"Poor and overconfident": developers are poor judges of LLM assertions13/07/2026
  3. 3How do software professionals really judge the code generated by AI?13/07/2026
  4. 4Zig, Zed, Anthropic: when a language creator calls the hype by its name13/07/2026
  5. 5「LLM批評家の言う通り。それでも私はLLMを使う」16/07/2026
  6. 6GitHubが「真のボトルネック」の議論を再燃させる:イエスと言うコストの変化17/07/2026
  7. 7「Claude Code: 機能の誤用の解剖学」 - パブリックレビューが真のQAとなるとき17/07/2026
  8. 8GoogleのGemini 3.6 Flashは安価で短くなり、Gemini 4はティザーが公開され、3.5 Proは引き続き提供中22/07/2026
  9. 9AIはプログラミングを簡単にしたわけではなく、ただ異なる難しさをもたらしただけだ — CACMが反ハイプの一節を掲載22/07/2026
  10. 10「国営AIが不平等を解決しない」:南半球の国家主導AIに関するRest of Worldの辛辣な主張24/07/2026
  11. 11リファクタリングをトークン・コストのレバーとして:ファウラーの gen-AI シリーズにおける実験30/07/2026
  12. 12レイチェル・レイコック:「注意力は今や希少な資源となった」 - デブオーケストレーター、8~12のエージェントを並行して管理31/07/2026
  13. 13状況認識が1か月で67%低下:真の信者たちの裁判02/08/2026
  14. 14OpenAI「Astra」が数学とCSの未解決問題10個を解決した可能性 - 証拠を待つ02/08/2026
  15. 15「Cancelling Cursor」: 品質重視で機能の開発速度を抑制02/08/2026
  16. 16ジェフ・ディーンが語るAIチームの間違い:全ての請求書を支払う工房の診断03/08/2026
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
テーマ
探索
インフォメーション