How do software professionals really judge the code generated by AI?

진행 중인 이슈 : Fatigue hype 2026 : le tri entre modèle et harness· 편 3/16

크래프트 Jul 13, 2026 at 09:1412북마크에 추가

How do software professionals really judge the code generated by AI?
삽화 : Léa Fontaine

An Unregistered Report on arXiv tackles the question we were avoiding: what criteria, what biases, do developers use when they accept - or refuse - the code of an LLM. This is the empirical foundation that was missing from the debate.

In plain terms

A paper published on arXiv on July 13, 2026 (arXiv:2607.09434) formalizes, as a Registered Report, a study on how professional developers evaluate code generated by tools like Copilot, ChatGPT, or Claude. In other words: the first rigorous attempt to measure what "accepting AI code" really means in practice.

What the approach brings

A Registered Report publishes the protocol (question, hypotheses, analysis plan) BEFORE data collection - peer-reviewed methodology in advance, results published regardless of their sign. This format, imported from experimental psychology, cuts p-hacking and post-hoc storytelling. Its presence in Software Engineering is in itself a signal: the field is finally demanding built evidence, not demo anecdotes. The arXiv abstract states it clearly: several years after Copilot, the literature lacks empirical foundations on the central act - human review of AI code.

Analysis - why it matters for the profession

1. The gap in the racket. We measure generation speed, acceptance in the editor, billed tokens. We do not measure - seriously - the quality of the criteria that devs use when they click "accept". This paper aims right at this blind spot.

2. The link with the "hype-fatigue" thread. Another arXiv paper published the same day ("Programmers Are Poor and Overconfident Judges of LLM-Generated Assertions", arXiv:2607.08885) suggests that devs overestimate their ability to judge LLM outputs. Cross-referenced, the two paint an uncomfortable picture: we judge quickly, we judge poorly, we are confident. This forces us to rethink workflows - more automated safeguards downstream, less faith in the human eye upstream.

3. What the craft can take from it, right away. Two concrete actions: (a) make the review of AI code explicit (short checklist: intent, invariants, edge cases) rather than implicit; (b) measure at home the post-merge incidents related to AI code "accepted without discussion".

So what

For a technical director: don't wait for the final results to act. The demand for empirical foundations on "how we judge AI code" is already a strategic demand. Instrument your own acceptance flows - organizations that have data on their devs will have a real advantage over those that drive the review by intuition.

Resources

인공지능이 작성하고 사람의 편집 감독하에 검수한 기사입니다.

편집팀
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
이 기사가 도움이 되었나요?

9 명이 이 기사를 좋아합니다

좋아요
M
Mateo RossiSoftware architect
🇬🇧 Architect, two decades of production systems.
공유:
댓글 (12)

토론에 참여하려면 로그인하세요.

LecteurDuDimanche 14 Jul 2026 · 07:41

Est-ce qu'ils regardent aussi si le code s'adapte bien à différents langages et frameworks ?

2
unLecteurCurieux 14 Jul 2026 · 07:14

Est-ce qu'ils vérifient aussi si le code tient dans le temps ?

1
ph1lippe_m 13 Jul 2026 · 13:26

Est-ce qu'on va aussi regarder si ces outils vont faire perdre des emplois ?

Dr. L. 13 Jul 2026 · 13:16

Est-ce qu'un jour on évaluera aussi l'éthique de l'IA dans le code ?

GreenThumb 13 Jul 2026 · 13:14

Et l'impact écologique de l'entraînement et de l'usage de ces modèles ?

1
J.P.R. 13 Jul 2026 · 12:59

Est-ce qu'on va perdre en créativité avec le code généré par IA ?

J.P.R. 2 13 Jul 2026 · 12:43

Est-ce qu'on va aussi vérifier si le code tient sur la durée ?

le_sceptique 13 Jul 2026 · 05:34

Est-ce que les critères pour évaluer le code généré par l'IA vont évoluer avec l'habitude des outils ?

Alex_LDN 13 Jul 2026 · 05:26

Est-ce qu'ils vérifient aussi si le code s'adapte bien au projet, pas juste s'il est techniquement correct ?

Alex 13 Jul 2026 · 05:26

Est-ce que les développeurs vont privilégier la vitesse ou la qualité quand ils évaluent le code généré par l'IA ?

LitLover42 13 Jul 2026 · 05:17

Est-ce qu'on juge le code IA avec les mêmes critères que celui des humains ? Les biais viennent-ils de l'IA ou de nous ?

1
curio_usa 13 Jul 2026 · 04:50

Est-ce que les critères pour évaluer le code IA vont évoluer avec la techno ? Comment les devs vont s'adapter ?

이슈 타임라인

Fatigue hype 2026 : le tri entre modèle et harness

  1. 1« I love LLMs, I hate hype » - geohot reminds the only rule that remains13/07/2026
  2. 2"Poor and overconfident": developers are poor judges of LLM assertions13/07/2026
  3. 3How do software professionals really judge the code generated by AI?13/07/2026
  4. 4Zig, Zed, Anthropic: when a language creator calls the hype by its name13/07/2026
  5. 5"LLM 비판자들은 옳아. 그래도 나는 LLMs를 사용해" - 재구성하는 목소리16/07/2026
  6. 6GitHub가 "예스"라고 말하는 비용이 변했습니다: GitHub가 진정한 병목 현상에 대한 논쟁을 재점화합니다17/07/2026
  7. 7« Claude Code: 해로운 기능의 해부 » - 공개 리뷰가 진정한 QA가 되는 순간17/07/2026
  8. 8Google의 Gemini 3.6 Flash는 더 저렴하고 짧아졌으며, Gemini 4는 teas를 받지만 3.5 Pro는 늦게 유지됩니다.22/07/2026
  9. 9AI가 프로그래밍을 더 쉽게 만들지 않았으며, 단지 다르게 어렵게 만들었을 뿐입니다 - CACM이 반하이프 라인을 제시합니다.22/07/2026
  10. 10국가 소유 AI가 불평등을 해결하지 못할 것이라는 레스트 오브 월드의 냉정한 thesis24/07/2026
  11. 11리팩토링을 토큰 비용 레버로: 파울러의 Gen-AI 시리즈 실험30/07/2026
  12. 12레이첼 레이콕 : “‘주의’가 희귀한 자원이 되었습니다” - 8~12명의 에이전트를 동시에 관리하는 개발 오케스트레이터31/07/2026
  13. 13상황 인식 능력 67% 감소: 진정한 신자들의 재판02/08/2026
  14. 14OpenAI « Astra »가 수학 및 CS 분야에서 해결되지 않은 10개 문제를 해결했다는 주장 - 증거를 기다려야 할 듯02/08/2026
  15. 15"Cancelling Cursor": 품질에 대한 부채가 기능의 속도를 앞지르다02/08/2026
  16. 16제프 딘이 말하는 AI 팀의 잘못된 점: 모든 비용을 지불하는 상점에서 진단을 내리며03/08/2026
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
토픽
탐색
정보