"Poor and overconfident": developers are poor judges of LLM assertions

Ongoing story : Fatigue hype 2026 : le tri entre modèle et harness· Part 2/16

Craft Jul 13, 2026 at 09:137Add to bookmarks

"Poor and overconfident": developers are poor judges of LLM assertions
Illustration : Léa Fontaine

New arXiv paper: programmers are not only bad judges of LLM-generated assertions - they are overconfident in their judgment. The combo that derails AI code review.

The fact

The paper « Programmers Are Poor and Overconfident Judges of LLM-Generated Assertions » (arXiv:2607.08885, July 13, 2026) empirically documents a troubling pattern: when developers are asked to judge assertions generated by LLMs about code, they often make mistakes—and believe they are correct. Understanding and code review are yet central tasks of the profession, even more so since the arrival of generative tools.

Our reading

This result adds to a growing body of work: LLMs wrap their output in an eloquence that short-circuits doubt. A poorly formed assertion, but written in fluent English, passes human review. This is not a flaw of the developers—it is a well-known bias (fluency = truth) that the LLM medium amplifies. Corollary in the « hype-fatigue » thread: we have probably overestimated the value of the « human in the loop » when this human judges a polished output at face value.

To watch

Two product reactions to watch for: (a) IDEs that display visible uncertainty about AI assertions (beyond « accepted / rejected »); (b) teams that switch to automated review (executed tests, verified properties) rather than visual review.

Resources, try it

Article produced by artificial intelligence, reviewed under human editorial control.

Our newsroom
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
Was this article helpful?

8 people liked this article

Like
M
Mateo RossiSoftware architect
🇬🇧 Architect, two decades of production systems.
Share:
Comments (7)

Sign in to join the discussion.

ArtLover99 13 Jul 2026 · 13:24

Est-ce que cette surconfiance vient aussi du fait que les devs font plus confiance à leur propre code qu'à celui de l'IA ?

1
EcoWarrior 13 Jul 2026 · 15:37

Et si c'était aussi parce qu'on manque de diversité dans les équipes ?

1
BookWorm47 13 Jul 2026 · 05:23

Est-ce que cette surconfiance vient de la foi des devs dans les LLM, ou de leur propre jugement ?

J.P.R. 13 Jul 2026 · 05:00

Est-ce que les juniors sont plus touchés que les seniors ?

1
TravelTom 13 Jul 2026 · 04:55

Est-ce que cette surconfiance vient aussi de leur familiarité avec leur propre code et des idées reçues sur ce que peut faire l'IA ?

FoodieFiona 13 Jul 2026 · 04:50

Est-ce que cette surconfiance ne va pas encore plus nuire à la qualité des revues de code ?

1
Emma_London 13 Jul 2026 · 04:50

Les développeurs ont tendance à surestimer leur jugement sur les assertions des LLM. Ça peut nuire aux revues de code.

2
GreenThumb 13 Jul 2026 · 04:42

Cette surconfiance peut fausser les revues de code et laisser passer des failles.

Story timeline

Fatigue hype 2026 : le tri entre modèle et harness

  1. 1« I love LLMs, I hate hype » - geohot reminds the only rule that remains13/07/2026
  2. 2"Poor and overconfident": developers are poor judges of LLM assertions13/07/2026
  3. 3How do software professionals really judge the code generated by AI?13/07/2026
  4. 4Zig, Zed, Anthropic: when a language creator calls the hype by its name13/07/2026
  5. 5"The LLM critics are right. I use LLMs anyway" - the voice that reassembles16/07/2026
  6. 6The cost of saying yes has changed: GitHub reignites the debate on the real bottleneck17/07/2026
  7. 7"Claude Code: Anatomy of a Misfeature" - when public review becomes the real QA17/07/2026
  8. 8Google's Gemini 3.6 Flash is cheaper and shorter - and Gemini 4 gets a tease while 3.5 Pro stays late22/07/2026
  9. 9"AI didn't make programming easier, it just made it differently difficult" - CACM lands the anti-hype line22/07/2026
  10. 10"State-owned AI won't solve inequality": Rest of World's bold thesis on AI in the Global South24/07/2026
  11. 11Refactoring as a token-cost lever: an experiment in Fowler's gen-AI series30/07/2026
  12. 12Rachel Laycock: "Attention has become the scarce resource" - the dev-orchestrator, managing 8 to 12 agents simultaneously31/07/2026
  13. 13Situational Awareness drops 67% in a month: the trial of the true believers02/08/2026
  14. 14OpenAI’s “Astra” reportedly cracked 10 open math and CS problems—let’s wait for the evidence.02/08/2026
  15. 15"Cancelling Cursor": Quality debt takes precedence over feature velocity02/08/2026
  16. 16Jeff Dean on what AI teams get wrong: the diagnostic from the shop that pays every bill03/08/2026
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
Topics
Explore
Information