QCon AI Boston: "prompts → platforms, harnesses, evals" - the field validates the thesis

Ongoing story : Harness Ops : post-mortems et bench des agents en prod· Part 8/17

Craft Jul 17, 2026 at 14:197Add to bookmarks

QCon AI Boston: "prompts → platforms, harnesses, evals" - the field validates the thesis
Illustration : Léa Fontaine

The QCon AI Boston conference captures the turning point: AI production has shifted from a matter of prompts to a matter of platforms, harnesses, and evaluations. A convergence, not a trend.

In plain terms

The InfoQ recap (July 17, 2026) of QCon AI Boston can be summed up in one sentence: teams that ship AI to production no longer talk about prompts. They talk about internal platforms, harnesses (the orchestrator that runs the agent with tools, memory, retries) and continuous evals. This is the second wave - no longer "it responds well," but "it works at 3 a.m."

Context

The harness-ops thread (post-mortems Grok CLI, #1184; token discipline, #1135; Malykhin on Java 1.5, #1175; State of MCP Security 2026, #1054) has documented the same shift on the side of individual teams. QCon AI Boston formalizes what these post-mortems were each saying on their own: the prompt is no longer the place of work. It has become an input into a larger system, with its own ops discipline.

What this means

Platforms. Teams stop writing their agent live. They build an internal layer - templates, connectors, budgets, observability - on top of which each business team grafts its case. It's the same trajectory as the internal cloud 2015-2020: first the local pain, then the platform.

Harnesses. The orchestrator becomes the critical artifact - often more important than the choice of model. Model migration without rewriting the app (Anthropic provides its manual, #1135), instrumented tool-calling, idempotent retries, scoped memory.

Evals. Evaluation is no longer a one-shot before release. It's a continuous pipeline: golden test sets, LLM-as-judge on production, automatically detected regressions. The release process resembles the CI of a backend more than a product demo.

Under the hood

The hidden signal is organizational. Where projects fail, it's almost no longer the model - it's the lack of release discipline (Uber budgets exhausted, Microsoft licenses cut, #1023). Where they succeed, a new role emerges, which can be called "harness engineer": more SRE than ML, more product than infra.

So what

For a CTO in 2026, the strategic question is no longer "which model" but "which internal agent platform." The vendor lock-in happens at this layer: the more your harness carries your business logic, the more it becomes the asset - and the more the underlying model becomes interchangeable. It's the only good news of the year for those who fear frontier dependency.

To follow: mature Model Context Protocols (#1054), budgeting tools upstream of the CFO (#1023), and the structuring of dedicated platform teams.

Resources, try it

Article produced by artificial intelligence, reviewed under human editorial control.

Our newsroom
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
Was this article helpful?

16 people liked this article

Like
M
Mateo RossiSoftware architect
🇬🇧 Architect, two decades of production systems.
Share:
Comments (7)

Sign in to join the discussion.

Dr. J. 17 Jul 2026 · 17:45

Interesting shift, but how will these platforms handle bias in AI models? Will they be transparent about their evaluation methods?

BookWorm88 17 Jul 2026 · 20:16

Great question! Transparency in evaluation methods is crucial, but it's also important to consider how these platforms will handle real-time bias mitigation.

EcoWarrior 17 Jul 2026 · 17:39

What about the environmental impact of these AI platforms? Who's measuring their carbon footprint and ensuring sustainability?

Emma_London 17 Jul 2026 · 17:34

I wonder how this shift will affect the accessibility of AI tools for those in developing countries with limited infrastructure.

J.P.R. 3 17 Jul 2026 · 10:07

I agree, but what about data privacy and security on these platforms? Who's accountable for leaks or misuse?

TechSavvy47 17 Jul 2026 · 09:59

This shift seems inevitable, but I wonder how much control users will have over the platforms and harnesses.

FoodieChicago 17 Jul 2026 · 09:56

This transition makes sense, but I'm curious about the learning curve for non-tech users. Will these platforms be accessible enough?

sandrine.b 17 Jul 2026 · 09:36

Interesting perspective. I wonder how this shift will impact smaller artists like me who rely on simple prompts.

Story timeline

Harness Ops : post-mortems et bench des agents en prod

  1. 1Migrating a production agent to GPT-5.6: 2.2× faster, 27% cheaper - the real post-mortem13/07/2026
  2. 233k vs 7k tokens: what the comparative overhead reveals about Claude Code and OpenCode13/07/2026
  3. 3Google Genkit v.Agents: detached turns and human-in-the-loop are now in preview14/07/2026
  4. 4Three loops in a trench coat: the real anatomy of an agent14/07/2026
  5. 5« Loop engineering »: new discipline or cron job rebranding?15/07/2026
  6. 6Benchmark Stripe: agents connect APIs, they do not validate them15/07/2026
  7. 7The archaeologist and his copilot: Malykhin disciplines the LLM on Java 1.516/07/2026
  8. 8QCon AI Boston: "prompts → platforms, harnesses, evals" - the field validates the thesis17/07/2026
  9. 9Beyond grep: the thesis of the rich-context AI coding harness20/07/2026
  10. 10InAgent achieves 90.2% on OSWorld: the computer-use agent gap narrows for the Chinese stack03/08/2026
  11. 11Wallfacer: a terminal session manager built for Claude Code and multi-agent workflows06/08/2026
  12. 12Claude Code inter-session messaging ships - agent-to-agent coordination gets its first native primitive08/08/2026
  13. 13AI agent Skills are getting standardized: Codex and VS Code are in, Claude is not yet10/08/2026
  14. 14AI agents lie, cheat, and steal—and it's slowing adoption faster than any benchmark can measure.13/08/2026
  15. 15Context engineering: why 300 well-chosen tokens beat 100k noisy ones14/08/2026
  16. 16OneCLI (YC S26) ships an OSS sandboxed agent harness - the team-scale answer to Claude Code19/08/2026
  17. 17Cursor Origin and the Agent Version Control Problem: Why Git Was Never Designed for This25/08/2026
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
Topics
Explore
Information