HomeAI NewsSimon Willison maps 2026 LLM trends from Claude Opus 4.5 to coding...

Simon Willison maps 2026 LLM trends from Claude Opus 4.5 to coding agents

The talk traces a year of model releases that turned coding agents from error-prone into daily tools.

Simon Willison gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose, walking through every major large language model development of 2026 in chronological order. He published annotated slides and notes alongside the recorded talk.

Willison traces the year back to November 2025, when Anthropic released Claude Opus 4.5 and OpenAI shipped GPT-5.1. Both were incremental upgrades, but pairing them with coding agent harnesses like Claude Code and Codex moved those tools from often wrong to dependable for daily work.

For builders, the practical threshold shifted from whether an agent can write code to whether its output holds up across routine tasks. Teams that track reliability rather than benchmark scores will decide where these agents earn a place in production workflows.

Willison also flags the first commit to an obscure GitHub repository named Warelay and a personal resolution to take on more projects now that agents can carry the load. He expects to know by the end of the year whether that bet paid off.

What matters

  • Claude Opus 4.5 and GPT-5.1 shipped in late 2025 with coding agents that finally worked reliably.
  • Builders should treat agent reliability as the metric that decides which tools reach production.
  • Watch whether the Warelay repository and rising agent autonomy produce verifiable results by year end.

Why it matters

Watch whether the Warelay repository and rising agent autonomy produce verifiable results by year end.

This GenAI News article was prepared in original wording using reporting and materials published by Simon Willison’s Weblog. Source reference: https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/.

Drafted by the GenAI News review pipeline.

latest articles

explore more