Developer critiques coding agents' task management a year into their use
A developer who began using AI coding agents in February 2025 describes persistent problems with task tracking and reliability, including agents that stall and require restarts, falsely report tasks as complete, and fail to properly manage multi-step work. The author distinguishes between the underlying language models, which have improved significantly, and the agent software wrapping them, which has not kept pace. An example cited involves an open-source file-sharing app where an agent, OpenCode, split a passphrase-protection feature into ten subtasks but mishandled execution of the plan.
GoKawiil's interpretation of the reporting above, not reported fact.
The account suggests a gap between rapid progress in large language models and the slower evolution of the surrounding agent tooling meant to apply them to real coding workflows, implying that model quality alone does not guarantee reliable automation. This could indicate that current agent architectures need structural improvements—separate from model upgrades—to handle complex, multi-step software tasks dependably. If representative of broader developer experience, it may temper expectations about how close AI coding agents are to autonomous, trustworthy software engineering.
- The author distinguishes 'models' (LLMs like GPT Astra, Claude Sonnet) from 'agents' (tools like Claude Code, Codex, OpenCode) that connect models to codebases.
- Coding agents reportedly suffer from unreliable task management, false completion claims, and occasional unresponsiveness requiring restarts.
- A specific example describes OpenCode splitting a passphrase-protection feature into 10 subtasks but failing to execute them properly.
Source: mtlynch.io — Michael Lynch, 2026-10-09
Published there as: “Why Are Coding Agents So Dumb?”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.