A blog post by Alex Ewerlof, a developer with two engineering degrees, argues that the widespread claim that “coding is solved” through large language models misunderstands fundamental software engineering challenges.

Ewerlof acknowledges that LLMs can generate code and that he has been an early adopter of AI coding tools, but contends that creation speed is only part of the equation. According to the post, the majority of software costs come from maintenance, reliability, security, and scalability—collectively known as non-functional requirements (NFRs)—which remain unsolved problems.
The author identifies three categories of software where code review is not strictly necessary: personal software and automation, proofs of concept, and deliberately weaponized AI. However, he notes that most professionally developed software—in healthcare, finance, automotive, defense, aviation, and manufacturing—operates under low risk tolerance, where mistakes can result in financial loss, loss of life, or legal consequences.
Ewerlof argues that AI systems cannot be held accountable for failures. “AI cannot suffer any consequences,” the post states. “You cannot be responsible for what you can’t control either.” This accountability gap, he contends, is crucial for understanding why LLMs struggle with production code in high-stakes industries.
The core technical limitation, according to Ewerlof, is that coding fundamentally depends on logic. LLMs are “stochastic and probabilistic,” meaning they excel at tasks related to natural language but struggle with logical rigor. While developers have created workarounds—feedback loops, testing harnesses, and techniques like chain-of-thought prompting—these are compensations for the underlying limitation. He notes that LLM accuracy degrades as input volume and context window usage increase.
Ewerlof criticizes those claiming LLM-generated code is production-ready, suggesting they either haven’t written code recently, lack quality standards, or don’t understand capability curves. He also expresses concern that companies like Google and GitHub are degrading their services by prioritizing AI velocity over reliability, citing examples from Anthropic’s Claude Code, which he says had quality issues.
The post concludes that AI has legitimate uses for proofs of concept, personal projects, and language-based tasks, but warns against forcing it into workflows where it doesn’t belong. Ewerlof frames his argument not as anti-AI but as a defense of accountability and long-term engineering quality.
Key facts
- LLMs excel at natural language tasks but struggle with logic, which is fundamental to coding
- Most software costs come from maintenance and reliability, not initial code creation
- AI systems cannot be held accountable for failures, which is problematic for high-stakes industries like healthcare and finance
- Accuracy of LLMs degrades as context window usage increases
- The author identifies legitimate uses for AI: POCs, personal software, and language-based tasks like translation
