# Perfect ranking. Missing answers.

Recorded local experiment, 30 September 2026. This is a synthetic correctness test, not a production latency or state-of-the-art benchmark.

## Question

Can retrieving the globally nearest 100 items and then filtering for eligibility reliably return the 10 nearest eligible items?

## Method

The probe generates 5,000 eight-dimensional Gaussian vectors per seed and varies eligibility across 10 seeds, seven eligible fractions, and three correlation conditions: 210 cases. Global ranking is exact, so approximate ranking errors cannot explain the failures. The target is recall@10 of at least 0.95.

| Strategy | Cases meeting the target |
| --- | ---: |
| Exact global top 100, then filter | 108 / 210 |
| Exact scan of eligible items | 210 / 210 |
| Idealized expanding global prefix | 210 / 210 |

All three strategies returned zero ineligible items. Failing recall and leaking unauthorized items are different failure modes.

## What changed

A fixed shortlist can discard the relevant eligible neighbors before filtering. These results support evaluating the eligible subset or expanding the candidate budget. They do not establish a production speed winner, performance of any ANN index, or a general research advantage over ordinary search.

The webpage replays these recorded results; it does not execute the probe. This particular experiment is a separate evaluation artifact, not a built-in evaluator shipped with the MCP runtime.

## Reproduce

Download [probe.py](probe.py) and [results.json](results.json). The probe uses Python's standard library. Run `python3 probe.py --help` for its output options. [Checksums](SHA256SUMS) cover the downloadable artifacts.

Local invocation and executable paths have been removed from the published results. The deterministic experiment data and original experiment hash are preserved.

## Comparison with ordinary research

We also compared ordinary research prompting with the original Midcenturion skill on one systems-design question. Both received 13/13 under the same qualitative rubric. The rubric allowed a detailed experiment plan instead of an executed test, so this score does not establish production readiness or superior research. Model tokens, elapsed time, and source exposure were not equalized.

After revising the skill, later runs found useful counterevidence and executed small synthetic probes. These single runs, including one related transfer case without its own ordinary-prompt baseline, do not establish a general improvement in answer quality or exhaustive discovery.

The MCP's state and acceptance rules were checked separately. Existing tests cover rejection after a failed check, required review, exact-input acceptance, invalidation after linked evidence changes, safe retries, and resuming in a fresh process. These properties make the investigation inspectable and resumable. They do not establish that supplied evidence is authentic or that a conclusion is true.
