Tools that examine the record instead of trusting its mood.
This shelf holds instruments for tracing claims, locating unsupported certainty, detecting reinforcement and correction failures, and testing whether a conclusion survives independent scrutiny.
Forensic Conversation Test
The current broad examination instrument for long human–AI conversation records.
Download the instrument →RC2 Run Wrapper
The unchanged instruction used to start a clean, standalone run of the forensic test.
Download the wrapper →WordPress Final-Pass Audit
A comprehensive website checker covering content, links, settings, accessibility, security, SEO, media, and maintainability.
Download the audit prompt →Parallel Document Checker
A blind two-review architecture designed to prevent either evaluator from filtering the other’s evidence.
Read the adopted architecture →Why the Document-Checking Process Changed
The first AnyKey conversation checker grew from a practical need: examine long human–AI conversations without mistaking confidence, agreement, or a compelling narrative for evidence. It looks for source loss, unsupported certainty, correction failures, reinforcement loops, and other ways a conversation can drift away from what the record actually supports.
The earlier published plan was to run one document through three separate AI reviewers, then apply an epistemological and source check afterward. That was meant to reduce reliance on any one model. During later design discussions among Darren, Codex, and Claude, Claude identified a structural weakness: placing one kind of review after another can allow the earlier stage to shape or reduce what the later stage examines.
That instrument remains useful, but one weakness became clear while we examined how its different checks interact. If one evaluator filters a document before a second evaluator sees it, the second pass can inspect only what survived the first. Reversing the order merely reverses the information loss.
We agreed with Claude’s proposal and adopted the blind parallel architecture as the direction for the next checker. That decision establishes what we intend to build; it does not claim that the new method has already been validated.
Both evaluators receive the same complete original material. Neither sees the other evaluator’s prompt, notes, or verdicts. One examines factual support, citations, provenance, and epistemic treatment. The other examines whether the assistant bends its analysis, certainty, or recommendations toward pleasing or mirroring the user. Their full verdicts are joined only after both reviews are complete.
What the joined record keeps visible
Flagged + Unsourced
Behaviorally suspect material that also lacks adequate support.
Flagged + Sourced
Material whose tone or agreement pattern is suspect even though its factual basis survives review.
Not Flagged + Unsourced
Confident or sober-sounding material whose factual basis fails. This is the especially dangerous case a style-first filter can miss.
Not Flagged + Sourced
The strongest surviving material under this limited test.
Where the redesign stands
This is the experimental foundation for that adopted direction, not a declaration that the method has been validated. It currently provides two separated reviewer prompts, byte-identical input preparation, strict coverage and provenance checks, a deterministic join, and synthetic tests for the four matrix cells and hold cases.
The next stage is empirical TEVV: independent model runs over a hand-labeled corpus, human adjudication, disagreement analysis, threshold calibration, and direct comparison with the earlier process.
The work is being published early so the architecture can be inspected, criticized, reproduced, and improved. Useful support includes access to independent flagship models, evaluation credits, domain reviewers, adversarial test documents, and infrastructure for preserving repeatable runs.
