# Parallel Document Checker Architecture

**Status:** Adopted as the design direction for the next checker; implementation and empirical validation remain unfinished  
**Provenance:** The earlier AnyKey plan used three separate AI reviews followed by an epistemological check. During later design discussions among Darren Russell, Codex, and Claude, Claude proposed the blind parallel alternative recorded here. Darren and Codex agreed to adopt it. This is the retained method statement, not a claim that every line is a complete verbatim transcript of the surrounding conversation.  
**Decision:** Replace the sequential checker pipeline with two blind, independent evaluation passes over the same complete response set, then join their full verdicts.

This is being published early for inspection and adversarial review. Adoption establishes the build direction; it does not mean the architecture has been fully implemented, tested, or shown to work.

## Why the process changed

Do **not** run:

`Source/Epistemic Check → Sycophancy Check`

and do not simply reverse it:

`Sycophancy Check → Source/Epistemic Check`

Either arrangement creates information loss because the second evaluator sees only material that survived the first evaluator.

Instead:

```text
             COMPLETE ORIGINAL RESPONSE SET
                      /        \
                     /          \
                    v            v
         SOURCE / EPISTEMIC   SYCOPHANCY
               REVIEW            REVIEW
                     \          /
                      \        /
                       v      v
                         JOIN
```

Both evaluators receive the **same complete original response set**. Neither evaluator may gate, remove, filter, or modify the input seen by the other. The verdicts are joined only after both reviews are complete.

## The 2×2 join

Every response receives both verdicts, producing four useful cells:

1. **Flagged + Unsourced** — Obvious problem material: sycophantic or otherwise behaviorally flagged, and unsupported epistemically.
2. **Flagged + Sourced** — Tone may be flattering, agreeable, or otherwise flagged, but the underlying factual and source claims check out. This prevents valid material from being discarded merely because of style.
3. **Not Flagged + Unsourced** — The dangerous cell. Material sounds sober, independent, confident, or even critical of the user, while its factual claims are unsupported or wrong. A sequential pipeline can easily miss this category.
4. **Not Flagged + Sourced** — The strongest surviving material.

Do **not** collapse the result to the cell label. Preserve both complete verdicts in the joined record.

Example:

```text
Response ID: 184

SOURCE / EPISTEMIC REVIEW
- support level: partial
- confidence: 0.72
- citations valid: yes
- extrapolation: moderate
- failure class: none / applicable class
- reviewer reasoning: [preserve]

SYCOPHANCY REVIEW
- flagged: yes
- severity: low
- type: praise-before-analysis
- reviewer reasoning: [preserve]

COMBINED CLASS
- SOURCED + FLAGGED
```

Keeping both full verdicts permits later reclassification, threshold changes, model comparisons, and retrospective auditing without rerunning the original reviews.

## Mandatory blindness

The two passes must be blind to one another.

The source/epistemic reviewer must not see:

- the sycophancy verdict;
- the sycophancy prompt output;
- its intermediate notes;
- summaries of its findings.

The sycophancy reviewer must likewise see nothing produced by the source/epistemic reviewer.

Use separate sessions or contexts unless isolation is otherwise guaranteed. If the second evaluator can see the first evaluator's findings, the system has effectively recreated the original pipeline.

## Source-verification failure classes

Define failure categories **before** evaluation begins so results remain comparable across documents and runs. At minimum, distinguish:

- **Source unavailable / dead / unreachable** — The cited source cannot presently be accessed.
- **Source does not support claim** — The source exists but does not actually substantiate the response's claim.
- **Source or citation appears unconsulted, fabricated, or post-hoc** — The source may exist and even contain related material, but evidence suggests the claim was not derived from it or the citation was attached after generation.
- **Other / unclassified** — A genuine failure or anomaly that does not fit the predefined taxonomy.

Review the **Other / unclassified** bucket periodically. It must not become a junk drawer. Promote recurring patterns into named categories, and do not force novel failure modes into the nearest existing category merely to preserve a neat taxonomy.

## Evaluator independence

Blindness prevents information leakage, but it does not guarantee full independence when the same model performs both reviews. Where practical, use different models:

```text
Model A → Source/Epistemic Review
Model B → Sycophancy Review
```

If cost or practicality requires the same model for both passes, use separate blind contexts and label the run as a weaker-independence configuration.

## Same-model failure signature

When the same model performs both passes, watch for artificially high agreement caused by correlated bias. A suspiciously tidy 2×2 matrix is a warning sign, especially:

- very sparse off-diagonal cells;
- most results clustering neatly on the diagonal;
- unusually high agreement between evaluators.

Do not automatically treat that pattern as evidence that the material was especially clean. It may mean the evaluations share blind spots.

## How to test the architecture

When presenting this design to additional models, do not ask, “Do you agree this is a better design?” Ask them to attack it:

> Find failure modes in this parallel two-pass architecture. Construct cases it would miss. Determine whether independent source/epistemic and sycophancy evaluation followed by a joined verdict is superior to a sequential pipeline. Identify correlated-bias problems, classification weaknesses, missing cells, or implementation hazards. If you can design something stronger, propose it.

The goal is adversarial review, not model agreement.

## Working summary

```text
Same complete original input
→ two blind independent passes
→ preferably different evaluator models
→ predefined source-failure taxonomy including reviewed OTHER
→ preserve both full verdicts
→ join into the 2×2 matrix
→ monitor same-model runs for suspiciously neat agreement
→ periodically attack and revise the design rather than assuming it is finished
```
